Gene copy number variation type detection method and related product

By acquiring and analyzing sequencing comparison data and calculating the corrected gene copy number of tumor cells, the problem that the NGS detection method cannot accurately identify homozygous gene deletions is solved, and accurate detection of gene copy number variation types is achieved.

CN120748484AActive Publication Date: 2025-10-03GUANGZHOU BURNING ROCK DX CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511142662.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-10-03
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing NGS detection methods cannot accurately assess the absolute copy number of genes in tumor cells, cannot identify homozygous gene deletions, and thus cannot accurately identify copy number variations of tumor suppressor genes.

Method used

By obtaining sequencing comparison data of the test sample and the baseline sample, the tumor purity and tumor ploidy are determined, and the correction formula is used to calculate the corrected tumor cell copy number of the test gene in the test sample to generate the detection result of the gene copy number variation type.

Benefits of technology

It achieves accurate assessment of gene copy number in tumor cells, can identify homozygous gene deletions and amplifications, and improves the accuracy and reliability of gene variation detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748484A_ABST
    Figure CN120748484A_ABST
Patent Text Reader

Abstract

The invention provides a gene copy number variation type detection method and a related product. According to a specific embodiment of the gene copy number variation type detection method, in the process of determining the corrected tumor cell copy number of the to-be-detected gene in the to-be-detected sample, the tumor ploidy of the to-be-detected sample is additionally considered; therefore, the accuracy of the corrected tumor cell copy number of the to-be-detected gene in the to-be-detected sample can be improved. Therefore, the accuracy of detecting the copy number variation state of the to-be-detected gene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of bioinformatics analysis technology, and specifically to a method for detecting gene copy number variation types and related products. Background Art

[0002] The healthy human genome consists of 23 pairs of chromosomes, each containing two copies, one from each parent, corresponding to diploidy. During the development and progression of tumors, genomic rearrangements can cause changes in the number of gene copies, such as copy number loss or copy number gain.

[0003] Homozygous copy number loss is a type of copy number loss in which both copies of a gene are lost. For tumor suppressor genes, such as TP53, RB1, PTEN, and MTAP, when both alleles are present, the protein is expressed normally, suppressing the development of malignant tumors. However, when homozygous or heterozygous loss occurs, the gene becomes inactive and no longer exerts its inhibitory effect, leading to cell transformation into cancer cells. Therefore, the detection of homozygous copy number loss in tumor suppressor genes as a tumor marker is crucial.

[0004] The significance of homozygous deletion of the tumor suppressor gene PTEN is illustrated by an example. PTEN, known in Chinese as phosphatase deleted on human chromosome 10, is a tumor suppressor gene located on the long arm of chromosome 10 (10q23.3). It comprises nine exons and is a phosphokinase with dual-specific phosphatase activity. PTEN is located upstream of the PIK3-AKT-mTOR pathway and acts as a negative regulator of this pathway. It participates in cellular regulation through dephosphorylation. Loss of PTEN function leads to pathway activation and drives tumor growth. Various types of PTEN mutations have been identified, with PTEN deletion being the most common, prevalent in breast and prostate cancers. PTEN mutations or deletions often result in inactivation or failure of PTEN protein expression, leading to activation of the PIK3-AKT-mTOR signaling pathway and enhanced cell proliferation, survival, and migration. Patients with PTEN mutations or deletions generally have a poor prognosis. Existing evidence also suggests that PTEN gene status is a potential predictive biomarker for therapeutic efficacy, highlighting the significant clinical value of PTEN gene testing.

[0005] Currently, the detection methods for gene copy number variations (including deletions and amplifications) are mainly based on methodologies such as IHC (ImmunoHistoChemistry), FISH (Fluorescence In Situ Hybridization), and NGS (Next Generation Sequencing, also known as high-throughput sequencing (HTS)).

[0006] IHC, a common clinical pathology test, is often used to detect protein expression. However, standardization is difficult, and test results are susceptible to human and subjective factors, especially in interpretation. Furthermore, IHC can only qualitatively determine protein expression and cannot provide information on specific gene mutations.

[0007] Compared to IHC, FISH offers the advantage of quantitative detection. However, it also faces challenges such as difficulty in standardization and susceptibility to human error. Furthermore, FISH is more complex to perform, requiring probe design and synthesis, and specialized instrumentation for fluorescence signal detection. Furthermore, FISH can only detect known mutations, not unknown ones. These factors pose significant challenges to the clinical application of FISH.

[0008] NGS sequencing technology overcomes the shortcomings of IHC and FISH, and can accurately, quantitatively and comprehensively detect all gene mutation information. At the same time, the detection throughput is greatly improved, and the operation methods and data analysis are standardized and streamlined, avoiding interference from human factors.

[0009] Given the diversity and complexity of genetic variation, NGS testing offers significant advantages. However, existing NGS methods generally assess gene copy number variation based on sequencing depth. This approach can only identify mixed copy numbers of normal and tumor cells in tumor samples, but cannot accurately assess the absolute copy number of genes in tumor cells, and consequently, cannot accurately identify homozygous gene deletions. Summary of the Invention

[0010] The embodiments of the present disclosure provide a method, apparatus, electronic device, storage medium, and computer program product for detecting gene copy number variation types.

[0011] In a first aspect, embodiments of the present disclosure provide a method for detecting a gene copy number variation type, the method comprising: Acquire sequencing comparison data of gene sequencing performed on the test sample and at least one baseline sample, wherein each sequencing comparison data includes sequencing comparison data of a preset set of analysis regions of the test gene, and the preset set of analysis regions of the test gene includes the region where the test gene is located; Determining the tumor purity and tumor ploidy of the sample to be tested based on the sequencing comparison data; Determining the average mixed copy number of the gene to be tested in the sample to be tested based on each of the sequencing comparison data; Determining the corrected tumor cell copy number of the gene to be tested in the test sample based on the tumor purity and tumor ploidy of the test sample and the average mixed copy number of the gene to be tested in the test sample; In response to the average mixed copy number of the gene to be tested in the test sample being less than 2, generating a detection result for the homozygous deletion of the gene to be tested in the test sample based on a comparison result of the corrected tumor cell copy number of the gene to be tested in the test sample and a positive judgment threshold for homozygous deletion of the gene to be tested; and / or, In response to the average mixed copy number of the gene to be tested in the sample to be tested being greater than 2, an amplification judgment value of the gene to be tested in the sample to be tested is determined according to the corrected tumor cell copy number of the gene to be tested in the sample to be tested, and a detection result of the amplification of the copy number of the gene to be tested in the sample to be tested is generated according to a comparison result of the amplification judgment value of the gene to be tested in the sample to be tested and an amplification positive judgment threshold of the gene to be tested.

[0012] In some optional embodiments, determining the corrected tumor cell copy number of the gene to be tested in the test sample based on the tumor purity and tumor ploidy of the test sample and the average mixed copy number of the gene to be tested in the test sample comprises: The tumor purity and tumor ploidy of the sample to be tested, and the average mixed copy number of the gene to be tested in the sample to be tested are substituted into the corrected tumor cell copy number solution to obtain the corrected tumor cell copy number of the gene to be tested.

[0013] In some optional embodiments, the formula for calculating the corrected tumor cell copy number is obtained by deriving a correction formula, wherein the correction formula is: ; The formula for calculating the corrected tumor cell copy number is: ; in, and are respectively the tumor purity and tumor ploidy of the sample to be tested, is the average mixed copy number of the gene to be tested in the sample to be tested, is the corrected tumor cell copy number of the gene to be tested in the sample to be tested.

[0014] In some optional embodiments, determining the amplification judgment value of the gene to be tested in the test sample according to the corrected tumor cell copy number of the gene to be tested in the test sample includes: Determining the corrected tumor cell copy number of the gene to be tested in the sample to be tested as the amplification judgment value of the gene to be tested in the sample to be tested, or, The difference between the corrected tumor cell copy number of the gene to be tested in the sample to be tested and the tumor ploidy of the sample to be tested is determined as the amplification judgment value of the gene to be tested in the sample to be tested, or, The amplification judgment value of the gene to be tested in the sample to be tested is determined as the ratio of the corrected tumor cell copy number of the gene to be tested in the sample to be tested divided by the tumor ploidy of the sample to be tested.

[0015] In some optional embodiments, determining the tumor purity and tumor ploidy of the test sample based on each sequencing comparison data includes: Determine the abundance of each SNP site in the preset whole-genome SNP site set of the sample to be tested based on the sequencing comparison data of the sample to be tested; For each predetermined whole-genome SNP site, the related gene analysis region is divided into at least one SNP analysis region according to a first predetermined length; For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each SNP analysis region; For each SNP analysis region, the mixed copy number of the sequencing alignment data of the test sample in the SNP analysis region is determined based on the corrected average coverage depth of the sequencing alignment data of the test sample in the SNP analysis region and the average of the corrected average coverage depth of the sequencing alignment data of each baseline sample in the SNP analysis region; Based on the abundance and mixed copy number of each SNP site in the sequencing comparison data of the sample to be tested, the tumor purity and tumor ploidy of the sample to be tested are determined, wherein the mixed copy number of each SNP site is the mixed copy number of the SNP analysis region where the SNP site is located.

[0016] In some optional embodiments, determining the average mixed copy number of the gene to be tested in the test sample based on each of the sequencing comparison data includes: For each preset gene analysis region to be tested, dividing the preset gene analysis region to be tested into at least one gene analysis window according to a second preset length; For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each gene analysis window; For each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, determine the reference baseline sample sequencing alignment data corresponding to the gene analysis window in the sequencing alignment data of each baseline sample, as well as the corrected average baseline sample sequencing depth and the corrected average baseline sample sequencing depth deviation of the gene analysis window; Determining a reference gene analysis window in each gene analysis window based on the corrected baseline sample average sequencing depth deviation of each gene analysis window; Determine each reference gene analysis window where the gene to be tested is located as the gene analysis window to be tested; For each gene analysis window to be tested, determining the mixed copy number of the gene analysis window to be tested in the sample to be tested based on the corrected average coverage depth of the sequencing alignment data of the sample to be tested in the gene analysis window to be tested and the corrected average sequencing depth of the baseline sample in the gene analysis window to be tested; The mean of the mixed copy numbers of each analysis window of the gene to be tested in the sample to be tested is determined as the average mixed copy number of the gene to be tested in the sample to be tested.

[0017] In some optional embodiments, for each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, determining the reference baseline sample sequencing alignment data corresponding to the gene analysis window, as well as the corrected baseline sample average sequencing depth and the corrected baseline sample average sequencing depth deviation of the gene analysis window in the sequencing alignment data of each baseline sample, includes: For each gene analysis window, perform the following baseline sample gene analysis window depth correction operation: Determine the normal range of the average sequencing depth of the baseline samples based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window; Determine the sequencing alignment data of each baseline sample whose average coverage depth after correction in the gene analysis window falls within the normal range of the average sequencing depth of the baseline samples as the reference baseline sample sequencing alignment data corresponding to the gene analysis window; Based on the corrected average coverage depth of the sequencing alignment data of each reference baseline sample in the gene analysis window, the corrected average sequencing depth of the baseline samples and the corrected average sequencing depth deviation of the baseline samples in the gene analysis window are determined.

[0018] In some optional embodiments, determining, for each sequencing alignment data, the corrected average coverage depth of the sequencing alignment data for each gene analysis window comprises: For each sequencing alignment data, each gene analysis window is used as the target analysis region, and the following target analysis region sequencing depth correction operation is performed to obtain the corrected average coverage depth of each gene analysis window. The target analysis region sequencing depth correction operation includes: The sequencing depth of each base site in each target analysis region is normalized according to the total sequencing depth of the sequencing alignment data to obtain the corresponding normalized sequencing depth; For each target analysis region, determining the average sequencing depth of the sequencing alignment data in the target analysis region based on the normalized sequencing depth of each base position in the target analysis region; Performing a local weighted regression correction on the average coverage depth of each target analysis region of the sequencing alignment data according to the number of probes covered in each target analysis region; According to the GC content ratio of the sequencing alignment data in each target analysis region, a local weighted regression correction is performed on the average coverage depth of the sequencing alignment data in each target analysis region to obtain the corrected average coverage depth of the corresponding target analysis region.

[0019] In some optional embodiments, determining, for each sequencing alignment data, the corrected average coverage depth of the sequencing alignment data for each SNP analysis region comprises: For each sequencing alignment data, each SNP analysis region is used as a target analysis region, and a sequencing depth correction operation of the target analysis region is performed to obtain the corrected average coverage depth of each SNP analysis region.

[0020] In some optional embodiments, before determining the tumor purity and tumor ploidy of the test sample based on each sequencing comparison data, the method further includes: For each sequencing alignment data, the following base site sequencing depth normalization operation is performed: for each preset gene analysis region to be tested, the sequencing depth of each base site of the sequencing alignment data in the preset gene analysis region to be tested is determined; the sum of the sequencing depths of all base sites of the sequencing alignment data in each of the preset gene analysis regions to be tested is determined as the total sequencing depth of the sequencing alignment data; for each base site in each preset gene analysis region to be tested, the sequencing depth of the base site is normalized based on the total sequencing depth of the sequencing alignment data.

[0021] In some optional embodiments, the homozygous deletion positive judgment threshold of the gene to be tested is predetermined by the following homozygous deletion positive judgment threshold determination step: Obtaining a homozygous deletion sample data set, wherein the homozygous deletion samples include negative sample data and homozygous deletion positive sample data for the gene to be tested, wherein each negative sample data is obtained based on sequencing comparison data of a benign sample, and each homozygous deletion positive sample data is sequencing comparison data of a homozygous deletion occurring in the gene to be tested; Determining the corrected tumor cell copy number of the gene to be tested for each homozygous deletion sample data; For each candidate homozygous deletion copy number in the candidate homozygous deletion copy number set, performing the following homozygous deletion positive judgment basis determination operation: determining a positive coincidence rate corresponding to the candidate homozygous deletion copy number based on a comparison result of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each of the homozygous deletion positive sample data; and / or determining a positive predictive value corresponding to the candidate homozygous deletion copy number based on a comparison result of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each of the homozygous deletion sample data; The positive judgment threshold for homozygous deletion of the gene to be tested is determined based on the positive coincidence rate and / or positive predictive value corresponding to each candidate homozygous deletion copy number.

[0022] In some optional embodiments, obtaining a homozygous deletion sample data set includes: Obtaining sequencing alignment data of N pairs of 100% tumor cell lines and their paired cell lines, wherein N is a positive integer; Based on the N pairs of sequencing comparison data, generating at least one homozygous deletion positive sample data for the gene to be tested, wherein the tumor purity of each homozygous deletion positive sample data is the corresponding annotated tumor purity; Sequencing comparison data of M benign samples are obtained, and at least one negative sample data is generated based on the sequencing comparison data of the M benign samples.

[0023] In some optional embodiments, determining the corrected tumor cell copy number of the gene to be tested for each homozygous deletion sample data includes: For each homozygous deletion sample data, determine the tumor purity and tumor ploidy corresponding to the homozygous deletion sample data and the average mixed copy number for the gene to be tested; based on the tumor purity and tumor ploidy corresponding to the homozygous deletion sample data and the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the homozygous deletion sample data for the gene to be tested.

[0024] In some optional embodiments, the comparison result of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number based on the homozygous deletion positive sample data, and determining the positive coincidence rate corresponding to the candidate homozygous deletion copy number, includes: According to the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number of each homozygous deletion positive sample data, the number of true positives and the number of false negatives corresponding to the candidate homozygous deletion copy number are counted; The positive coincidence rate corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and the number of false negatives corresponding to the candidate homozygous deletion copy number.

[0025] In some optional embodiments, determining the positive predictive value corresponding to the candidate homozygous deletion copy number based on the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number in each of the homozygous deletion sample data includes: According to the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number of each homozygous deletion positive sample data, the number of true positives corresponding to the candidate homozygous deletion copy number is counted; According to the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number in each of the negative sample data, counting the number of false positives corresponding to the candidate homozygous deletion copy number; and A positive predictive value corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and the number of false positives corresponding to the candidate homozygous deletion copy number.

[0026] In some optional embodiments, generating at least one negative sample data based on the sequencing comparison data of the M benign samples includes: For each sequencing alignment data of a benign sample, different degrees of noise are added to the sequencing depth of the base sites in the sequencing alignment data of the benign sample to obtain at least one negative sample data corresponding to the sequencing alignment data of the benign sample.

[0027] In some optional embodiments, the amplification positive judgment threshold of the gene to be tested is predetermined by the following amplification positive judgment threshold determination step: Obtaining a set of amplified sample data, wherein the amplified sample data includes negative sample data and amplified positive sample data for the gene to be tested, wherein each negative sample data is obtained based on sequencing comparison data of a benign sample, and each amplified positive sample data is sequencing comparison data for copy number amplification of the gene to be tested; Determine the corrected tumor cell copy number for each amplified sample data; Determining the amplification judgment value of the corresponding amplified sample data for the gene to be tested according to the corrected tumor cell copy number of each amplified sample data; For each candidate amplification positive judgment value in the candidate amplification positive judgment value set, performing the following amplification positive judgment basis determination operation: determining a positive coincidence rate corresponding to the candidate amplification positive judgment value based on a comparison result of the amplification judgment value of each amplification positive sample data for the gene to be tested and the candidate amplification positive judgment value; and / or determining a positive predictive value corresponding to the candidate amplification positive judgment value based on a comparison result of the amplification judgment value of each amplification sample data for the gene to be tested and the candidate amplification positive judgment value; The amplification positive judgment threshold of the gene to be tested is determined according to the positive coincidence rate and / or positive predictive value corresponding to each candidate amplification positive judgment value.

[0028] In some optional embodiments, obtaining the amplified sample data set includes: Obtaining sequencing alignment data of I pairs of 100% tumor cell lines and their paired cell lines, where I is a positive integer; Based on the I pair of sequencing comparison data, generating at least one amplification-positive sample data for the gene to be tested, wherein the tumor purity of each amplification-positive sample data is the corresponding annotated tumor purity; Sequencing comparison data of J benign samples are obtained, and at least one negative sample data is generated based on the sequencing comparison data of the J benign samples.

[0029] In some optional embodiments, determining the corrected tumor cell copy number of each amplified sample data includes: For each amplified sample data, determine the tumor purity and tumor ploidy corresponding to the amplified sample data and the average mixed copy number for the gene to be tested; based on the tumor purity and tumor ploidy corresponding to the amplified sample data and the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the gene to be tested in the amplified sample data.

[0030] In some optional embodiments, the determining of the positive coincidence rate corresponding to the candidate amplification positive judgment value based on the comparison result of the amplification judgment value of the gene to be tested with the candidate amplification positive judgment value of each amplification positive sample data includes: According to the comparison result of the amplification judgment value of each amplification-positive sample data for the gene to be tested and the candidate amplification-positive judgment value, counting the number of true positives and the number of false negatives corresponding to the candidate amplification-positive judgment value; The positive coincidence rate corresponding to the candidate amplification positive judgment value is determined based on the number of true positives and the number of false negatives corresponding to the candidate amplification positive judgment value.

[0031] In some optional embodiments, the determining of a positive predictive value corresponding to the candidate amplification positive judgment value based on a comparison result of the amplification judgment value of the gene to be tested with the candidate amplification positive judgment value of each amplification sample data includes: According to the comparison result of the amplification judgment value of each amplification-positive sample data for the gene to be tested and the candidate amplification-positive judgment value, counting the number of true positives corresponding to the candidate amplification-positive judgment value; According to the comparison result of the amplification judgment value of each negative sample data for the gene to be tested and the candidate amplification positive judgment value, counting the number of false positives corresponding to the candidate amplification positive judgment value; and The positive predictive value corresponding to the candidate amplification positive judgment value is determined based on the number of true positives and the number of false positives corresponding to the candidate amplification positive judgment value.

[0032] In some optional embodiments, determining the amplification judgment value of the corresponding amplified sample data for the gene to be tested based on the corrected tumor cell copy number of each amplified sample data includes: Determine the corrected tumor cell copy number of each amplified sample data for the gene to be tested as the amplification judgment value of the amplified sample data, or, The difference between the corrected tumor cell copy number of each amplified sample data for the gene to be tested and the tumor ploidy of the amplified sample data is determined as the amplification judgment value of the amplified sample data, or The amplification judgment value of the amplified sample data is determined by dividing the ratio of the corrected tumor cell copy number of the gene to be tested for each amplified sample data by the tumor ploidy of the amplified sample data.

[0033] In a second aspect, an embodiment of the present disclosure provides a device for detecting a gene copy number variation type, the device comprising: A sequencing comparison data acquisition module is configured to acquire sequencing comparison data of gene sequencing performed on the test sample and at least one baseline sample, wherein each sequencing comparison data includes sequencing comparison data of a preset set of gene analysis regions to be tested, and the preset set of gene analysis regions to be tested includes the region where the gene to be tested is located; a purity and ploidy determination module, configured to determine the tumor purity and tumor ploidy of the sample to be tested based on each of the sequencing comparison data; an average mixed copy number determination module, configured to determine the average mixed copy number of the gene to be tested in the sample to be tested based on each of the sequencing comparison data; a tumor cell copy number correction module, configured to determine the corrected tumor cell copy number of the gene to be tested in the sample to be tested based on the tumor purity and tumor ploidy of the sample to be tested and the average mixed copy number of the gene to be tested in the sample to be tested; A homozygous deletion detection module is configured to generate a detection result for a homozygous deletion of the gene to be tested in the test sample in response to an average mixed copy number of the gene to be tested in the test sample being less than 2, based on a comparison result of the corrected tumor cell copy number of the gene to be tested in the test sample with a homozygous deletion positive judgment threshold of a maximum value of the homozygous deletion positive copy number of the gene to be tested; and / or, The copy number amplification detection module is configured to determine, in response to an average mixed copy number of the gene to be tested in the sample to be tested being greater than 2, an amplification judgment value of the gene to be tested in the sample to be tested based on the corrected tumor cell copy number of the gene to be tested in the sample to be tested, and generate a detection result of the copy number amplification of the gene to be tested in the sample to be tested based on a comparison result of the amplification judgment value of the gene to be tested in the sample to be tested with an amplification positive judgment threshold value, which is a minimum value of the amplification positive judgment value of the gene to be tested.

[0034] In some optional embodiments, the tumor cell copy number correction module is further configured to: The tumor purity and tumor ploidy of the sample to be tested, and the average mixed copy number of the gene to be tested in the sample to be tested are substituted into the corrected tumor cell copy number solution to obtain the corrected tumor cell copy number of the gene to be tested.

[0035] In some optional embodiments, the formula for calculating the corrected tumor cell copy number is obtained by deriving a correction formula, wherein the correction formula is: ; The formula for calculating the corrected tumor cell copy number is: ; in, and are respectively the tumor purity and tumor ploidy of the sample to be tested, is the average mixed copy number of the gene to be tested in the sample to be tested, is the corrected tumor cell copy number of the gene to be tested in the sample to be tested.

[0036] In some optional embodiments, determining the amplification judgment value of the gene to be tested in the test sample according to the corrected tumor cell copy number of the gene to be tested in the test sample includes: Determining the corrected tumor cell copy number of the gene to be tested in the sample to be tested as the amplification judgment value of the gene to be tested in the sample to be tested, or, The difference between the corrected tumor cell copy number of the gene to be tested in the sample to be tested and the tumor ploidy of the sample to be tested is determined as the amplification judgment value of the gene to be tested in the sample to be tested, or, The amplification judgment value of the gene to be tested in the sample to be tested is determined as the ratio of the corrected tumor cell copy number of the gene to be tested in the sample to be tested divided by the tumor ploidy of the sample to be tested.

[0037] In some optional embodiments, the purity and ploidy determination module is further configured to: Determine the abundance of each SNP site in the preset whole-genome SNP site set of the sample to be tested based on the sequencing comparison data of the sample to be tested; For each predetermined whole-genome SNP site, the related gene analysis region is divided into at least one SNP analysis region according to a first predetermined length; For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each SNP analysis region; For each SNP analysis region, the mixed copy number of the sequencing alignment data of the test sample in the SNP analysis region is determined based on the corrected average coverage depth of the sequencing alignment data of the test sample in the SNP analysis region and the average of the corrected average coverage depth of the sequencing alignment data of each baseline sample in the SNP analysis region; Based on the abundance and mixed copy number of each SNP site in the sequencing comparison data of the sample to be tested, the tumor purity and tumor ploidy of the sample to be tested are determined, wherein the mixed copy number of each SNP site is the mixed copy number of the SNP analysis region where the SNP site is located.

[0038] In some optional embodiments, the average mixed copy number determination module is further configured to: For each preset gene analysis region to be tested, dividing the preset gene analysis region to be tested into at least one gene analysis window according to a second preset length; For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each gene analysis window; For each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, determine the reference baseline sample sequencing alignment data corresponding to the gene analysis window in the sequencing alignment data of each baseline sample, as well as the corrected average baseline sample sequencing depth and the corrected average baseline sample sequencing depth deviation of the gene analysis window; Determining a reference gene analysis window in each gene analysis window based on the corrected baseline sample average sequencing depth deviation of each gene analysis window; Determine each reference gene analysis window where the gene to be tested is located as the gene analysis window to be tested; For each gene analysis window to be tested, determining the mixed copy number of the gene analysis window to be tested in the sample to be tested based on the corrected average coverage depth of the sequencing alignment data of the sample to be tested in the gene analysis window to be tested and the corrected average sequencing depth of the baseline sample in the gene analysis window to be tested; The mean of the mixed copy numbers of each analysis window of the gene to be tested in the sample to be tested is determined as the average mixed copy number of the gene to be tested in the sample to be tested.

[0039] In some optional embodiments, for each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, determining the reference baseline sample sequencing alignment data corresponding to the gene analysis window, as well as the corrected baseline sample average sequencing depth and the corrected baseline sample average sequencing depth deviation of the gene analysis window in the sequencing alignment data of each baseline sample, includes: For each gene analysis window, perform the following baseline sample gene analysis window depth correction operation: Determine the normal range of the average sequencing depth of the baseline samples based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window; Determine the sequencing alignment data of each baseline sample whose average coverage depth after correction in the gene analysis window falls within the normal range of the average sequencing depth of the baseline samples as the reference baseline sample sequencing alignment data corresponding to the gene analysis window; Based on the corrected average coverage depth of the sequencing alignment data of each reference baseline sample in the gene analysis window, the corrected average sequencing depth of the baseline samples and the corrected average sequencing depth deviation of the baseline samples in the gene analysis window are determined.

[0040] In some optional embodiments, determining, for each sequencing alignment data, the corrected average coverage depth of the sequencing alignment data for each gene analysis window comprises: For each sequencing alignment data, each gene analysis window is used as the target analysis region, and the following target analysis region sequencing depth correction operation is performed to obtain the corrected average coverage depth of each gene analysis window. The target analysis region sequencing depth correction operation includes: The sequencing depth of each base site in each target analysis region is normalized according to the total sequencing depth of the sequencing alignment data to obtain the corresponding normalized sequencing depth; For each target analysis region, determining the average sequencing depth of the sequencing alignment data in the target analysis region based on the normalized sequencing depth of each base position in the target analysis region; Performing a local weighted regression correction on the average coverage depth of each target analysis region of the sequencing alignment data according to the number of probes covered in each target analysis region; According to the GC content ratio of the sequencing alignment data in each target analysis region, a local weighted regression correction is performed on the average coverage depth of the sequencing alignment data in each target analysis region to obtain the corrected average coverage depth of the corresponding target analysis region.

[0041] In some optional embodiments, determining, for each sequencing alignment data, the corrected average coverage depth of the sequencing alignment data for each SNP analysis region comprises: For each sequencing alignment data, each SNP analysis region is used as a target analysis region, and a sequencing depth correction operation of the target analysis region is performed to obtain the corrected average coverage depth of each SNP analysis region.

[0042] In some optional embodiments, the apparatus further comprises a normalization module configured to, before determining the tumor purity and tumor ploidy of the test sample based on each sequencing comparison data: For each sequencing alignment data, the following base site sequencing depth normalization operation is performed: for each preset gene analysis region to be tested, the sequencing depth of each base site of the sequencing alignment data in the preset gene analysis region to be tested is determined; the sum of the sequencing depths of all base sites of the sequencing alignment data in each of the preset gene analysis regions to be tested is determined as the total sequencing depth of the sequencing alignment data; for each base site in each preset gene analysis region to be tested, the sequencing depth of the base site is normalized based on the total sequencing depth of the sequencing alignment data.

[0043] In some optional embodiments, the homozygous deletion positive judgment threshold of the gene to be tested is predetermined by the following homozygous deletion positive judgment threshold determination step: Obtaining a homozygous deletion sample data set, wherein the homozygous deletion samples include negative sample data and homozygous deletion positive sample data for the gene to be tested, wherein each negative sample data is obtained based on sequencing comparison data of a benign sample, and each homozygous deletion positive sample data is sequencing comparison data of a homozygous deletion occurring in the gene to be tested; Determining the corrected tumor cell copy number of the gene to be tested for each homozygous deletion sample data; For each candidate homozygous deletion copy number in the candidate homozygous deletion copy number set, performing the following homozygous deletion positive judgment basis determination operation: determining a positive coincidence rate corresponding to the candidate homozygous deletion copy number based on a comparison result of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each of the homozygous deletion positive sample data; and / or determining a positive predictive value corresponding to the candidate homozygous deletion copy number based on a comparison result of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each of the homozygous deletion sample data; The positive judgment threshold for homozygous deletion of the gene to be tested is determined based on the positive coincidence rate and / or positive predictive value corresponding to each candidate homozygous deletion copy number.

[0044] In some optional embodiments, obtaining a homozygous deletion sample data set includes: Obtaining sequencing alignment data of N pairs of 100% tumor cell lines and their paired cell lines, wherein N is a positive integer; Based on the N pairs of sequencing comparison data, generating at least one homozygous deletion positive sample data for the gene to be tested, wherein the tumor purity of each homozygous deletion positive sample data is the corresponding annotated tumor purity; Sequencing comparison data of M benign samples are obtained, and at least one negative sample data is generated based on the sequencing comparison data of the M benign samples.

[0045] In some optional embodiments, determining the corrected tumor cell copy number of the gene to be tested for each homozygous deletion sample data includes: For each homozygous deletion sample data, determine the tumor purity and tumor ploidy corresponding to the homozygous deletion sample data and the average mixed copy number for the gene to be tested; based on the tumor purity and tumor ploidy corresponding to the homozygous deletion sample data and the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the homozygous deletion sample data for the gene to be tested.

[0046] In some optional embodiments, the comparison result of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number based on the homozygous deletion positive sample data, and determining the positive coincidence rate corresponding to the candidate homozygous deletion copy number, includes: According to the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number of each homozygous deletion positive sample data, the number of true positives and the number of false negatives corresponding to the candidate homozygous deletion copy number are counted; The positive coincidence rate corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and the number of false negatives corresponding to the candidate homozygous deletion copy number.

[0047] In some optional embodiments, determining the positive predictive value corresponding to the candidate homozygous deletion copy number based on the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number in each of the homozygous deletion sample data includes: According to the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number of each homozygous deletion positive sample data, the number of true positives corresponding to the candidate homozygous deletion copy number is counted; According to the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number in each of the negative sample data, counting the number of false positives corresponding to the candidate homozygous deletion copy number; and A positive predictive value corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and the number of false positives corresponding to the candidate homozygous deletion copy number.

[0048] In some optional embodiments, generating at least one negative sample data based on the sequencing comparison data of the M benign samples includes: For each sequencing alignment data of a benign sample, different degrees of noise are added to the sequencing depth of the base sites in the sequencing alignment data of the benign sample to obtain at least one negative sample data corresponding to the sequencing alignment data of the benign sample.

[0049] In some optional embodiments, the amplification positive judgment threshold of the gene to be tested is predetermined by the following amplification positive judgment threshold determination step: Obtaining a set of amplified sample data, wherein the amplified sample data includes negative sample data and amplified positive sample data for the gene to be tested, wherein each negative sample data is obtained based on sequencing comparison data of a benign sample, and each amplified positive sample data is sequencing comparison data for copy number amplification of the gene to be tested; Determine the corrected tumor cell copy number for each amplified sample data; Determining the amplification judgment value of the corresponding amplified sample data for the gene to be tested according to the corrected tumor cell copy number of each amplified sample data; For each candidate amplification positive judgment value in the candidate amplification positive judgment value set, performing the following amplification positive judgment basis determination operation: determining a positive coincidence rate corresponding to the candidate amplification positive judgment value based on a comparison result of the amplification judgment value of each amplification positive sample data for the gene to be tested and the candidate amplification positive judgment value; and / or determining a positive predictive value corresponding to the candidate amplification positive judgment value based on a comparison result of the amplification judgment value of each amplification sample data for the gene to be tested and the candidate amplification positive judgment value; The amplification positive judgment threshold of the gene to be tested is determined according to the positive coincidence rate and / or positive predictive value corresponding to each candidate amplification positive judgment value.

[0050] In some optional embodiments, obtaining the amplified sample data set includes: Obtaining sequencing alignment data of I pairs of 100% tumor cell lines and their paired cell lines, where I is a positive integer; Based on the I pair of sequencing comparison data, generating at least one amplification-positive sample data for the gene to be tested, wherein the tumor purity of each amplification-positive sample data is the corresponding annotated tumor purity; Sequencing comparison data of J benign samples are obtained, and at least one negative sample data is generated based on the sequencing comparison data of the J benign samples.

[0051] In some optional embodiments, determining the corrected tumor cell copy number of each amplified sample data includes: For each amplified sample data, determine the tumor purity and tumor ploidy corresponding to the amplified sample data and the average mixed copy number for the gene to be tested; based on the tumor purity and tumor ploidy corresponding to the amplified sample data and the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the gene to be tested in the amplified sample data.

[0052] In some optional embodiments, the determining of the positive coincidence rate corresponding to the candidate amplification positive judgment value based on the comparison result of the amplification judgment value of the gene to be tested with the candidate amplification positive judgment value of each amplification positive sample data includes: According to the comparison result of the amplification judgment value of each amplification-positive sample data for the gene to be tested and the candidate amplification-positive judgment value, counting the number of true positives and the number of false negatives corresponding to the candidate amplification-positive judgment value; The positive coincidence rate corresponding to the candidate amplification positive judgment value is determined based on the number of true positives and the number of false negatives corresponding to the candidate amplification positive judgment value.

[0053] In some optional embodiments, the determining of a positive predictive value corresponding to the candidate amplification positive judgment value based on a comparison result of the amplification judgment value of the gene to be tested with the candidate amplification positive judgment value of each amplification sample data includes: According to the comparison result of the amplification judgment value of each amplification-positive sample data for the gene to be tested and the candidate amplification-positive judgment value, counting the number of true positives corresponding to the candidate amplification-positive judgment value; According to the comparison result of the amplification judgment value of each negative sample data for the gene to be tested and the candidate amplification positive judgment value, counting the number of false positives corresponding to the candidate amplification positive judgment value; and The positive predictive value corresponding to the candidate amplification positive judgment value is determined based on the number of true positives and the number of false positives corresponding to the candidate amplification positive judgment value.

[0054] In some optional embodiments, determining the amplification judgment value of the corresponding amplified sample data for the gene to be tested based on the corrected tumor cell copy number of each amplified sample data includes: Determine the corrected tumor cell copy number of each amplified sample data for the gene to be tested as the amplification judgment value of the amplified sample data, or, The difference between the corrected tumor cell copy number of each amplified sample data for the gene to be tested and the tumor ploidy of the amplified sample data is determined as the amplification judgment value of the amplified sample data, or The amplification judgment value of the amplified sample data is determined by dividing the ratio of the corrected tumor cell copy number of the gene to be tested for each amplified sample data by the tumor ploidy of the amplified sample data.

[0055] In a third aspect, an embodiment of the present disclosure provides an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.

[0056] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by one or more processors, implements the method described in any implementation manner in the first aspect.

[0057] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program / instruction, which implements the method described in any implementation manner in the first aspect when the computer program / instruction is executed by a processor.

[0058] In order to solve the problem that the existing gene copy number variation type detection method based on NGS sequencing technology cannot accurately identify the absolute copy number of genes in tumor cells and cannot distinguish between homozygous deletions and heterozygous deletions, the embodiments of the present disclosure provide a gene copy number variation type detection method, device, electronic device, storage medium and computer program product, which first obtains sequencing comparison data of gene sequencing of a test sample and at least one baseline sample, wherein each sequencing comparison data includes sequencing comparison data of a preset set of analysis regions of the gene to be tested, and the preset set of analysis regions of the gene to be tested includes the region where the gene to be tested is located; then, based on each sequencing comparison data, determines the tumor purity and tumor ploidy of the sample to be tested and the average mixed copy number of the gene to be tested in the sample to be tested; then, based on the tumor purity and tumor ploidy of the sample to be tested, determines the average mixed copy number of the gene to be tested in the sample to be tested; and then, based on the tumor purity and tumor ploidy of the sample to be tested, determines the average mixed copy number of the gene to be tested in the sample to be tested. The method comprises the steps of: determining the corrected tumor cell copy number of the gene to be tested in the test sample based on the tumor ploidy and the average mixed copy number of the gene to be tested in the test sample; determining the corrected tumor cell copy number of the gene to be tested in the test sample; if the average mixed copy number of the gene to be tested in the test sample is less than 2, determining the test result of the test sample for homozygous deletion of the gene to be tested based on the comparison result of the corrected tumor cell copy number of the gene to be tested in the test sample and the homozygous deletion positive judgment threshold of the gene to be tested; and / or, if the average mixed copy number of the gene to be tested in the test sample is greater than 2, determining the amplification judgment value of the gene to be tested in the test sample based on the corrected tumor cell copy number of the gene to be tested in the test sample, and determining the test result of the test sample for copy number amplification of the gene to be tested based on the comparison result of the amplification judgment value of the gene to be tested in the test sample and the amplification positive judgment threshold. By additionally considering the ploidy of the tumor in the test sample during the process of determining the corrected tumor cell copy number of the test gene in the test sample, the accuracy of the corrected tumor cell copy number of the test gene in the test sample (i.e., the absolute copy number of the test gene in the tumor cells) can be improved due to the consideration of more factors. This, in turn, can improve the accuracy of detecting the copy number variation status of the test gene. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Other features, objects, and advantages of the present disclosure will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings. The drawings are for illustration purposes only and are not to be considered as limiting the present disclosure. In the drawings: Figure 1 is an exemplary system architecture diagram in which an embodiment of the present disclosure may be applied; Figure 2A is a flow chart of an embodiment of a method for detecting gene copy number variation types according to the present disclosure; Figure 2B is a decomposed flow chart of one embodiment of step 202 according to the present disclosure; Figure 2Cis a decomposed flow chart of one embodiment of step 203 according to the present disclosure; Figure 2D is a decomposed flow chart of one embodiment of a baseline sample genetic analysis window depth correction operation according to the present disclosure; Figure 3A is a flowchart of an embodiment of the steps for determining a homozygous deletion positive judgment threshold according to the present disclosure; Figure 3B is a decomposed flow chart of one embodiment of step 301 according to the present disclosure; Figure 3C is a decomposed flow chart of one embodiment of step 3012 according to the present disclosure; Figure 3D is a flow chart of one embodiment of the operation of determining a homozygous deletion positive judgment basis according to the present disclosure; Figure 4A is a flow chart of an embodiment of the steps for determining the amplification positive judgment threshold according to the present disclosure; Figure 4B is a decomposed flow chart according to one embodiment of step 401 of the present disclosure; Figure 4C is a decomposed flow chart of one embodiment of the amplification positive sample data generation operation according to the present disclosure; Figure 4D is a flow chart of one embodiment of the operation of determining the basis for determining amplification positivity according to the present disclosure; Figure 5 It is a Probit curve diagram of the tested gene PTEN under two homozygous deletion variation types: 1-2 exon deletion type and 3-9 exon deletion type; Figure 6A 、 Figure 6B and Figure 6C These are the Probit curves of the tested genes MTAP, CDKN2A, and CDKN2B; Figure 7A 、 Figure 7B and Figure 7C These are the Probit curves of diploid CCNE1 amplification, triploid CCNE1 amplification, and tetraploid CCNE1 amplification; Figure 8 A schematic structural diagram of an embodiment of a device for detecting gene copy number variation types according to the present disclosure; Figure 9 A schematic diagram of the structure of a computer system of an electronic device suitable for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION

[0060] The present disclosure will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the relevant invention are shown in the accompanying drawings.

[0061] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0062] Figure 1 An exemplary system architecture 100 is shown to which embodiments of the disclosed gene copy number variation type detection method, apparatus, electronic device, and storage medium may be applied.

[0063] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0064] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as gene copy number variation type detection applications, high-throughput sequencing applications, bioinformatics applications, etc.

[0065] Terminal devices 101, 102, and 103 can be either hardware or software. When hardware is used, terminal devices 101, 102, and 103 can be various electronic devices with information input devices (e.g., keyboard, mouse, touch screen, microphone, camera, etc.) and information output devices (e.g., display screen, speakers, etc.), including but not limited to smartphones, tablet computers, e-book readers, laptop computers, and desktop computers. When software is used, terminal devices 101, 102, and 103 can be installed in the terminal devices listed above. They can be implemented as multiple software programs or software modules (e.g., to provide distributed services) or as a single software program or software module. This is not specifically limited here.

[0066] In some cases, the gene copy number variation type detection method provided by the present disclosure can be executed by the terminal devices 101, 102, and 103. Accordingly, the gene copy number variation type detection device can be set in the terminal devices 101, 102, and 103. In this case, the system architecture 100 may also not include the server 105.

[0067] In some cases, the gene copy number variation type detection method provided in the present disclosure can be jointly performed by the terminal devices 101, 102, 103 and the server 105. For example, the step of "obtaining sequencing comparison data of gene sequencing of the test sample and at least one baseline sample" can be performed by the terminal devices 101, 102, 103, and the steps of "determining the tumor purity and tumor ploidy of the test sample based on each sequencing comparison data" can be performed by the server 105. The present disclosure is not limited to this. Accordingly, the gene copy number variation type detection device can also be respectively provided in the terminal devices 101, 102, 103 and the server 105.

[0068] In some cases, the gene copy number variation type detection method provided in the present disclosure can be executed by the server 105. Accordingly, the gene copy number variation type detection device can also be set in the server 105. In this case, the system architecture 100 may also not include the terminal devices 101, 102, and 103.

[0069] It should be noted that server 105 can be either hardware or software. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules (for example, to provide distributed services), or as a single software program or software module. This is not specifically limited here.

[0070] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0071] Continue to refer Figure 2A , which shows a process 200 of an embodiment of a method for detecting a gene copy number variation type according to the present disclosure, the method for detecting a gene copy number variation type comprises the following steps: Step 201: Acquire sequencing comparison data of gene sequencing of a test sample and at least one baseline sample.

[0072] Here, the sample to be tested may be, for example, a tumor tissue (such as lung cancer tissue or breast cancer tissue) obtained by biopsy or surgical resection, and is used to detect whether copy number variation occurs in the sample to be tested.

[0073] The baseline sample can be normal tissue cells from the same patient as the sample to be tested, or a group of normal tissue (or benign tissue) cells from another sample. It is usually obtained from the patient's peripheral blood leukocytes (germline DNA) or normal tissue adjacent to the cancer (collected during surgery).

[0074] First, high-throughput sequencing can be performed on the sample to obtain raw sequencing data. This raw sequencing data can then be aligned (for example, using tools like BWA or Bowtie2) to a human reference genome (such as hg38 / GRCh38) to obtain sequencing alignment data for the sample. This sequencing alignment data can be, for example, a BAM (Binary Alignment Map) file.

[0075] Alternatively, high-throughput sequencing can be performed on each baseline sample to obtain raw sequencing data for the corresponding baseline sample. The raw sequencing data for each baseline sample can then be aligned (e.g., using tools like BWA or Bowtie2) to the human reference genome (e.g., hg38 / GRCh38) to obtain sequencing alignment data for the corresponding baseline sample.

[0076] As an example, suppose high-throughput sequencing is performed on each of B baseline samples. Then, after step 201 , sequencing alignment data for one test sample and sequencing alignment data for each of the B baseline samples can be obtained, resulting in a total of (B+1) sequencing alignment data. For example, a total of (B+1) BAM files can be obtained.

[0077] It should be noted that the (B+1) sequencing alignment data mentioned above all include sequencing alignment data of a preset set of gene analysis regions to be tested, and the preset set of gene analysis regions to be tested includes the region where the gene to be tested is located.

[0078] In the above sequencing process, various library construction and sequencing methods can be used. For example, hybridization capture method, amplicon method or whole genome sequencing method, etc., as long as the above library construction and sequencing method can sequence each predetermined gene region to be tested and obtain sequencing comparison data for each predetermined gene region to be tested.

[0079] It should be noted that the sequencing methods for the test sample and each baseline sample can be the same library construction method and the same sequencing method. For example, if hybridization capture sequencing is used, the capture probe settings can be the same.

[0080] Step 202: Determine the tumor purity and tumor ploidy of the sample to be tested based on the sequencing comparison data.

[0081] Here, the tumor purity and tumor ploidy of the test sample can be determined based on the sequencing alignment data of a test sample and the sequencing alignment data of each baseline sample obtained in step 201, for example, the sequencing alignment data of 1 test sample and the sequencing alignment data of B baseline samples.

[0082] Here, the tumor purity of the sample to be tested is the proportion of tumor cells in the sample to be tested, and its value can be between 0 and 100%.

[0083] The tumor ploidy of the sample to be tested is the average copy number of the tumor cell genome (usually 2 (diploid), 3, 4, etc.).

[0084] In some optional implementations, step 202 may include: Figure 2B The following steps 2021 to 2025 are shown: Step 2021: Based on the sequencing comparison data of the sample to be tested, determine the abundance of each SNP site in the preset whole-genome SNP site set of the sample to be tested.

[0085] Specifically, step 2021 may be performed as follows: First, variant calling is to detect SNP sites in the sequencing data of the sample to be tested and generate a VCF file: For example, you can use GATK (gold standard), BCFtools (lightweight), or FreeBayes to generate VCF files.

[0086] Then, calculate the SNP abundance (Allele Frequency), that is, extract the allele frequency (AF) of each SNP site from the VCF file.

[0087] For example, this can be achieved using BCFtools, GATK VariantsToTable, or through custom scripts (using scripting languages ​​such as Python / R).

[0088] Here, the preset whole genome SNP site set may be SNP sites pre-set according to actual detection needs.

[0089] Step 2022: For each preset whole-genome SNP site, the related gene analysis region is divided into at least one SNP analysis region according to a first preset length.

[0090] Here, various methods can be used to first determine the relevant gene analysis region for each whole-genome SNP site. As an example, the whole-genome SNP site can be used as a reference position, and the relevant gene analysis region for the whole-genome SNP site can be determined according to a preset gene analysis region determination method. For example, the region between the first base site at a third preset length forward from the reference position and the second base site at a fourth preset length backward from the reference position is used as the relevant gene analysis region for the whole-genome SNP site.

[0091] Then, the gene analysis region related to the whole genome SNP site is divided into at least one SNP analysis region according to a first preset length.

[0092] Here, the first preset length can be pre-set according to actual detection needs. For example, when the test sample and the baseline sample are sequenced using a hybrid capture method, the first preset length can be determined based on the probe length during the hybrid capture method sequencing process. Generally speaking, the probe length of each capture region is the same, such as 120 bp, and the length of each capture probe can be used as the first preset length. That is, as an example, the first preset length can be 120 bp.

[0093] Step 2023: For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each SNP analysis region.

[0094] Here, for each of the sequencing alignment data of the test sample and the sequencing alignment data of the B baseline samples obtained in step 201, the corrected average coverage depth for each SNP analysis region in the sequencing alignment data can be determined. Specifically, the average coverage depth of each SNP analysis region in the sequencing alignment data can be first calculated, and then various coverage depth correction methods can be used to correct the average coverage depth of each SNP analysis region.

[0095] Optionally, step 2023 may be performed as follows: For each sequencing alignment data, each SNP analysis region is used as the target analysis region, and the target analysis region sequencing depth correction operation is performed to obtain the corrected average coverage depth of each SNP analysis region. Here, the target analysis region sequencing depth correction operation can be performed as follows: First, the sequencing depth of each base site in each target analysis region is normalized according to the total sequencing depth of the sequencing alignment data to obtain the corresponding normalized sequencing depth.

[0096] For example, the normalized sequencing depth of a base site within a target analysis region can be calculated as follows: the ratio of the sequencing depth of the base site divided by the total sequencing depth of the sequencing alignment data multiplied by the global scaling factor. Here, the global scaling factor is usually the median or mean of the total depth of all samples (such as 100×) to make the normalized sequencing depth more intuitive. The global scaling factor is related to the size of the capture region. Specifically, it is generally based on the size of the panel (gene combination / detection package) used for sequencing, such as using 10 to the power of 7 for a large panel and 10 to the power of 6 for a small panel, with the aim of keeping the normalized sequencing depth within 100.

[0097] Normalizing sequencing depth can eliminate differences in sequencing data volume between different samples.

[0098] Then, for each target analysis region, the average sequencing depth of the sequencing alignment data in the target analysis region is determined based on the normalized sequencing depth of each base site in the target analysis region.

[0099] That is, the ratio of the sum of the normalized sequencing depths of all base sites in the target analysis region divided by the number of base sites in the target analysis region can be used as the average sequencing depth of the sequencing alignment data in the target analysis region.

[0100] Then, based on the number of probes covering each target analysis region, the average coverage depth of the sequencing alignment data in each target analysis region is corrected by local weighted regression.

[0101] Here, in order to eliminate the difference in coverage depth caused by the difference in the number of probes covering different target analysis regions, the average coverage depth of the sequencing alignment data in each target analysis region can be locally weighted regression corrected according to the number of probes covering each target analysis region.

[0102] Finally, based on the GC content ratio of the sequencing alignment data in each target analysis region, the average coverage depth of the sequencing alignment data in each target analysis region was corrected by local weighted regression to obtain the corrected average coverage depth of the corresponding target analysis region.

[0103] Similarly, in order to eliminate the coverage depth differences caused by differences in GC content ratios in different target analysis regions, the average coverage depth of the sequencing alignment data in each target analysis region can be locally weighted regression corrected based on the GC content ratio of the sequencing alignment data in each target analysis region to obtain the corrected average coverage depth of the corresponding target analysis region.

[0104] After step 2023 , the corrected average coverage depth of each sequencing data for each SNP analysis region in the (B+1) sequencing comparison data of 1 test sample and B baseline samples can be obtained.

[0105] Step 2024, for each SNP analysis region, determine the mixed copy number of the sequencing alignment data of the sample to be tested in the SNP analysis region based on the corrected average coverage depth of the sequencing alignment data of the sample to be tested in the SNP analysis region and the average of the corrected average coverage depth of the sequencing alignment data of each baseline sample in the SNP analysis region.

[0106] That is, here, for each SNP analysis region, first, obtain the corrected average coverage depth of the test sample obtained in step 2023 in the SNP analysis region. Then, calculate the average coverage depth after correction of the sequencing alignment data of the B baseline samples obtained in step 2023 in the SNP analysis region. , and finally according to and Determine the mixed copy number of the sequencing alignment data of the sample to be tested in the SNP analysis region. Specifically, the corrected average coverage depth of the sample to be tested in the SNP analysis region can be Divide by the average coverage depth of the sequencing alignment data of each baseline sample in the SNP analysis region after correction The ratio of the two samples is multiplied by 2 to determine the mixed copy number of the sequencing data of the sample to be tested in the SNP analysis region. , that is, it can be calculated using the following formula:

[0107] If there is SNP analysis regions, after step 2024, can be obtained The mixed copy number of each SNP analysis region in the SNP analysis region can be obtained. Mixed copy number .

[0108] Step 2025 , based on the abundance and mixed copy number of each SNP site in the sequencing comparison data of the sample to be tested, determine the tumor purity and tumor ploidy of the sample to be tested.

[0109] Here, the mixed copy number of each SNP site is the mixed copy number of the SNP analysis region where the SNP site is located.

[0110] If there is a SNP sites, various implementation methods can be used here, based on The abundance and mixed copy number of each SNP site in the SNP sites are used to determine the tumor purity and tumor ploidy of the sample to be tested.

[0111] As an example, this can be achieved through existing mature bioinformatics analysis tools or equivalent technical means. For example, using bioinformatics analysis tools with similar functions (ABSOLUTE, Sequenza, ASCAT, FACETS, etc.), based on their principles and processes, and combined with the specific needs of the present disclosure, adaptive adjustments and implementation can be made to determine the tumor purity of the sample to be tested based on the abundance and mixed copy number of each SNP site in the sequencing comparison data of the sample to be tested. and tumor ploidy It is understood that other methods can also be used to determine the tumor purity of the sample to be tested. and tumor ploidy This is not the focus of the present invention and will not be elaborated here.

[0112] The optional method of the above steps 2021 to 2025 can improve the accuracy of the average coverage depth of each SNP analysis region by dividing the relevant gene analysis region of each preset whole-genome SNP site into at least one SNP analysis region and then correcting the average coverage depth of each SNP analysis region. And according to the corrected average coverage depth of the sample to be tested and the baseline sample, the mixed copy number of each SNP analysis region is obtained, and then the mixed copy number of each SNP analysis region is used as the mixed copy number of all SNP sites in the SNP analysis region, which can greatly reduce the amount of calculation of the mixed copy number of all SNP sites. Finally, according to the abundance and mixed copy number of each SNP site, the tumor purity of the sample to be tested is obtained. and tumor ploidy .

[0113] Step 203: Determine the average mixed copy number of the gene to be tested in the sample to be tested based on each sequencing comparison data.

[0114] If the genetic sequencing process for the test sample and at least one baseline sample is targeted sequencing (e.g., panel sequencing, exome sequencing), the genes to be tested can be those selected during the experimental design (e.g., cancer-related genes such as PTEN, TP53, ERBB2, EGFR, and BRCA1 / 2). In targeted sequencing, the genes to be tested are determined by the coverage of the probes or capture kit (e.g., Illumina TruSeq, Agilent SureSelect). For example, tumor panels (e.g., MSK-IMPACT) include hundreds of cancer driver genes, while clinical diagnostic panels (e.g., FoundationOne CDx) cover key genes for specific cancer types.

[0115] If the genetic sequencing process for the test sample and at least one baseline sample is whole genome sequencing (WGS), in theory, any gene can be analyzed for the test gene, but usually priority is given to genes known to be associated with diseases (such as tumors).

[0116] The average combined copy number of the gene being tested in a test sample can be defined as the weighted average copy number of the gene being tested across all cells (tumor cells + normal cells) in the test sample (e.g., tumor tissue). Because tumor samples typically contain tumor cells and a mixture of normal cells (e.g., stromal cells, immune cells), the average combined copy number is a composite of the copy numbers of these two cell populations.

[0117] In some optional implementations, step 203 may include: Figure 2CThe following steps 2031 to 2037 are shown: Step 2031 : For each preset gene analysis region to be tested, the preset gene analysis region to be tested is divided into at least one gene analysis window according to a second preset length.

[0118] Here, the second preset length can be set according to actual detection needs. For example, when the hybridization capture method is used to sequence the test sample and each baseline sample, the second preset length can be determined based on the average probe laying multiplier and the average probe length of each capture region.

[0119] Here, executing step 2031 can achieve two technical effects: First, generally speaking, the length of the pre-set gene analysis region to be tested is 120 bp. If the pre-set gene analysis region to be tested is not divided, each pre-set gene analysis region as an analysis interval corresponds to only one data point, and the statistical test in the subsequent step 2036 cannot be performed. Therefore, it is necessary to divide it into smaller regions for the subsequent statistical test.

[0120] Second, in the subsequent process of correcting the average coverage depth using the number of covered probes, the accuracy of the corrected average coverage depth can be improved by dividing the preset gene analysis region to be tested into gene analysis regions with smaller granularity.

[0121] As an example, assuming that the average probe laying multiplier of each capture area is 5 and the average probe length is 120BP, the second preset length can be the ratio of the average probe length 120BP divided by the average probe laying multiplier 5, that is, the probe step length is 24BP.

[0122] Step 2032: For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each gene analysis window.

[0123] Here, for each of the sequencing alignment data of the test sample and the sequencing alignment data of the B baseline samples obtained in step 201, the corrected average coverage depth for each gene analysis window in the sequencing alignment data can be determined. Specifically, the average coverage depth of each gene analysis window in the sequencing alignment data can be first calculated, and then various coverage depth correction methods can be used to correct the average coverage depth of each gene analysis window.

[0124] Optionally, step 2032 may be performed as follows: for each sequencing alignment data, each gene analysis window is used as the target analysis region, and a sequencing depth correction operation of the target analysis region is performed to obtain the corrected average coverage depth of each gene analysis window.

[0125] Specifically, the specific execution process of the target analysis region sequencing depth correction operation can refer to the relevant records about step 2023 above, which will not be repeated here.

[0126] After step 2032 , the corrected average coverage depth of each sequencing data for each gene analysis window in the (B+1) sequencing alignment data of 1 test sample and B baseline samples can be obtained.

[0127] Step 2033, for each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, determine the reference baseline sample sequencing alignment data corresponding to the gene analysis window in the sequencing alignment data of each baseline sample, as well as the corrected baseline sample average sequencing depth and the corrected baseline sample average sequencing depth deviation of the gene analysis window.

[0128] After step 2032 , the corrected average coverage depth of the sequencing alignment data of each baseline sample in each gene analysis window has been obtained.

[0129] Optionally, step 2033 may be performed as follows: for each gene analysis window, perform the following baseline sample gene analysis window depth correction operation, the baseline sample gene analysis window depth correction operation includes: Figure 2D The following steps 20331 to 20333 are shown: Step 20331: Determine the normal range of the average sequencing depth of the baseline samples based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window.

[0130] In practice, due to low capture efficiency, high GC content, repetitive sequences, etc., the coverage depth of some gene analysis windows in some baseline samples may drop / suddenly increase. Therefore, the normal range of the average sequencing depth of the baseline samples can be determined based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window. For example, the median or mean of the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window can be first determined. Then, the absolute value of the difference between the above median or mean is less than a specified threshold (for example, 3 times the standard deviation) as the normal range of the average sequencing depth of the baseline samples.

[0131] In step 20332, the sequencing alignment data of each baseline sample whose corrected average coverage depth in the gene analysis window falls within the normal range of the average depth of baseline sample sequencing is determined as the reference baseline sample sequencing alignment data corresponding to the gene analysis window.

[0132] In other words, the reference baseline sample sequencing alignment data corresponding to the gene analysis window excludes those baseline sample sequencing alignment data whose corrected average coverage depth values ​​of the gene analysis window are in the abnormal range (i.e., the absolute value of the difference between the median or mean is greater than or equal to the specified threshold). The remaining baseline samples that can represent the gene analysis window can be used as the reference baseline sample sequencing alignment data corresponding to the gene analysis window.

[0133] Step 20333, based on the corrected average coverage depth of the sequencing alignment data of each reference baseline sample in the gene analysis window in the gene analysis window, determine the corrected baseline sample average sequencing depth and the corrected baseline sample average sequencing depth deviation of the gene analysis window.

[0134] After step 20332, the baseline sample sequencing alignment data with a corrected average coverage depth value within an abnormal range has been removed from the reference baseline sample sequencing alignment data for the gene analysis window. Here, the mean of the corrected average coverage depth of the reference baseline sample sequencing alignment data for the gene analysis window can be calculated, and the calculated mean is used as the corrected baseline sample average sequencing depth for the gene analysis window. Based on the corrected average coverage depth of the reference baseline sample sequencing alignment data for the gene analysis window and the corrected baseline sample average sequencing depth for the gene analysis window, the corrected baseline sample average sequencing depth deviation is determined. The corrected baseline sample average sequencing depth deviation can be, for example, a standard deviation, a coefficient of variation, an absolute median difference, a Z-score (normalized deviation), and the like.

[0135] Step 2034 : Based on the corrected baseline sample average sequencing depth deviation of each gene analysis window, a reference gene analysis window is determined in each gene analysis window.

[0136] After step 2033, the corrected average baseline sample sequencing depth deviation of each gene analysis window can be obtained. Here, various abnormality determination methods can be used to determine the statistically abnormal corrected average baseline sample sequencing depth deviation from the corrected average baseline sample sequencing depth deviation of each gene analysis window. For example, it can be the corrected average baseline sample sequencing depth deviation with a Z-score greater than 3. Then, the gene analysis window whose corrected average baseline sample sequencing depth deviation in each gene analysis window does not belong to the above-mentioned statistically abnormal corrected average baseline sample sequencing depth deviation is determined as the reference gene analysis window.

[0137] Through steps 2031 to 2034, the sequencing comparison data of each baseline sample is screened in each gene analysis window to determine the reference gene analysis window, and for each reference gene analysis window, the baseline sample sequencing comparison data is screened to obtain the reference baseline sample sequencing comparison data.

[0138] Step 2035: determine each reference gene analysis window where the gene to be tested is located as the gene to be tested analysis window.

[0139] In order to determine the mixed copy number of the gene to be tested, the region where the gene to be tested is located is first required, that is, each reference gene analysis window where the gene to be tested is located is determined as the analysis window of the gene to be tested.

[0140] Step 2036, for each gene analysis window to be tested, determine the mixed copy number of the gene analysis window to be tested in the sample to be tested based on the corrected average coverage depth of the sequencing alignment data of the sample to be tested in the gene analysis window to be tested and the corrected average sequencing depth of the baseline samples in the gene analysis window to be tested.

[0141] Specifically, the corrected average coverage depth of the sample to be tested in the analysis window of the gene to be tested can be Divide by the average sequencing depth of the corrected baseline samples in the analysis window of the gene to be tested The ratio of the two products is multiplied by 2 to determine the mixed copy number of the gene analysis window in the sample to be tested. , that is, it can be calculated using the following formula:

[0142] If there is After step 2036, the gene analysis window to be tested can be obtained. The mixed copy number of each gene analysis window to be tested can be obtained The mixed copy number of the gene analysis window to be tested.

[0143] Step 2037: Determine the mean of the mixed copy numbers of each analysis window of the gene to be tested in the sample to be tested as the average mixed copy number of the gene to be tested in the sample to be tested.

[0144] The optional method of the above steps 2031 to 2037 can improve the accuracy of the average coverage depth of each gene analysis window by dividing each preset capture area into smaller gene analysis windows and then correcting the average coverage depth of each gene analysis window in each sequencing alignment data. And based on the sequencing alignment data of each baseline sample, the reference gene analysis window is screened in each gene analysis window, and for each reference gene analysis window, the baseline sample sequencing alignment data is screened to obtain the reference baseline sample sequencing alignment data, which can remove abnormal gene analysis windows and abnormal baseline samples. Then, according to the corrected average coverage depth of the sample to be tested and the baseline sample, the mixed copy number of the reference gene analysis window where each gene to be tested in the sample to be tested is located is obtained, and finally the average mixed copy number of the gene to be tested in the sample to be tested is determined. Combining the above steps, the average mixed copy number of the gene to be tested in the sample to be tested can be increased. Calculation accuracy.

[0145] Step 204 : determining the corrected tumor cell copy number of the gene to be tested in the sample to be tested based on the tumor purity and tumor ploidy of the sample to be tested and the average mixed copy number of the gene to be tested in the sample to be tested.

[0146] Typically, to determine the corrected tumor cell copy number of the gene being tested in the sample being tested, the tumor purity of the sample being tested and the average mixed copy number of the gene being tested in the sample being tested are required. By solving the traditional correction formula, the corrected tumor cell copy number of the gene being tested in the sample being tested can be obtained. The traditional correction formula can be expressed as follows:

[0147] in, is the tumor purity of the sample to be tested, is the average mixed copy number of the gene to be tested in the sample to be tested, is the corrected tumor cell copy number of the gene to be tested in the sample to be tested. Solving the above traditional correction formula can obtain the following formula:

[0148] The above formula can be used to calculate the corrected tumor cell copy number of the gene to be tested in the sample to be tested: . However, the above calculation method does not take into account the tumor ploidy of the sample to be tested, which will cause polyploidy to interfere with the calculation of gene copy number, and thus cannot accurately reveal the genomic variation characteristics of tumor cells. Therefore, in the process of determining the corrected tumor cell copy number of the gene to be tested in the sample to be tested, in addition to considering the tumor purity of the sample to be tested and the average mixed copy number of the gene to be tested in the sample to be tested, the tumor ploidy of the sample to be tested must also be considered. By using tumor ploidy and tumor purity to correct the tumor cell copy number of the gene to be tested in the sample to be tested, the accuracy of copy number variation detection can be improved, and the whole genome ploidy (such as triploidy) can be distinguished from the amplification of local chromosome regions.

[0149] Optionally, step 204 can be performed as follows: substitute the tumor purity and tumor ploidy of the sample to be tested, and the average mixed copy number of the gene to be tested in the sample to be tested into the corrected tumor cell copy number solution formula to obtain the corrected tumor cell copy number of the gene to be tested.

[0150] Optionally, here, the formula for solving the corrected tumor cell copy number can be obtained by deriving a correction formula, wherein the correction formula is as follows:

[0151] in, is the tumor purity of the sample to be tested, The tumor ploidy of the sample to be tested, is the average mixed copy number of the gene to be tested in the sample to be tested, is the corrected tumor cell copy number of the gene to be tested in the sample to be tested.

[0152] The following is an explanation of the above correction formula: The numerator on the left side of the equal sign in the corrected formula is: .in, is the tumor purity of the sample to be tested, is the corrected tumor cell copy number of the gene to be tested in the sample to be tested, so It represents the theoretical average mixed copy number of the gene to be tested in the tumor cells of the sample to be tested. is the theoretical average mixed copy number of the gene to be tested in normal cells in the sample to be tested, It can be the theoretical average mixed copy number of the gene to be tested in the sample to be tested.

[0153] The denominator on the left side of the equal sign in the correction formula is: .in, and Represent the tumor purity and tumor ploidy of the sample to be tested, respectively. is the theoretical average copy number of tumor cells in the sample to be tested, is the theoretical average copy number of normal cells in the sample to be tested, It represents the theoretical average mixed copy number of the sample to be tested.

[0154] Therefore, the correction formula to the left of the equal sign is: It represents the ratio of the theoretical average mixed copy number of the gene to be tested in the sample to the average mixed copy number of the sample to be tested.

[0155] Correct the numerator on the right side of the equal sign in the formula is the average mixed copy number of the gene to be tested in the actual sample to be tested, and 2 is the average mixed copy number of the gene to be tested in the actual sample to be tested.

[0156] Therefore, the right side of the equal sign in the correction formula is: It represents the ratio of the average mixed copy number of the gene to be tested in the sample to be tested to the average mixed copy number of the sample to be tested.

[0157] In other words, the numerators on both sides of the equal sign correspond to the average mixed copy number of the gene to be tested in the sample to be tested, and the denominator corresponds to the average mixed copy number of the sample to be tested. The left side of the equal sign corresponds to the theoretical data, and the right side of the equal sign corresponds to the actual data.

[0158] Because it is known 、 and By deducing the above correction formula, we can obtain the formula for solving the corrected tumor cell copy number, which is as follows:

[0159] In this way, the corrected tumor cell copy number of the gene to be tested in the test sample can be calculated.

[0160] Here, after executing step 204, the process may proceed to step 205 or step 206.

[0161] Step 205 , in response to the average mixed copy number of the gene to be tested in the sample to be tested being less than 2, generating a detection result for homozygous deletion of the gene to be tested in the sample to be tested based on a comparison result of the corrected tumor cell copy number of the gene to be tested in the sample to be tested and a positive judgment threshold for homozygous deletion of the gene to be tested.

[0162] Here, if the average mixed copy number of the gene to be tested in the test sample is less than 2, it can be first determined whether the corrected tumor cell copy number of the gene to be tested in the test sample is greater than the homozygous deletion positive judgment threshold of the gene to be tested.

[0163] If it is determined to be not greater than, a test result indicating a positive homozygous deletion of the gene to be tested in the test sample may be generated.

[0164] If it is determined to be greater than, a test result indicating that the homozygous deletion of the gene to be tested in the test sample is negative can be generated.

[0165] Here, the positive judgment threshold for homozygous deletion of the gene to be tested can be predetermined in various implementations, for example, it can be manually set by professional technicians.

[0166] Alternatively, the positive judgment threshold for homozygous deletion of the gene to be tested can be determined by Figure 3A The homozygous deletion positive judgment threshold determination step 300 is predetermined, and the homozygous deletion positive judgment threshold determination step 300 may include the following steps 301 to 304: Step 301: Obtain a homozygous deletion sample data set.

[0167] Here, the homozygous deletion samples may include negative sample data and homozygous deletion positive sample data for the gene to be tested, wherein each negative sample data is obtained based on sequencing comparison data of benign samples, and each homozygous deletion positive sample data is sequencing comparison data of homozygous deletion for the gene to be tested.

[0168] Optionally, step 301 may include: Figure 3B The following steps 3011 to 3013 are shown: Step 3011: Obtain sequencing alignment data of N pairs of 100% tumor cell lines and their paired cell lines.

[0169] Here, N is a positive integer.

[0170] A 100% tumor cell line refers to a cell line composed of pure tumor cells and does not contain other non-tumor cells (such as stromal cells, immune cells, etc.).

[0171] The paired cell line of a 100% tumor cell line refers to a control cell line that is derived from the same individual as the tumor cell line but has a normal phenotype.

[0172] Here, high-throughput sequencing can be performed on each 100% tumor cell line and each paired cell line to generate the corresponding raw sequencing data. The raw sequencing data can then be aligned (for example, using tools such as BWA and Bowtie2) to the human reference genome (such as hg38 / GRCh38) to generate the corresponding sequencing alignment data. Here, the sequencing alignment data can be, for example, a BAM file.

[0173] Step 3012: Generate at least one homozygous deletion positive sample data for the gene to be tested based on N pairs of sequencing comparison data.

[0174] Here, the tumor purity corresponding to each homozygous deletion positive sample data generated is the corresponding annotated tumor purity.

[0175] Optionally, step 3012 may be performed as follows: for each tumor cell line sequencing comparison data, perform a homozygous deletion positive sample data generation operation. The homozygous deletion positive sample data generation operation may include: Figure 3C The following steps 30121 and 30122 are shown: Step 30121: Delete the read segments corresponding to the gene to be tested in the sequencing alignment data of the tumor cell line.

[0176] The reads corresponding to the gene to be tested in the sequencing alignment data of the tumor cell line may include, but are not limited to, reads covering specific regions (such as exons (general), introns, or regulatory regions) of the gene to be tested in the sequencing alignment data.

[0177] That is, the homozygous deletion of the gene to be tested is simulated by deleting the corresponding read segments of the gene to be tested.

[0178] Step 30122, for each preset tumor purity gradient in the preset tumor purity gradient set, based on the sequencing alignment data of the tumor cell line and the sequencing alignment data of the corresponding paired cell line, generate homozygous deletion positive sample data of the tumor cell line under the preset tumor purity gradient for the gene to be tested, and determine the preset tumor gradient as the annotated tumor purity of the generated homozygous deletion positive sample data.

[0179] That is, through step 302, the sequencing comparison data of each tumor cell line are simulated to generate homozygous deletion positive data with different tumor purity gradients, and the generated homozygous deletion positive sample data are labeled with corresponding tumor purity.

[0180] For example, the preset tumor purity gradient is , can be randomly selected from the sequencing alignment data of the tumor cell line Read segments are randomly selected from the sequencing alignment data of the paired cell lines of this tumor cell line The two selected reads are mixed to generate homozygous deletion positive sample data, and the tumor purity can be obtained as Homozygous deletion positive sample data.

[0181] Step 3013: Acquire sequencing comparison data of M benign samples, and generate at least one negative sample data based on the sequencing comparison data of the M benign samples.

[0182] Here, the benign sample can be benign tissue or adjacent tissue. M is a positive integer.

[0183] Specifically, high-throughput sequencing can be performed on each benign sample to generate raw sequencing data. This raw sequencing data can then be aligned (e.g., using tools like BWA or Bowtie2) to a human reference genome (e.g., hg38 / GRCh38) to generate sequencing alignment data. The sequencing alignment data can be, for example, a BAM file.

[0184] In order to expand the number of benign samples, various methods can be used to generate at least one negative sample data based on the sequencing comparison data of M benign samples.

[0185] In practice, since the sequencing process may have various noises, in order to adapt to the actual situation in practice, step 3013 may optionally be performed as follows: For each sequencing alignment data of a benign sample, noise of different proportions is added to the sequencing depth of the base sites in the sequencing alignment data of the benign sample to obtain at least one negative sample data corresponding to the sequencing alignment data of the benign sample.

[0186] As an example, you can add noise as follows: Assuming that the sequencing depth of each base site follows a log-normal distribution, we first take the logarithmic value of the sequencing depth of each base site in the sequencing alignment data of the benign sample. Then, we calculate the standard deviation (SD) of the logarithm of the sequencing depth of all base sites. Next, we randomly add sigma to the logarithm of the base site depth of the corresponding proportion in the sequencing alignment data of the benign sample, where sigma ~ rnorm(0, sd). To prevent the noise of all base sites from being too consistent, we assume sd 2 Finally, by exponentially calculating the logarithm of the sequencing depth at each base site, we can add varying degrees of noise to the sequencing depth of the base site in the sequencing comparison data of the benign sample.

[0187] Step 302: Determine the corrected tumor cell copy number of the gene to be tested for each homozygous deletion sample data.

[0188] Here, the same method as that described in steps 202, 203, and 204 for determining the corrected tumor cell copy number of the gene to be tested in the sample to be tested can be used to determine the corrected tumor cell copy number of the gene to be tested for each homozygous deletion sample data.

[0189] Specifically, for each homozygous deletion sample data, the tumor purity and tumor ploidy of the homozygous deletion sample data can first be determined based on the sequencing comparison data of the homozygous deletion sample data and at least one baseline sample obtained in step 201 (using the method described in step 202); then, based on the sequencing comparison data of the homozygous deletion sample data and at least one baseline sample obtained in step 201, the average mixed copy number of the gene to be tested in the homozygous deletion sample data is determined; finally, based on the tumor purity and tumor ploidy of the homozygous deletion sample data and the average mixed copy number of the gene to be tested in the homozygous deletion sample data, the corrected tumor cell copy number of the gene to be tested in the homozygous deletion sample data is determined.

[0190] Step 303 : For each candidate homozygous deletion copy number in the candidate homozygous deletion copy number set, a homozygous deletion positive judgment basis determination operation is performed.

[0191] Here, the candidate homozygous deletion copy number set may be pre-set by professional technicians based on the possible value range of the homozygous deletion copy number in practice.

[0192] Here, the homozygous deletion positive judgment basis determination operation may include the following Figure 3D The following steps 3031 and / or 3032 are shown: Step 3031 : Determine the positive coincidence rate corresponding to the candidate homozygous deletion copy number based on the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number in each homozygous deletion positive sample data.

[0193] Specifically, step 3031 may be performed as follows: First, based on the comparison results of the corrected tumor cell copy number of the gene to be tested for each homozygous deletion positive sample data and the candidate homozygous deletion copy number, the number of true positives and false negatives corresponding to the candidate homozygous deletion copy number are counted. First, for each homozygous deletion positive sample data, it is determined whether the corrected tumor cell copy number of the homozygous deletion positive sample data for the gene to be tested is greater than the candidate homozygous deletion positive copy number. If it is determined that it is not greater than, it can be determined that the homozygous deletion positive sample data is detected as homozygous deletion positive. If it is determined that it is greater than, it can be determined that the homozygous deletion positive sample data is detected as homozygous deletion negative.

[0194] Then, the number of homozygous deletion positive sample data detected as homozygous deletion positive in each homozygous deletion positive sample data is counted as the number of true positives corresponding to the candidate homozygous deletion copy number. . Count the number of homozygous deletion negative samples detected in each homozygous deletion positive sample data as the number of false negatives corresponding to the candidate homozygous deletion copy number .

[0195] Finally, the positive coincidence rate corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and the number of false negatives corresponding to the candidate homozygous deletion copy number.

[0196] Here, PPA (Positive Percent Agreement) can be calculated according to the following formula:

[0197] in, and are the number of true positives and false negatives corresponding to the candidate homozygous deletion copy number, is the calculated positive coincidence rate corresponding to the candidate homozygous deletion copy number.

[0198] Step 3032: Determine the positive predictive value corresponding to the candidate homozygous deletion copy number based on the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number in each homozygous deletion sample data.

[0199] Specifically, step 3032 may be performed as follows: First, based on the comparison results of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number of each homozygous deletion positive sample data, the number of true positives corresponding to the candidate homozygous deletion copy number is counted.

[0200] For details, please refer to the relevant description in step 3031, which will not be repeated here.

[0201] Then, based on the comparison results of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number for each negative sample data, the number of false positives corresponding to the candidate homozygous deletion copy number is counted.

[0202] First, for each negative sample data, determine whether the corrected tumor cell copy number of the negative sample data for the gene to be tested is greater than the candidate homozygous deletion positive copy number. If it is determined that it is not greater than, it can be determined that the negative sample data is detected as homozygous deletion positive. If it is determined that it is greater, it can be determined that the negative sample data is detected as homozygous deletion negative. Then, count the number of negative sample data detected as homozygous deletion positive in each negative sample data as the number of false positive data corresponding to the candidate homozygous deletion copy number. .

[0203] Finally, the positive predictive value corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and the number of false positives corresponding to the candidate homozygous deletion copy number.

[0204] Specifically, for a candidate homozygous deletion copy number, the positive predictive value corresponding to the candidate homozygous deletion copy number can be calculated according to the following formula: (Positive Predictive Value):

[0205] in, and are the number of true positives and false positives corresponding to the candidate homozygous deletion copy number, is the calculated positive predictive value corresponding to the copy number of the candidate homozygous deletion.

[0206] Step 304 : Determine a homozygous deletion positive judgment threshold for the gene to be tested based on the positive coincidence rate and / or positive predictive value corresponding to each candidate homozygous deletion copy number.

[0207] Here, the preferred homozygous deletion copy number whose corresponding positive compliance rate and / or positive predictive value meets the preset detection standard can be first determined in the candidate homozygous deletion copy number set. Among them, the preset detection standard can be pre-set according to the needs of the specific detection application scenario. As an example, the preset detection standard can be: PPV≥99% and PPA≥95%. Then, based on the preferred homozygous deletion copy numbers determined above, the homozygous deletion positive judgment threshold of the gene to be tested is determined. For example, the maximum value among the determined preferred homozygous deletion copy numbers can be determined as the homozygous deletion positive judgment threshold of the gene to be tested.

[0208] The optional methods of steps 301 to 304 above can achieve, including but not limited to, the following technical effects: First, by generating homozygous deletion-positive sample data for the gene to be tested based on sequencing comparison data of a small number of 100% tumor cell lines and their paired cell lines, the sequencing cost required to generate homozygous deletion-positive sample data was reduced.

[0209] Secondly, by generating a large amount of negative sample data based on sequencing comparison data of a small number of benign samples, the sequencing cost of benign samples is reduced.

[0210] In addition, by additionally considering the tumor ploidy of the homozygous deletion positive sample data and the negative sample data in the process of determining the corrected tumor cell copy number of the homozygous deletion positive sample data and the negative sample data for the gene to be tested, the interference of aneuploidy (polyploidy or haploidy) on the test results can be eliminated.

[0211] Step 206, in response to the average mixed copy number of the gene to be tested in the sample to be tested being greater than 2, determining an amplification judgment value of the gene to be tested in the sample to be tested based on the corrected tumor cell copy number of the gene to be tested in the sample to be tested, and generating a detection result for the amplification of the copy number of the gene to be tested in the sample to be tested based on a comparison result of the amplification judgment value of the gene to be tested in the sample to be tested and an amplification positive judgment threshold of the gene to be tested.

[0212] Here, if the average mixed copy number of the gene to be tested in the sample to be tested is greater than 2, various implementation methods can be used to first determine the amplification judgment value of the gene to be tested in the sample to be tested based on the corrected tumor cell copy number of the gene to be tested in the sample to be tested.

[0213] Alternatively, the corrected tumor cell copy number of the gene to be tested in the test sample can be The amplification judgment value of the gene to be tested in the sample to be tested is determined, or the difference between the corrected tumor cell copy number of the gene to be tested in the sample to be tested and the tumor ploidy of the sample to be tested is subtracted ( ) is determined as the amplification judgment value of the gene to be tested in the sample to be tested, or the ratio of the corrected tumor cell copy number of the gene to be tested in the sample to be tested divided by the tumor ploidy of the sample to be tested Determine the amplification judgment value of the gene to be tested in the sample to be tested.

[0214] Preferably, the difference in tumor ploidy of the sample to be tested can be obtained by subtracting the corrected tumor cell copy number of the gene to be tested in the sample to be tested from the tumor ploidy of the sample to be tested ( ) is determined as the amplification judgment value of the gene to be tested in the sample to be tested. This is because if the entire genome of the tumor patient is amplified, for example, amplified to tetraploid, hexaploid or octaploid, then the corrected tumor cell copy number of the gene to be tested in the sample to be tested It is also amplified accordingly. For example, if the amplification is 8, if the corrected tumor cell copy number of the gene to be tested in the sample to be tested is If the amplification judgment value of the gene to be tested in the sample to be tested is determined, it is impossible to distinguish whether the gene to be tested itself has been amplified. ) is determined as the amplification judgment value of the gene to be tested in the sample to be tested, which can solve the problem of determining whether the gene to be tested has been amplified individually when the entire genome has been amplified. This is because It is aimed at the gene to be tested in the sample to be tested, and is for the entire sample to be tested, ( ) can reflect whether the copy number of the gene to be tested in the test sample is amplified relative to the entire genome of the test sample.

[0215] With the adoption of ( ) as the amplification judgment value of the gene to be tested in the sample to be tested, it can also be As the amplification judgment value of the gene to be tested in the sample to be tested.

[0216] After obtaining the amplification judgment value of the gene to be tested in the sample to be tested, it can be determined whether the amplification judgment value of the gene to be tested in the sample to be tested is greater than an amplification positive judgment threshold.

[0217] If it is determined to be greater than, a test result indicating that the copy number of the gene to be tested in the test sample is positively amplified can be generated.

[0218] If it is determined that it is not greater than, a test result indicating that the copy number amplification of the gene to be tested in the test sample is negative can be generated.

[0219] In some optional embodiments, the positive judgment threshold for amplification of the gene to be tested can be determined by Figure 4A The amplification positive judgment threshold determination step 400 is predetermined, and the amplification positive judgment threshold determination step 400 may include the following steps: Figure 4A Steps 401 to 405 are shown: Step 401: Acquire an amplified sample data set.

[0220] Here, the amplified sample data includes negative sample data and amplified positive sample data for the gene to be tested, wherein each negative sample data is obtained based on sequencing comparison data of benign samples, and each amplified positive sample data is sequencing comparison data for the copy number amplification of the gene to be tested.

[0221] Optionally, step 401 may include: Figure 4B The following steps 4011 and 4012 are shown: Step 4011, obtain sequencing alignment data of 1 pair of 100% tumor cell lines and their paired cell lines.

[0222] Here, I is a positive integer.

[0223] Here, the specific operation of step 4011 and the technical effect produced are the same as Figure 3A The operation and effect of step 3011 in the illustrated embodiment are substantially the same and will not be described in detail here.

[0224] Step 4012: Based on I pair of sequencing comparison data, generate at least one amplification positive sample data for the gene to be tested.

[0225] Here, the tumor purity of each amplification-positive sample data generated is the corresponding annotated tumor purity.

[0226] Optionally, step 4012 may be performed as follows: for each tumor cell line sequencing comparison data, perform an amplification positive sample data generation operation, which may include: Figure 4CThe following steps 40121 and 40122 are shown: Step 40121: For each preset copy number gradient in the preset copy number gradient set, add the corresponding read segments of the gene to be tested to the sequencing alignment data of the tumor cell line to generate amplification positive sample data of the tumor cell line under the preset copy number gradient.

[0227] Here, the tumor cell copy number of the amplified positive sample data of the tumor cell line under the preset copy number gradient is the preset copy number gradient.

[0228] The reads corresponding to the gene to be tested in the sequencing alignment data of the tumor cell line may include, but are not limited to, reads covering specific regions (such as exons (general), introns, or regulatory regions) of the gene to be tested in the sequencing alignment data.

[0229] That is, the copy number amplification of the gene to be tested at different copy numbers is simulated by adding read segments corresponding to the gene to be tested at different ratios to the sequencing alignment data of the tumor cell line.

[0230] For example, the original sequencing data of the tumor cell line has the corresponding read segments of the gene to be tested. segment, if the preset copy number gradient is , here you can add the sequencing alignment data of the tumor cell line The number of reads corresponding to the genes to be tested. That is, the number of reads to be added Satisfy the following formula:

[0231] Here, the default copy number of the gene to be tested in the tumor cell line is 2, and the preset copy number gradient is simulated on this basis .

[0232] Step 40122: For each preset tumor purity gradient in the preset tumor purity gradient set, based on the sequencing alignment data of the tumor cell line and the sequencing alignment data of the corresponding paired cell line, generate amplification-positive sample data of the tumor cell line under each tumor cell copy number gradient and the preset tumor purity gradient, and determine the preset tumor gradient as the annotated tumor purity of the generated amplification-positive sample data.

[0233] That is, through step 4012, the sequencing comparison data of each tumor cell line is simulated to generate amplification simulation positive data with different tumor purity gradients, and the corresponding tumor purity is annotated for the generated amplification positive sample data.

[0234] For example, the preset tumor purity gradient is , can be randomly selected from the sequencing alignment data of the tumor cell line Read segments are randomly selected from the sequencing alignment data of the paired cell lines of this tumor cell line The two selected reads are mixed to generate amplified positive sample data, and the tumor purity is obtained. Amplification positive sample data.

[0235] Step 4013: Acquire sequencing comparison data of J benign samples, and generate at least one negative sample data based on the sequencing comparison data of the J benign samples.

[0236] Here, J is a positive integer.

[0237] Here, the specific operation of step 4013 and the technical effect produced are the same as Figure 3B The operation and effect of step 3013 in the illustrated embodiment are substantially the same and will not be described in detail here.

[0238] Step 402: Determine the corrected tumor cell copy number of each amplified sample data.

[0239] Here, the same method as that described in steps 202, 203 and 204 for determining the corrected tumor cell copy number of the gene to be tested in the sample to be tested can be used to determine the corrected tumor cell copy number of the gene to be tested for each amplified sample data.

[0240] Specifically, for each amplified sample data, the tumor purity and tumor ploidy of the amplified sample data can first be determined based on the sequencing comparison data of the amplified sample data and at least one baseline sample obtained in step 201 (using the method described in step 202); then, based on the sequencing comparison data of the amplified sample data and at least one baseline sample obtained in step 201, the average mixed copy number of the gene to be tested in the amplified sample data is determined; finally, based on the tumor purity and tumor ploidy of the amplified sample data and the average mixed copy number of the gene to be tested in the amplified sample data, the corrected tumor cell copy number of the gene to be tested in the amplified sample data is determined.

[0241] Step 403 : determining the amplification judgment value of the corresponding amplified sample data for the gene to be tested according to the corrected tumor cell copy number of each amplified sample data.

[0242] Here, various methods can be used to determine the amplification judgment value of each amplification-positive sample data for the gene to be tested based on the corrected tumor cell copy number of the gene to be tested.

[0243] It can be understood that the specific operation of determining the amplification judgment value of the amplified positive sample data for the gene to be tested based on the corrected tumor cell copy number of the gene to be tested based on the amplified positive sample data and the technical effect produced can be basically the same as the operation and effect of determining the amplification judgment value of the gene to be tested in the sample to be tested based on the corrected tumor cell copy number of the gene to be tested in the sample to be tested in step 206, and will not be repeated here.

[0244] Step 404 : For each candidate amplification positive judgment value in the candidate amplification positive judgment value set, perform an amplification positive judgment basis determination operation.

[0245] Here, the candidate amplification positive judgment value set may be pre-set by professional technicians based on the possible value range of amplification positive judgment values ​​in practice.

[0246] Here, the determination operation based on the positive amplification judgment may include the following steps: Figure 4D The following steps 4041 and / or 4042 are shown: Step 4041 : Determine the positive coincidence rate corresponding to the candidate amplification positive judgment value based on the comparison result of the amplification judgment value of each amplification positive sample data for the gene to be tested and the candidate amplification positive judgment value.

[0247] Specifically, step 4041 may be performed as follows: First, based on the comparison result of the amplification positive judgment value of each amplification positive sample data for the gene to be tested and the candidate amplification positive judgment value, the number of true positives and false negatives corresponding to the candidate amplification positive judgment value is counted.

[0248] First, for each amplification-positive sample data, determine whether the amplification-positive judgment value of the amplification-positive sample data for the gene to be tested is greater than the candidate amplification-positive judgment value. If it is determined to be greater than, it can be determined that the amplification-positive sample data is detected as amplification-positive. If it is determined not to be greater than, it can be determined that the amplification-positive sample data is detected as amplification-negative.

[0249] Then, the number of amplification positive sample data detected as amplification positive in each amplification positive sample data is counted as the number of true positives corresponding to the candidate amplification positive judgment value. . Count the number of amplification-positive samples that are detected as amplification-negative, as the number of false negatives corresponding to the candidate amplification-positive judgment value .

[0250] Finally, the positive coincidence rate corresponding to the candidate amplification positive judgment value is determined based on the number of true positives and the number of false negatives corresponding to the candidate amplification positive judgment value.

[0251] here, The Positive Percent Agreement (Positive Percent Agreement) can be calculated using the following formula:

[0252] in, and are the number of true positives and false negatives corresponding to the candidate amplification positive judgment value, It is the calculated positive coincidence rate corresponding to the amplification positive judgment value.

[0253] Step 4042: Determine a positive prediction value corresponding to the candidate amplification positive judgment value based on a comparison result of the amplification judgment value of each amplification sample data for the gene to be tested and the candidate amplification positive judgment value.

[0254] Specifically, step 4042 may be performed as follows: First, based on the comparison result of the amplification positive judgment value of each amplification positive sample data for the gene to be tested and the candidate amplification positive judgment value, the number of true positives corresponding to the candidate amplification positive judgment value is counted.

[0255] For details, please refer to the relevant description in step 4041, which will not be repeated here.

[0256] Then, based on the comparison results of the amplification positive judgment value of each negative sample data for the gene to be tested and the candidate amplification positive judgment value, the number of false positives corresponding to the candidate amplification positive judgment value is counted. First, for each negative sample data, determine whether the amplification positive judgment value of the negative sample data for the gene to be tested is greater than the candidate amplification positive judgment value. If it is determined to be greater than, it can be determined that the negative sample data is detected as amplification positive. If it is determined not to be greater than, it can be determined that the negative sample data is detected as amplification negative. Then, the number of negative sample data detected as amplification positive in each negative sample data is counted as the number of false positive data corresponding to the candidate amplification positive judgment value. .

[0257] Finally, the positive predictive value corresponding to the candidate amplification positive judgment value is determined based on the number of true positives and the number of false positives corresponding to the candidate amplification positive judgment value.

[0258] Specifically, for a candidate amplification positive judgment value, the positive predictive value corresponding to the candidate amplification positive judgment value can be calculated according to the following formula: (Positive Predictive Value):

[0259] in, and are the number of true positives and false positives corresponding to the candidate amplification positive judgment value, is the calculated positive predictive value corresponding to the candidate amplification positive judgment value.

[0260] Step 405 : Determine the amplification positive judgment threshold of the gene to be tested according to the positive prediction value corresponding to each candidate amplification positive judgment value.

[0261] Here, a preferred amplification positive judgment value whose corresponding positive coincidence rate and / or positive predictive value meets a preset detection standard can be first determined from the set of candidate amplification positive judgment values. The preset detection standard can be pre-set according to the needs of a specific detection application scenario. As an example, the preset detection standard can be: PPV ≥ 99% and PPA ≥ 95%.

[0262] Then, the amplification positive judgment threshold of the gene to be tested is determined based on the above-determined preferred amplification positive judgment values. For example, the minimum value among the determined preferred amplification positive judgment values ​​can be determined as the amplification positive judgment threshold of the gene to be tested.

[0263] The optional methods of steps 401 to 405 above can achieve the following technical effects including but not limited to: First, by generating amplification-positive sample data for the gene to be tested based on sequencing comparison data of a small number of 100% tumor cell lines and their paired cell lines, the sequencing cost required to generate amplification-positive sample data is reduced.

[0264] Secondly, by generating a large amount of negative sample data based on sequencing comparison data of a small number of benign samples, the sequencing cost of benign samples is reduced.

[0265] In addition, by additionally considering the tumor ploidy of the amplified positive sample data and the negative sample data in the process of determining the corrected tumor cell copy number of the amplified sample data for the gene to be tested, the interference of aneuploidy (polyploidy or haploidy) on the test results can be eliminated.

[0266] In some optional embodiments, before step 202, the above method may further perform the following base site sequencing depth normalization operation for each sequencing alignment data: First, for each preset capture region, the sequencing depth of each base site in the preset capture region of the sequencing alignment data is determined.

[0267] Then, the sum of the sequencing depths of all base sites in each preset capture region of the sequencing alignment data is determined as the total sequencing depth of the sequencing alignment data.

[0268] Finally, for each base site in each preset capture region, the sequencing depth of the base site is normalized based on the total sequencing depth of the sequencing alignment data.

[0269] Specifically, for each base site in each preset capture region, the sequencing depth of the base site divided by the ratio of the total sequencing depth of the sequencing alignment data multiplied by the global scaling factor can be used as the sequencing depth of the base site after normalization.

[0270] By adopting the above optional implementation method, the difference in sequencing data volume between different samples can be eliminated.

[0271] The above-mentioned embodiments of the present disclosure provide a method for detecting gene copy number variation types, which first obtains sequencing comparison data of gene sequencing of a test sample and at least one baseline sample, wherein the capture region of the gene sequencing process includes a preset capture region set; then, based on each sequencing comparison data, determines the tumor purity and tumor ploidy of the test sample and the average mixed copy number of the test gene in the test sample; then, based on the tumor purity and tumor ploidy of the test sample and the average mixed copy number of the test gene in the test sample, determines the corrected tumor cell copy number of the test gene in the test sample; next, if If the average mixed copy number of the gene to be tested is less than 2, the test result for homozygous deletion of the gene to be tested is determined based on the comparison of the corrected tumor cell copy number of the gene to be tested in the test sample with the positive judgment threshold for homozygous deletion of the gene to be tested. If the average mixed copy number of the gene to be tested is greater than 2, the amplification judgment value of the gene to be tested in the test sample is determined based on the corrected tumor cell copy number of the gene to be tested in the test sample, and the test result for copy number amplification of the gene to be tested is determined based on the comparison of the amplification judgment value of the gene to be tested in the test sample with the positive judgment threshold for amplification. Thus, by additionally considering the tumor ploidy of the test sample in the process of determining the corrected tumor cell copy number of the gene to be tested in the test sample, the accuracy of the corrected tumor cell copy number of the gene to be tested in the test sample can be improved due to the richer factors considered. In turn, the accuracy of detecting the copy number variation status of the gene to be tested can be improved.

[0272] Example 1: use Figure 3A The steps for determining the homozygous deletion positive judgment threshold are as follows: 1. Select two pairs of cell lines: a tumor cell line with 100% tumor content and its paired cell line. Perform high-throughput sequencing, align the raw sequencing data with the human hg19 genome, and obtain sequencing alignment data (BAM files). This yields four BAM files: two tumor cell line sequencing alignment data (BAM files) and two paired cell line sequencing alignment data (BAM files).

[0273] 2. Generate homozygous deletion positive sample data: (1) For the 9 exons of the PTEN gene, a total of 45 types of continuous exon homozygous deletions to be investigated were simulated, including: 9 types of single exon deletions, 8 types of continuous double exon deletions, 7 types of continuous three exon deletions, ..., and 1 type of whole gene deletion.

[0274] That is, for each tumor cell line sequencing alignment data set, by removing one consecutive exon, nine different homozygous deletion-positive sample data sets can be obtained; by removing two consecutive exons, eight different homozygous deletion-positive sample data sets can be obtained; and so on, until nine consecutive exons are removed, one homozygous deletion-positive sample data set can be obtained. In total, 45 homozygous deletion-positive sample data sets of consecutive exon homozygous deletion types to be examined can be obtained. Since there are two tumor cell line sequencing alignment data sets, 90 homozygous deletion-positive sample data sets (2 tumor cell lines * 45 homozygous deletion types) can be obtained (BAM files).

[0275] (2) Simulating different tumor proportion gradients: For each of the 90 homozygous deletion-positive sample data, reads were randomly sampled from the sequencing alignment data of its paired cell line and mixed into different gradients of tumor proportion, specifically including the following 8 tumor proportion gradients: 20%, 30%, 40%, 50%, 60%, 70%, 80%, and 90%. In addition, for different types of homozygous deletions of consecutive exons, the above tumor gradient mixing operation was repeated different times, generating a total of 3744 homozygous deletion-positive sample data. The specific distribution of PTEN homozygous deletion-positive sample data is shown in Table 1.

[0276] Table 1 PTEN deletion positive sample data

[0277] As can be seen from Table 1, 10 tumor gradient mixing operations were performed for each homozygous deletion-positive sample data with consecutive exon deletions of 1, 3, 6, and 9. In other words, mixing was performed 10 times according to each gradient ratio. However, because reads were randomly sampled from the sequencing alignment data of the paired cell lines each time, the 10 mixed homozygous deletion-positive sample data generated by the 10 tumor gradient mixing operations for the same consecutive exon deletion type and the same tumor ratio gradient were also different.

[0278] For the remaining homozygous deletion types, one tumor gradient mixing operation was performed. This is because if the tumor gradient mixing operation is repeated ten times for all deletion types, the amount of data will be too large.

[0279] 3. Negative sample data generation: 85 benign tissue samples with normal PTEN expression and no tumor cells were selected for high-throughput sequencing. The raw sequencing data were aligned with the human hg19 genome, and sequencing alignment data (BAM files) were obtained. In other words, 85 sequencing alignment data sets were obtained.

[0280] For the 85 sequencing alignments, noise was then added to the sequencing depths of base sites at different noise gradients to simulate varying noise levels. Specifically, four noise gradients were used: 25%, 33%, 50%, and 100%. This yielded 340 negative sample data with added noise. Together with the original 85 negative data without added noise, this yielded a total of 425 negative sample data. Detailed information about the negative sample data can be found in Table 2.

[0281] Table 2 PTEN deletion negative data table

[0282] 4. Use the above Figure 3A In the method of step 304 and step 305 in the illustrated embodiment, for each of the 3744 homozygous deletion positive sample data and the 425 negative sample data, the corrected tumor cell copy number for PTEN is determined.

[0283] 5. Use the above Figure 3A In the illustrated embodiment, the method of step 306 performs a homozygous deletion positive predictive value determination operation for each candidate homozygous deletion copy number in the candidate homozygous deletion copy number set, ultimately determining a homozygous deletion positive threshold of 0.6 for the gene to be tested, PTEN. Table 3 shows the PPA and PPV values ​​for 425 negative sample data for homozygous deletion-positive sample data of different tumor proportions and different homozygous deletion types, when the candidate homozygous deletion copy number is 0.6.

[0284] Table 3 Detection performance of different PTEN gene deletion types and tumor purity

[0285] As can be seen from Table 3, with the tumor proportion (also known as tumor purity) of 0.3 as the cutoff point, when the tumor proportion is lower than 0.3 (among the 3744 homozygous deletion-positive sample data, the tumor proportion lower than 0.3 is 3744÷8=468 homozygous deletion-positive sample data with a 20% tumor proportion gradient), all deletion types (i.e., 468 homozygous deletion-positive sample data) cannot meet PPV≥99% and PPA≥95%.

[0286] When the tumor proportion was greater than or equal to 0.3 (including 3744-468 ​​= 3276 homozygous deletion-positive sample data), the positive judgment value criteria (PPV ≥ 99% and PPA ≥ 95%) could not be met only when one exon was deleted (including 2 (2 tumor cell lines) * 9 (9 types of deleted exons) * 10 (10 repeated tumor proportion mixing operations) * 7 (7 tumor proportion gradients with tumor proportion greater than or equal to 0.3) = 1260 homozygous deletion-positive sample data); for other deletion types (including 3276-1260 = 2016 homozygous deletion-positive sample data), PPV ≥ 99% and PPA ≥ 95% were met.

[0287] 6. Confirmation of the minimum detection limit: The corrected tumor cell copy number GCN ≤ 0.6 was used as the threshold for positive judgment of homozygous deletion of PTEN. Probit analysis was performed for 1-2 exon deletion type and 3-9 exon deletion type respectively. Probit curves were drawn to obtain Figure 5 The two Probit curves in . Figure 5 The left graph corresponds to 1-2 exon deletion types, and the right graph corresponds to 3-9 exon deletion types. The horizontal axis of each Probit curve corresponds to the tumor proportion, and the vertical axis is PPA. Figure 5 As shown in Table 4, when PTEN gene deletions occur in 1-2 exons, the minimum detection limit is 43.61%; when PTEN deletions occur in 3-9 exons, the minimum detection limit is 31.71%. The minimum detection limits for different deletion types are shown in Table 4.

[0288] Table 4 Minimum detection limits for different deletion types of PTEN gene

[0289] 7. Third-party IHC verification: 97 cancer tissue samples were selected and the Figure 2AThe described method detects homozygous deletion of the PTEN gene according to the threshold for homozygous deletion positivity determined above (i.e., 0.6). This method was also validated using a third-party IHC method. The IHC criteria are: an IHC staining index of 0 indicates IHC-positive PTEN homozygous deletion; otherwise, IHC-negative PTEN homozygous deletion. The criteria for this example are: a corrected tumor cell copy number greater than or equal to 0.6 indicates PTEN homozygous deletion positivity; otherwise, PTEN homozygous deletion positivity. Results showed a high concordance of 85.57% between the detection method of this example and IHC. The concordance results are shown in Table 5.

[0290] Table 5: Consistency between the test results disclosed in this paper and IHC

[0291] Example 2: use Figure 3A The steps for determining the homozygous deletion positive judgment threshold are shown as follows: determining the homozygous deletion positive judgment threshold of the genes to be tested MTAP, CDKN2A and CDKN2B, and the specific steps are as follows: 1. Obtain sequencing data: Select 13 pairs of cell lines, including tumor cell lines with 100% tumor proportion and their paired cell lines, for high-throughput sequencing. Align the raw sequencing data with the human hg19 genome, and obtain sequencing alignment data (BAM files).

[0292] 2. Generate homozygous deletion positive sample data (1) The sequencing comparison data of the above 13 tumor cell lines were removed respectively from the gene reads of the genes to be tested, MTAP, CDKN2A and CDKN2B, and the homozygous deletion positive sample data for MTAP, CDKN2A and CDKN2B were generated respectively, that is, 13*3=39 homozygous deletion positive sample data were obtained.

[0293] (2) Simulating different tumor proportion gradients: The 39 homozygous deletion-positive sample data were randomly sampled from the sequencing alignment data (BAM files) of their paired cell lines and mixed into different gradients of tumor proportion, specifically including five tumor proportion gradients: 20%, 30%, 35%, 40%, and 50%. The above tumor proportion gradient mixing operation was repeated 12 times for each homozygous deletion-positive sample data, resulting in 780*3=2340 homozygous deletion-positive sample data. The distribution of homozygous deletion-positive sample data for MTAP, CDKN2A, and CDKN2B genes is shown in Table 6.

[0294] Table 6 Data table of positive samples with MTAP / CDKN2A / CDKN2B deletion

[0295] 3. Negative sample data generation: 86 benign tissue samples with normal expression of the MTAP, CDKN2A, and CDKN2B genes, as determined by pathological assessment, were selected. These samples, along with the paired cell lines from the aforementioned 13 cell line pairs, were subjected to high-throughput sequencing. The raw sequencing data were aligned with the human hg19 genome, and sequencing alignment data (BAM files) were generated. This resulted in sequencing alignment data for 99 negative samples, also referred to as 99 negative sample data. The distribution of negative sample data is shown in Table 7.

[0296] Table 7 Negative data of MTAP / CDKN2A / CDKN2B gene deletion

[0297] 4. Use the above Figure 3A In the method of step 304 and step 305 in the illustrated embodiment, for each of the 2340 homozygous deletion positive sample data and the 99 negative sample data, the corrected tumor cell copy number for MTAP, CDKN2A, and CDKN2B is determined.

[0298] 5. Use the above Figure 3A In the method of step 306 in the embodiment shown, for each candidate homozygous deletion copy number in the candidate homozygous deletion copy number set, a homozygous deletion positive prediction value determination operation is performed, and the homozygous deletion positive judgment threshold of the tested genes MTAP, CDKN2A and CDKN2B is finally determined to be 0.6.

[0299] Table 8 shows the PPA and PPV values ​​of 99 negative sample data corresponding to the number of homozygous deletion-positive sample data for different genes to be tested and different tumor proportions when the candidate homozygous deletion copy number is 0.6.

[0300] Table 8 Detection performance of MTAP and CDKN2A genes when the positive judgment value was 0.6

[0301] As can be seen from Table 8, with the tumor proportion (also known as tumor purity) of 0.3 as the cutoff point, when the tumor proportion is lower than 0.3, the three tested genes MTAP, CDKN2A, and CDKN2B cannot meet the positive judgment value standards PPV ≥ 99% and PPA ≥ 95%.

[0302] When the tumor ratio was greater than or equal to 0.3, all three genes tested met PPV ≥ 99% and PPA ≥ 95%. Therefore, a corrected tumor cell copy number GCN ≤ 0.6 was selected as the threshold for determining homozygous deletions of MTAP, CDKN2A, and CDKN2B.

[0303] 6. Confirmation of the minimum detection limit: The corrected tumor cell copy number GCN≤0.6 was used as the threshold for positive judgment of homozygous deletion of MTAP, CDKN2A and CDKN2B. Probit analysis was performed on the three genes to be tested, MTAP, CDKN2A and CDKN2B, and Probit curves were drawn to obtain the following results: Figure 6A 、 Figure 6B and Figure 6C The Probit curve in is used to evaluate the minimum detection limit of sample data. Figure 6A 、 Figure 6B and Figure 6C The Probit curves in the figure correspond to the Probit curves of the three genes to be tested, MTAP, CDKN2A and CDKN2B. The horizontal axis of the Probit curve is the tumor proportion and the vertical axis is PPA.

[0304] like Figure 6A 、 Figure 6B and Figure 6C As shown, the minimum detection limit of MTAP was 27%, the minimum detection limit of CDKN2A was 21%, and the minimum detection limit of CDKN2B was 21%.

[0305] 7. Confirmation of the incidence of MTAP homozygous deletion: Sequencing data of 2401 cancer tissues were selected and used Figure 2A The method shown in this example detects homozygous deletion of the MTAP gene using the threshold for positive homozygous deletion (i.e., 0.6) determined above. The detection rate of this example method is then compared with the incidence rates of various cancers in the cBiPortal public database to assess whether it meets expectations. The criteria used in this example are: a tumor cell copy number greater than or equal to 0.6 is considered positive for MTAP homozygous deletion; otherwise, it is considered negative for MTAP homozygous deletion. The results demonstrate that the incidence rates of MTAP homozygous deletion detected by this example method are generally consistent with those in the public database, meeting expectations. The consistency results are shown in Table 9.

[0306] Table 9 Incidence of homozygous deletion of MTAP gene

[0307] In Table 9, the 27% (28 / 105) incidence of homozygous deletion of the MTAP gene in this example means the following: The denominator 105 indicates that a total of 105 samples diagnosed with bladder urothelial cancer were tested using the method of this embodiment; Molecule 27 indicates that 28 samples were found to be positive for MTAP homozygous deletion using the method of this example.

[0308] Based on Example 1 and Example 2, it can be seen that the gene copy number variation type detection method provided by the embodiments of the present disclosure performs well in detecting homozygous deletions of the gene to be tested.

[0309] Example 3: use Figure 4A The amplification positive judgment threshold determination step shown in the figure determines the amplification positive judgment threshold of the gene to be tested CCNE1, and the specific steps are as follows: 1. Obtaining Sequencing Data: Seven cell line pairs were selected, each containing a 100% tumor cell line and its corresponding paired cell line. High-throughput sequencing was performed, and the raw sequencing data were aligned with the human hg19 genome. Sequencing alignment data (BAM files) were obtained. These seven cell line pairs included three diploids (i.e., with a Ploidy of 2), two triploids (i.e., with a Ploidy of 3), and four tetraploids (i.e., with a Ploidy of 4).

[0310] 2. Generate amplification positive sample data (1) For the sequencing comparison data of the seven tumor cell lines mentioned above, the read segments of the test gene CCNE1 were inserted according to different amplification ratios to generate amplification-positive sample data for the test gene CCNE1, that is, 7 amplification-positive sample data were obtained. Specifically, there are 4 different amplification copy numbers: 5, 6, 7, and 8, and thus 7*4=28 amplification-positive sample data were obtained.

[0311] (2) Simulating different tumor proportion gradients: The 28 amplification-positive sample data were randomly sampled from the sequencing alignment data (BAM files) of their paired cell lines and mixed into different gradients of tumor proportion, including four tumor proportion gradients: 15%, 20%, 25%, and 30%. The above tumor proportion gradient mixing operation was repeated 7 times for each amplification-positive sample data, resulting in 7 (7 tumor cell lines) * 4 (4 simulated copy number gradients) * 4 (4 tumor proportion gradients) * 7 (repeated 7 times tumor proportion mixing) = 784 amplification-positive sample data. The distribution of CCNE1 gene amplification-positive sample data is shown in Table 10.

[0312] Table 10 Data of positive samples of CCNE1 gene copy number amplification

[0313] 3. Negative sample data generation: 78 benign tissue samples with normal CCNE1 gene expression, as determined by pathological assessment and lacking tumor cells, were selected for high-throughput sequencing. The raw sequencing data were aligned with the human hg19 genome, and sequencing alignment data (BAM files) were obtained. In other words, data from 78 negative samples were obtained.

[0314] 4. Use the above Figure 4A In the illustrated embodiment, steps 404 and 405 determine an amplification judgment value for the gene CCNE1 to be tested for each of the 784 homozygous deletion-positive sample data and the 78 homozygous deletion-negative sample data. The amplification judgment value is the difference between the corrected tumor cell copy number and the tumor ploidy, i.e., (GCN-ploidy).

[0315] 5. Use the above Figure 4A In the method of step 406 in the embodiment shown, for each candidate amplification positive judgment value in the candidate amplification positive judgment value set, an amplification positive prediction value determination operation is performed, and the amplification positive judgment threshold of the gene to be tested CCNE1 is finally determined to be 2.5.

[0316] Table 11 shows the number of true positives and false negatives detected for 784 amplification-positive samples with different tumor cell copy numbers and tumor proportions when the candidate amplification-positive judgment value is 2.5, as well as the number of true negatives and false positives detected for 78 negative samples, and the corresponding PPA and PPV values.

[0317] Table 11 Detection performance of the test gene CCNE1 when the positive judgment value is 2.5

[0318] As shown in Table 11, when the candidate amplification positive judgment value is 2.5 and the tumor proportion is greater than or equal to 0.3, the amplification positive detection of the test gene CCNE1 meets the PPV ≥ 99% and PPA ≥ 95%. Therefore, the amplification positive judgment threshold of the test gene CCNE1 can be 2.5.

[0319] 6. Confirmation of the minimum detection limit: The difference between the corrected tumor cell copy number and the tumor ploidy (GCN-ploidy) ≥ 2.5 is used as the threshold for positive amplification of the test gene CCNE1. Probit analysis is performed on the test gene CCNE1, and the Probit curve is drawn to obtain Figure 7A 、 Figure 7B and Figure 7C The Probit curve in is used to evaluate the minimum detection limit of sample data. Figure 7A 、 Figure 7B and Figure 7CThe Probit curves in the figure correspond to the Probit curves of diploid CCNE1 amplification, triploid CCNE1 amplification, and tetraploid CCNE1 amplification, respectively. In the Probit curves, the horizontal axis is the tumor proportion and the vertical axis is PPA.

[0320] like Figure 7A 、 Figure 7B and Figure 7C As shown, the minimum detection limits of diploid, triploid and tetraploid CCNE1 genes to be tested are 24%, 20% and 16% respectively. Therefore, it can be determined that the minimum detection limit of CCNE1 genes to be tested is 24%.

[0321] 7. Confirmation of CCNE1 copy number amplification incidence: Sequencing data of 1528 breast cancer samples were selected and used Figure 2A The method shown in this example detected CCNE1 gene copy number amplification using the previously determined threshold for positive CCNE1 gene amplification (i.e., 2.5). The detection rate of this example method was then compared with the breast cancer incidence rate in the cBiPortal public database to examine whether it met expectations. The results showed that the CCNE1 gene copy number amplification rate detected by this example method was 5.24%, which is basically consistent with the public database (3.04%) and meets expectations.

[0322] As can be seen from Example 3, the gene copy number variation type detection method provided in the embodiments of the present disclosure performs well in detecting the amplification of the copy number of the gene to be tested.

[0323] Further references Figure 8 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a gene copy number variation type detection device. Figure 2A Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0324] like Figure 8As shown, the gene copy number variation type detection device 800 of this embodiment includes: a sequencing comparison data acquisition module 801, a purity and ploidy determination module 802, an average mixed copy number determination module 803, a tumor cell copy number correction module 804, a homozygous deletion detection module 805 and a copy number amplification detection module 806. Among them, the sequencing comparison data acquisition module 801 is configured to obtain sequencing comparison data of gene sequencing performed on the test sample and at least one baseline sample respectively, wherein each sequencing comparison data includes sequencing comparison data of a preset set of gene analysis regions to be tested, and the preset set of gene analysis regions to be tested includes the region where the gene to be tested is located; the purity and ploidy determination module 802 is configured to determine the tumor purity and tumor ploidy of the test sample based on each sequencing comparison data; the average mixed copy number determination module 803 is configured to determine the average mixed copy number of the gene to be tested in the test sample based on each sequencing comparison data; the tumor cell copy number correction module 804 is configured to determine the corrected tumor cell copy number of the gene to be tested in the test sample based on the tumor purity and tumor ploidy of the test sample and the average mixed copy number of the gene to be tested in the test sample; The homozygous deletion detection module 805 is configured to generate a detection result of the homozygous deletion of the gene to be tested in the test sample in response to the average mixed copy number of the gene to be tested in the test sample being less than 2, based on the comparison result of the corrected tumor cell copy number of the gene to be tested in the test sample and the homozygous deletion positive judgment threshold of the maximum value of the homozygous deletion positive copy number of the gene to be tested; and / or, the copy number amplification detection module 806 is configured to determine the amplification judgment value of the gene to be tested in the test sample in response to the average mixed copy number of the gene to be tested in the test sample being greater than 2, based on the corrected tumor cell copy number of the gene to be tested in the test sample, and generate a detection result of the copy number amplification of the gene to be tested in the test sample in response to the average mixed copy number of the gene to be tested in the test sample being greater than 2, based on the correction of the tumor cell copy number of the gene to be tested in the test sample, and based on the comparison result of the amplification judgment value of the gene to be tested in the test sample and the amplification positive judgment threshold of the minimum value of the amplification positive judgment value of the gene to be tested.

[0325] In this embodiment, the specific processing and technical effects of the sequencing comparison data acquisition module 801, the purity and ploidy determination module 802, the average mixed copy number determination module 803, the tumor cell copy number correction module 804, the homozygous deletion detection module 805 and the copy number amplification detection module 806 of the gene copy number variation type detection device 800 can be referred to respectively. Figure 2A The relevant descriptions of step 201, step 202, step 203, step 204, step 205 and step 206 in the corresponding embodiment are not repeated here.

[0326] In some optional embodiments, the tumor cell copy number correction module 804 may be further configured to: The tumor purity and tumor ploidy of the sample to be tested, and the average mixed copy number of the gene to be tested in the sample to be tested are substituted into the corrected tumor cell copy number solution to obtain the corrected tumor cell copy number of the gene to be tested.

[0327] In some optional embodiments, the corrected tumor cell copy number solution formula may be obtained by deriving a correction formula, wherein the correction formula is: ; The formula for calculating the corrected tumor cell copy number is: ; in, and are respectively the tumor purity and tumor ploidy of the sample to be tested, is the average mixed copy number of the gene to be tested in the sample to be tested, is the corrected tumor cell copy number of the gene to be tested in the sample to be tested.

[0328] In some optional embodiments, determining the amplification judgment value of the gene to be tested in the test sample according to the corrected tumor cell copy number of the gene to be tested in the test sample may include: Determining the corrected tumor cell copy number of the gene to be tested in the sample to be tested as the amplification judgment value of the gene to be tested in the sample to be tested, or, The difference between the corrected tumor cell copy number of the gene to be tested in the sample to be tested and the tumor ploidy of the sample to be tested is determined as the amplification judgment value of the gene to be tested in the sample to be tested, or, The amplification judgment value of the gene to be tested in the sample to be tested is determined as the ratio of the corrected tumor cell copy number of the gene to be tested in the sample to be tested divided by the tumor ploidy of the sample to be tested.

[0329] In some optional embodiments, the purity and ploidy determination module 802 may be further configured to: Determine the abundance of each SNP site in the preset whole-genome SNP site set of the sample to be tested based on the sequencing comparison data of the sample to be tested; For each predetermined whole-genome SNP site, the related gene analysis region is divided into at least one SNP analysis region according to a first predetermined length; For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each SNP analysis region; For each SNP analysis region, the mixed copy number of the sequencing alignment data of the test sample in the SNP analysis region is determined based on the corrected average coverage depth of the sequencing alignment data of the test sample in the SNP analysis region and the average of the corrected average coverage depth of the sequencing alignment data of each baseline sample in the SNP analysis region; Based on the abundance and mixed copy number of each SNP site in the sequencing comparison data of the sample to be tested, the tumor purity and tumor ploidy of the sample to be tested are determined, wherein the mixed copy number of each SNP site is the mixed copy number of the SNP analysis region where the SNP site is located.

[0330] In some optional embodiments, the average mixed copy number determination module 803 may be further configured to: For each preset gene analysis region to be tested, dividing the preset gene analysis region to be tested into at least one gene analysis window according to a second preset length; For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each gene analysis window; For each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, determine the reference baseline sample sequencing alignment data corresponding to the gene analysis window in the sequencing alignment data of each baseline sample, as well as the corrected average baseline sample sequencing depth and the corrected average baseline sample sequencing depth deviation of the gene analysis window; Determining a reference gene analysis window in each gene analysis window based on the corrected baseline sample average sequencing depth deviation of each gene analysis window; Determine each reference gene analysis window where the gene to be tested is located as the gene analysis window to be tested; For each gene analysis window to be tested, determining the mixed copy number of the gene analysis window to be tested in the sample to be tested based on the corrected average coverage depth of the sequencing alignment data of the sample to be tested in the gene analysis window to be tested and the corrected average sequencing depth of the baseline sample in the gene analysis window to be tested; The mean of the mixed copy numbers of each analysis window of the gene to be tested in the sample to be tested is determined as the average mixed copy number of the gene to be tested in the sample to be tested.

[0331] In some optional embodiments, for each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, determining the reference baseline sample sequencing alignment data corresponding to the gene analysis window, as well as the corrected baseline sample average sequencing depth and the corrected baseline sample average sequencing depth deviation of the gene analysis window in the sequencing alignment data of each baseline sample may include: For each gene analysis window, perform the following baseline sample gene analysis window depth correction operation: Determine the normal range of the average sequencing depth of the baseline samples based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window; Determine the sequencing alignment data of each baseline sample whose average coverage depth after correction in the gene analysis window falls within the normal range of the average sequencing depth of the baseline samples as the reference baseline sample sequencing alignment data corresponding to the gene analysis window; Based on the corrected average coverage depth of the sequencing alignment data of each reference baseline sample in the gene analysis window, the corrected average sequencing depth of the baseline samples and the corrected average sequencing depth deviation of the baseline samples in the gene analysis window are determined.

[0332] In some optional embodiments, determining, for each sequencing alignment data, the corrected average coverage depth of the sequencing alignment data for each gene analysis window may include: For each sequencing alignment data, each gene analysis window is used as the target analysis region, and the following target analysis region sequencing depth correction operation is performed to obtain the corrected average coverage depth of each gene analysis window. The target analysis region sequencing depth correction operation includes: The sequencing depth of each base site in each target analysis region is normalized according to the total sequencing depth of the sequencing alignment data to obtain the corresponding normalized sequencing depth; For each target analysis region, determining the average sequencing depth of the sequencing alignment data in the target analysis region based on the normalized sequencing depth of each base position in the target analysis region; Performing a local weighted regression correction on the average coverage depth of each target analysis region of the sequencing alignment data according to the number of probes covered in each target analysis region; According to the GC content ratio of the sequencing alignment data in each target analysis region, a local weighted regression correction is performed on the average coverage depth of the sequencing alignment data in each target analysis region to obtain the corrected average coverage depth of the corresponding target analysis region.

[0333] In some optional embodiments, determining, for each sequencing alignment data, the corrected average coverage depth of the sequencing alignment data for each SNP analysis region may include: For each sequencing alignment data, each SNP analysis region is used as a target analysis region, and a sequencing depth correction operation of the target analysis region is performed to obtain the corrected average coverage depth of each SNP analysis region.

[0334] In some optional embodiments, the apparatus 800 further includes a standardization module ( Figure 8 ), is configured to, before determining the tumor purity and tumor ploidy of the sample to be tested based on each sequencing comparison data: For each sequencing alignment data, the following base site sequencing depth normalization operation is performed: for each preset gene analysis region to be tested, the sequencing depth of each base site of the sequencing alignment data in the preset gene analysis region to be tested is determined; the sum of the sequencing depths of all base sites of the sequencing alignment data in each of the preset gene analysis regions to be tested is determined as the total sequencing depth of the sequencing alignment data; for each base site in each preset gene analysis region to be tested, the sequencing depth of the base site is normalized based on the total sequencing depth of the sequencing alignment data.

[0335] In some optional embodiments, the homozygous deletion positive judgment threshold of the gene to be tested can be predetermined by the following homozygous deletion positive judgment threshold determination step: Obtaining a homozygous deletion sample data set, wherein the homozygous deletion samples include negative sample data and homozygous deletion positive sample data for the gene to be tested, wherein each negative sample data is obtained based on sequencing comparison data of a benign sample, and each homozygous deletion positive sample data is sequencing comparison data of a homozygous deletion occurring in the gene to be tested; Determining the corrected tumor cell copy number of the gene to be tested for each homozygous deletion sample data; For each candidate homozygous deletion copy number in the candidate homozygous deletion copy number set, performing the following homozygous deletion positive judgment basis determination operation: determining a positive coincidence rate corresponding to the candidate homozygous deletion copy number based on a comparison result of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each of the homozygous deletion positive sample data; and / or determining a positive predictive value corresponding to the candidate homozygous deletion copy number based on a comparison result of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each of the homozygous deletion sample data; The positive judgment threshold for homozygous deletion of the gene to be tested is determined based on the positive coincidence rate and / or positive predictive value corresponding to each candidate homozygous deletion copy number.

[0336] In some optional embodiments, obtaining a homozygous deletion sample data set may include: Obtaining sequencing alignment data of N pairs of 100% tumor cell lines and their paired cell lines, wherein N is a positive integer; Based on the N pairs of sequencing comparison data, generating at least one homozygous deletion positive sample data for the gene to be tested, wherein the tumor purity of each homozygous deletion positive sample data is the corresponding annotated tumor purity; Sequencing comparison data of M benign samples are obtained, and at least one negative sample data is generated based on the sequencing comparison data of the M benign samples.

[0337] In some optional embodiments, determining the corrected tumor cell copy number of the gene to be tested for each homozygous deletion sample data may include: For each homozygous deletion sample data, determine the tumor purity and tumor ploidy corresponding to the homozygous deletion sample data and the average mixed copy number for the gene to be tested; based on the tumor purity and tumor ploidy corresponding to the homozygous deletion sample data and the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the homozygous deletion sample data for the gene to be tested.

[0338] In some optional embodiments, the comparing the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number based on the homozygous deletion positive sample data may include: According to the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number of each homozygous deletion positive sample data, the number of true positives and the number of false negatives corresponding to the candidate homozygous deletion copy number are counted; The positive coincidence rate corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and the number of false negatives corresponding to the candidate homozygous deletion copy number.

[0339] In some optional embodiments, determining the positive predictive value corresponding to the candidate homozygous deletion copy number based on the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number for each of the homozygous deletion sample data may include: According to the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number of each homozygous deletion positive sample data, the number of true positives corresponding to the candidate homozygous deletion copy number is counted; According to the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number in each of the negative sample data, counting the number of false positives corresponding to the candidate homozygous deletion copy number; and A positive predictive value corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and the number of false positives corresponding to the candidate homozygous deletion copy number.

[0340] In some optional embodiments, generating at least one negative sample data based on the sequencing comparison data of the M benign samples may include: For each sequencing alignment data of a benign sample, different degrees of noise are added to the sequencing depth of the base sites in the sequencing alignment data of the benign sample to obtain at least one negative sample data corresponding to the sequencing alignment data of the benign sample.

[0341] In some optional embodiments, the amplification positive judgment threshold of the gene to be tested can be predetermined by the following amplification positive judgment threshold determination step: Obtaining a set of amplified sample data, wherein the amplified sample data includes negative sample data and amplified positive sample data for the gene to be tested, wherein each negative sample data is obtained based on sequencing comparison data of a benign sample, and each amplified positive sample data is sequencing comparison data for copy number amplification of the gene to be tested; Determine the corrected tumor cell copy number for each amplified sample data; Determining the amplification judgment value of the corresponding amplified sample data for the gene to be tested according to the corrected tumor cell copy number of each amplified sample data; For each candidate amplification positive judgment value in the candidate amplification positive judgment value set, performing the following amplification positive judgment basis determination operation: determining a positive coincidence rate corresponding to the candidate amplification positive judgment value based on a comparison result of the amplification judgment value of each amplification positive sample data for the gene to be tested and the candidate amplification positive judgment value; and / or determining a positive predictive value corresponding to the candidate amplification positive judgment value based on a comparison result of the amplification judgment value of each amplification sample data for the gene to be tested and the candidate amplification positive judgment value; The amplification positive judgment threshold of the gene to be tested is determined according to the positive coincidence rate and / or positive predictive value corresponding to each candidate amplification positive judgment value.

[0342] In some optional embodiments, obtaining the amplified sample data set may include: Obtaining sequencing alignment data of I pairs of 100% tumor cell lines and their paired cell lines, where I is a positive integer; Based on the I pair of sequencing comparison data, generating at least one amplification-positive sample data for the gene to be tested, wherein the tumor purity of each amplification-positive sample data is the corresponding annotated tumor purity; Sequencing comparison data of J benign samples are obtained, and at least one negative sample data is generated based on the sequencing comparison data of the J benign samples.

[0343] In some optional embodiments, determining the corrected tumor cell copy number of each amplified sample data may include: For each amplified sample data, determine the tumor purity and tumor ploidy corresponding to the amplified sample data and the average mixed copy number for the gene to be tested; based on the tumor purity and tumor ploidy corresponding to the amplified sample data and the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the gene to be tested in the amplified sample data.

[0344] In some optional embodiments, the determining of the positive coincidence rate corresponding to the candidate amplification positive judgment value based on the comparison result of the amplification judgment value of the gene to be tested with the candidate amplification positive judgment value for each of the amplification positive sample data may include: According to the comparison result of the amplification judgment value of each amplification-positive sample data for the gene to be tested and the candidate amplification-positive judgment value, counting the number of true positives and the number of false negatives corresponding to the candidate amplification-positive judgment value; The positive coincidence rate corresponding to the candidate amplification positive judgment value is determined based on the number of true positives and the number of false negatives corresponding to the candidate amplification positive judgment value.

[0345] In some optional embodiments, determining the positive predictive value corresponding to the candidate amplification positive judgment value based on the comparison result of the amplification judgment value of the gene to be tested with the candidate amplification positive judgment value of each amplification sample data may include: According to the comparison result of the amplification judgment value of each amplification-positive sample data for the gene to be tested and the candidate amplification-positive judgment value, counting the number of true positives corresponding to the candidate amplification-positive judgment value; According to the comparison result of the amplification judgment value of each negative sample data for the gene to be tested and the candidate amplification positive judgment value, counting the number of false positives corresponding to the candidate amplification positive judgment value; and The positive predictive value corresponding to the candidate amplification positive judgment value is determined based on the number of true positives and the number of false positives corresponding to the candidate amplification positive judgment value.

[0346] In some optional embodiments, determining the amplification judgment value of the corresponding amplified sample data for the gene to be tested based on the corrected tumor cell copy number of each amplified sample data may include: Determine the corrected tumor cell copy number of each amplified sample data for the gene to be tested as the amplification judgment value of the amplified sample data, or, The difference between the corrected tumor cell copy number of each amplified sample data for the gene to be tested and the tumor ploidy of the amplified sample data is determined as the amplification judgment value of the amplified sample data, or The amplification judgment value of the amplified sample data is determined by dividing the ratio of the corrected tumor cell copy number of the gene to be tested for each amplified sample data by the tumor ploidy of the amplified sample data.

[0347] It should be noted that the implementation details and technical effects of each module in the gene copy number variation type detection device provided in the embodiments of the present disclosure can be referred to the description of other embodiments in the present disclosure and will not be repeated here.

[0348] Reference below Figure 9 , which shows a schematic structural diagram of a computer system 900 suitable for implementing the electronic device of the present disclosure. Figure 9 The computer system 900 shown is only an example and should not limit the functions and usage scope of the embodiments of the present disclosure.

[0349] like Figure 9 As shown, computer system 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. RAM 903 also stores various programs and data required for the operation of computer system 900. Processing device 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to bus 904.

[0350] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the computer system 900 to communicate with other devices wirelessly or by wire to exchange data. Figure 9 The computer system 900 of the electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0351] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0352] It should be noted that the computer-readable medium described above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.

[0353] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0354] The computer readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device can realize the following operation: Figure 2AThe illustrated embodiment and its optional implementations illustrate a method for detecting gene copy number variation types.

[0355] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Python, Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0356] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0357] The modules described in the embodiments of the present disclosure may be implemented in software or hardware. In some cases, the name of a module does not limit the module itself. For example, a sequencing alignment data acquisition module may also be described as a "module for acquiring sequencing alignment data for genetic sequencing of a test sample and at least one baseline sample."

[0358] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the scope of the above disclosure. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

Claims

1. A method for detecting gene copy number variation types, comprising: Acquire sequencing comparison data of gene sequencing performed on the test sample and at least one baseline sample, wherein each sequencing comparison data includes sequencing comparison data of a preset set of analysis regions of the test gene, and the preset set of analysis regions of the test gene includes the region where the test gene is located; Determining the tumor purity and tumor ploidy of the sample to be tested based on the sequencing comparison data; Determining the average mixed copy number of the gene to be tested in the sample to be tested based on each of the sequencing comparison data; Determining the corrected tumor cell copy number of the gene to be tested in the test sample based on the tumor purity and tumor ploidy of the test sample and the average mixed copy number of the gene to be tested in the test sample; In response to the average mixed copy number of the gene to be tested in the test sample being less than 2, generating a detection result for the homozygous deletion of the gene to be tested in the test sample based on a comparison result of the corrected tumor cell copy number of the gene to be tested in the test sample and a positive judgment threshold for homozygous deletion of the gene to be tested; and / or, In response to the average mixed copy number of the gene to be tested in the sample to be tested being greater than 2, an amplification judgment value of the gene to be tested in the sample to be tested is determined according to the corrected tumor cell copy number of the gene to be tested in the sample to be tested, and a detection result of the amplification of the copy number of the gene to be tested in the sample to be tested is generated according to a comparison result of the amplification judgment value of the gene to be tested in the sample to be tested and an amplification positive judgment threshold of the gene to be tested.

2. The method according to claim 1, wherein Determining the corrected tumor cell copy number of the gene to be tested in the test sample based on the tumor purity and tumor ploidy of the test sample and the average mixed copy number of the gene to be tested in the test sample comprises: The tumor purity and tumor ploidy of the sample to be tested, and the average mixed copy number of the gene to be tested in the sample to be tested are substituted into the corrected tumor cell copy number solution to obtain the corrected tumor cell copy number of the gene to be tested.

3. The method according to claim 2, wherein: The formula for calculating the corrected tumor cell copy number is obtained by deriving a correction formula, wherein the correction formula is: ; The formula for calculating the corrected tumor cell copy number is: ; in, and are respectively the tumor purity and tumor ploidy of the sample to be tested, is the average mixed copy number of the gene to be tested in the sample to be tested, is the corrected tumor cell copy number of the gene to be tested in the sample to be tested.

4. The method according to claim 1, wherein Determining the amplification judgment value of the gene to be tested in the sample to be tested according to the corrected tumor cell copy number of the gene to be tested in the sample to be tested includes: Determining the corrected tumor cell copy number of the gene to be tested in the sample to be tested as the amplification judgment value of the gene to be tested in the sample to be tested, or, The difference between the corrected tumor cell copy number of the gene to be tested in the sample to be tested and the tumor ploidy of the sample to be tested is determined as the amplification judgment value of the gene to be tested in the sample to be tested, or, The amplification judgment value of the gene to be tested in the sample to be tested is determined as the ratio of the corrected tumor cell copy number of the gene to be tested in the sample to be tested divided by the tumor ploidy of the sample to be tested.

5. The method according to claim 1, wherein Determining the tumor purity and tumor ploidy of the sample to be tested based on each of the sequencing comparison data includes: Determine the abundance of each SNP site in the preset whole-genome SNP site set of the sample to be tested based on the sequencing comparison data of the sample to be tested; For each predetermined whole-genome SNP site, the related gene analysis region is divided into at least one SNP analysis region according to a first predetermined length; For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each SNP analysis region; For each SNP analysis region, the mixed copy number of the sequencing alignment data of the test sample in the SNP analysis region is determined based on the corrected average coverage depth of the sequencing alignment data of the test sample in the SNP analysis region and the average of the corrected average coverage depth of the sequencing alignment data of each baseline sample in the SNP analysis region; Based on the abundance and mixed copy number of each SNP site in the sequencing comparison data of the sample to be tested, the tumor purity and tumor ploidy of the sample to be tested are determined, wherein the mixed copy number of each SNP site is the mixed copy number of the SNP analysis region where the SNP site is located.

6. The method according to claim 5, wherein: Determining the average mixed copy number of the gene to be tested in the sample to be tested based on each of the sequencing comparison data includes: For each preset gene analysis region to be tested, dividing the preset gene analysis region to be tested into at least one gene analysis window according to a second preset length; For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each gene analysis window; For each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, determine the reference baseline sample sequencing alignment data corresponding to the gene analysis window in the sequencing alignment data of each baseline sample, as well as the corrected average baseline sample sequencing depth and the corrected average baseline sample sequencing depth deviation of the gene analysis window; Determining a reference gene analysis window in each gene analysis window based on the corrected baseline sample average sequencing depth deviation of each gene analysis window; Determine each reference gene analysis window where the gene to be tested is located as the gene analysis window to be tested; For each gene analysis window to be tested, determining the mixed copy number of the gene analysis window to be tested in the sample to be tested based on the corrected average coverage depth of the sequencing alignment data of the sample to be tested in the gene analysis window to be tested and the corrected average sequencing depth of the baseline sample in the gene analysis window to be tested; The mean of the mixed copy numbers of each analysis window of the gene to be tested in the sample to be tested is determined as the average mixed copy number of the gene to be tested in the sample to be tested.

7. The method according to claim 6, wherein: For each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, determining the reference baseline sample sequencing alignment data corresponding to the gene analysis window, as well as the corrected baseline sample average sequencing depth and the corrected baseline sample average sequencing depth deviation of the gene analysis window in the sequencing alignment data of each baseline sample, including: For each gene analysis window, perform the following baseline sample gene analysis window depth correction operation: Determine the normal range of the average sequencing depth of the baseline samples based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window; Determine the sequencing alignment data of each baseline sample whose average coverage depth after correction in the gene analysis window falls within the normal range of the average sequencing depth of the baseline samples as the reference baseline sample sequencing alignment data corresponding to the gene analysis window; Based on the corrected average coverage depth of the sequencing alignment data of each reference baseline sample in the gene analysis window, the corrected average sequencing depth of the baseline samples and the corrected average sequencing depth deviation of the baseline samples in the gene analysis window are determined.

8. The method according to claim 6, wherein: For each sequencing alignment data, determining the corrected average coverage depth of the sequencing alignment data for each gene analysis window includes: For each sequencing alignment data, each gene analysis window is used as the target analysis region, and the following target analysis region sequencing depth correction operation is performed to obtain the corrected average coverage depth of each gene analysis window. The target analysis region sequencing depth correction operation includes: The sequencing depth of each base site in each target analysis region is normalized according to the total sequencing depth of the sequencing alignment data to obtain the corresponding normalized sequencing depth; For each target analysis region, determining the average sequencing depth of the sequencing alignment data in the target analysis region based on the normalized sequencing depth of each base position in the target analysis region; Performing a local weighted regression correction on the average coverage depth of each target analysis region of the sequencing alignment data according to the number of probes covered in each target analysis region; According to the GC content ratio of the sequencing alignment data in each target analysis region, a local weighted regression correction is performed on the average coverage depth of the sequencing alignment data in each target analysis region to obtain the corrected average coverage depth of the corresponding target analysis region.

9. The method according to claim 8, wherein For each sequencing alignment data, determining the corrected average coverage depth of the sequencing alignment data for each SNP analysis region comprises: For each sequencing alignment data, each SNP analysis region is used as a target analysis region, and a sequencing depth correction operation of the target analysis region is performed to obtain the corrected average coverage depth of each SNP analysis region.

10. The method according to claim 1, wherein Before determining the tumor purity and tumor ploidy of the sample to be tested based on each sequencing comparison data, the method further includes: For each sequencing alignment data, the following base site sequencing depth normalization operation is performed: for each preset gene analysis region to be tested, the sequencing depth of each base site of the sequencing alignment data in the preset gene analysis region to be tested is determined; the sum of the sequencing depths of all base sites of the sequencing alignment data in each of the preset gene analysis regions to be tested is determined as the total sequencing depth of the sequencing alignment data; for each base site in each preset gene analysis region to be tested, the sequencing depth of the base site is normalized based on the total sequencing depth of the sequencing alignment data.

11. The method according to claim 1, wherein The homozygous deletion positive judgment threshold of the gene to be tested is predetermined by the following homozygous deletion positive judgment threshold determination step: Obtaining a homozygous deletion sample data set, wherein the homozygous deletion samples include negative sample data and homozygous deletion positive sample data for the gene to be tested, wherein each negative sample data is obtained based on sequencing comparison data of a benign sample, and each homozygous deletion positive sample data is sequencing comparison data of a homozygous deletion occurring in the gene to be tested; Determining the corrected tumor cell copy number of the gene to be tested for each homozygous deletion sample data; For each candidate homozygous deletion copy number in the candidate homozygous deletion copy number set, performing the following homozygous deletion positive judgment basis determination operation: determining a positive coincidence rate corresponding to the candidate homozygous deletion copy number based on a comparison result of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each of the homozygous deletion positive sample data; and / or determining a positive predictive value corresponding to the candidate homozygous deletion copy number based on a comparison result of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each of the homozygous deletion sample data; The positive judgment threshold for homozygous deletion of the gene to be tested is determined based on the positive coincidence rate and / or positive predictive value corresponding to each candidate homozygous deletion copy number.

12. The method according to claim 11, wherein The obtaining of a homozygous deletion sample data set includes: Obtaining sequencing alignment data of N pairs of 100% tumor cell lines and their paired cell lines, wherein N is a positive integer; Based on the N pairs of sequencing comparison data, generating at least one homozygous deletion positive sample data for the gene to be tested, wherein the tumor purity of each homozygous deletion positive sample data is the corresponding annotated tumor purity; Sequencing comparison data of M benign samples are obtained, and at least one negative sample data is generated based on the sequencing comparison data of the M benign samples.

13. The method according to claim 11, wherein Determining the corrected tumor cell copy number of the gene to be tested for each homozygous deletion sample data includes: For each homozygous deletion sample data, determine the tumor purity and tumor ploidy corresponding to the homozygous deletion sample data and the average mixed copy number for the gene to be tested; based on the tumor purity and tumor ploidy corresponding to the homozygous deletion sample data and the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the homozygous deletion sample data for the gene to be tested.

14. The method according to claim 11, wherein The step of comparing the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number based on the homozygous deletion positive sample data to determine the positive coincidence rate corresponding to the candidate homozygous deletion copy number comprises: According to the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number of each homozygous deletion positive sample data, the number of true positives and the number of false negatives corresponding to the candidate homozygous deletion copy number are counted; The positive coincidence rate corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and the number of false negatives corresponding to the candidate homozygous deletion copy number.

15. The method according to claim 11, wherein The step of comparing the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number based on the homozygous deletion sample data to determine the positive predictive value corresponding to the candidate homozygous deletion copy number comprises: According to the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number of each homozygous deletion positive sample data, the number of true positives corresponding to the candidate homozygous deletion copy number is counted; According to the comparison result of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number in each of the negative sample data, counting the number of false positives corresponding to the candidate homozygous deletion copy number; and A positive predictive value corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and the number of false positives corresponding to the candidate homozygous deletion copy number.

16. The method according to claim 12, wherein: The step of generating at least one negative sample data based on the sequencing comparison data of the M benign samples includes: For each sequencing alignment data of a benign sample, different degrees of noise are added to the sequencing depth of the base sites in the sequencing alignment data of the benign sample to obtain at least one negative sample data corresponding to the sequencing alignment data of the benign sample.

17. The method according to claim 1, wherein The amplification positive judgment threshold of the gene to be tested is predetermined by the following amplification positive judgment threshold determination step: Obtaining a set of amplified sample data, wherein the amplified sample data includes negative sample data and amplified positive sample data for the gene to be tested, wherein each negative sample data is obtained based on sequencing comparison data of a benign sample, and each amplified positive sample data is sequencing comparison data for copy number amplification of the gene to be tested; Determine the corrected tumor cell copy number for each amplified sample data; Determining the amplification judgment value of the corresponding amplified sample data for the gene to be tested according to the corrected tumor cell copy number of each amplified sample data; For each candidate amplification positive judgment value in the candidate amplification positive judgment value set, performing the following amplification positive judgment basis determination operation: determining a positive coincidence rate corresponding to the candidate amplification positive judgment value based on a comparison result of the amplification judgment value of each amplification positive sample data for the gene to be tested and the candidate amplification positive judgment value; and / or determining a positive predictive value corresponding to the candidate amplification positive judgment value based on a comparison result of the amplification judgment value of each amplification sample data for the gene to be tested and the candidate amplification positive judgment value; The amplification positive judgment threshold of the gene to be tested is determined according to the positive coincidence rate and / or positive predictive value corresponding to each candidate amplification positive judgment value.

18. The method according to claim 17, wherein The obtaining of the amplified sample data set includes: Obtaining sequencing alignment data of I pairs of 100% tumor cell lines and their paired cell lines, where I is a positive integer; Based on the I pair of sequencing comparison data, generating at least one amplification-positive sample data for the gene to be tested, wherein the tumor purity of each amplification-positive sample data is the corresponding annotated tumor purity; Sequencing comparison data of J benign samples are obtained, and at least one negative sample data is generated based on the sequencing comparison data of the J benign samples.

19. The method according to claim 17, wherein Determining the corrected tumor cell copy number of each amplified sample data comprises: For each amplified sample data, determine the tumor purity and tumor ploidy corresponding to the amplified sample data and the average mixed copy number for the gene to be tested; based on the tumor purity and tumor ploidy corresponding to the amplified sample data and the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the gene to be tested in the amplified sample data.

20. The method according to claim 17, wherein The step of comparing the amplification judgment value of the gene to be tested with the candidate amplification positive judgment value based on the amplification positive sample data to determine the positive coincidence rate corresponding to the candidate amplification positive judgment value includes: According to the comparison result of the amplification judgment value of each amplification-positive sample data for the gene to be tested and the candidate amplification-positive judgment value, counting the number of true positives and the number of false negatives corresponding to the candidate amplification-positive judgment value; The positive coincidence rate corresponding to the candidate amplification positive judgment value is determined based on the number of true positives and the number of false negatives corresponding to the candidate amplification positive judgment value.

21. The method according to claim 17, wherein The step of determining a positive prediction value corresponding to the candidate amplification positive judgment value based on a comparison result of the amplification judgment value of the gene to be tested with the candidate amplification positive judgment value of each amplification sample data comprises: According to the comparison result of the amplification judgment value of each amplification-positive sample data for the gene to be tested and the candidate amplification-positive judgment value, counting the number of true positives corresponding to the candidate amplification-positive judgment value; According to the comparison result of the amplification judgment value of each negative sample data for the gene to be tested and the candidate amplification positive judgment value, counting the number of false positives corresponding to the candidate amplification positive judgment value; and The positive predictive value corresponding to the candidate amplification positive judgment value is determined based on the number of true positives and the number of false positives corresponding to the candidate amplification positive judgment value.

22. The method according to claim 17, wherein The step of determining the amplification judgment value of the corresponding amplified sample data for the gene to be tested based on the corrected tumor cell copy number of each amplified sample data includes: Determine the corrected tumor cell copy number of each amplified sample data for the gene to be tested as the amplification judgment value of the amplified sample data, or, The difference between the corrected tumor cell copy number of each amplified sample data for the gene to be tested and the tumor ploidy of the amplified sample data is determined as the amplification judgment value of the amplified sample data, or The amplification judgment value of the amplified sample data is determined by dividing the ratio of the corrected tumor cell copy number of the gene to be tested for each amplified sample data by the tumor ploidy of the amplified sample data.

23. A device for detecting gene copy number variation types, comprising: A sequencing comparison data acquisition module is configured to acquire sequencing comparison data of gene sequencing performed on the test sample and at least one baseline sample, wherein each sequencing comparison data includes sequencing comparison data of a preset set of analysis regions of the test gene, and the preset set of analysis regions of the test gene includes the region where the test gene is located; a purity and ploidy determination module, configured to determine the tumor purity and tumor ploidy of the sample to be tested based on each of the sequencing comparison data; an average mixed copy number determination module, configured to determine the average mixed copy number of the gene to be tested in the sample to be tested based on each of the sequencing comparison data; a tumor cell copy number correction module, configured to determine the corrected tumor cell copy number of the gene to be tested in the sample to be tested based on the tumor purity and tumor ploidy of the sample to be tested and the average mixed copy number of the gene to be tested in the sample to be tested; a homozygous deletion detection module configured to, in response to the average mixed copy number of the gene to be tested in the sample to be tested being less than 2, generate a detection result for the homozygous deletion of the gene to be tested in the sample to be tested based on a comparison result of the corrected tumor cell copy number of the gene to be tested in the sample to be tested and a homozygous deletion positive judgment threshold of the gene to be tested; and / or, The copy number amplification detection module is configured to, in response to an average mixed copy number of the gene to be tested in the sample to be tested being greater than 2, determine an amplification judgment value of the gene to be tested in the sample to be tested based on the corrected tumor cell copy number of the gene to be tested in the sample to be tested, and generate a detection result of the copy number amplification of the gene to be tested in the sample to be tested based on a comparison result of the amplification judgment value of the gene to be tested in the sample to be tested and an amplification positive judgment threshold of the gene to be tested.

24. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 22.

25. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by one or more processors, the method according to any one of claims 1 to 22 is implemented.

26. A computer program product comprising a computer program / instructions, wherein when the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 22 is implemented.

Citation Information

Patent Citations

  • Single-sample whole-genome allele specific copy number variation prediction method

    CN112802548A

  • Single sample allele copy number variation detection method, probe set and kit

    CN113889187A

  • Method and device for deducing tumor purity and ploidy

    CN113990389A

  • Method and device for detecting copy number variation and computer readable medium

    CN117524301A

  • Copy number variation detection method and application thereof

    WO2023030233A1