Method and device for detecting embryo system variation by single sample
By genome sequencing and comparison of tumor samples, combined with copy number interval division and cluster analysis, the problem of germline variant recognition in the absence of control was solved, and tumor gene detection was achieved with high accuracy.
Patent Information
- Application Number
- CN202411945308.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-26
AI Technical Summary
In tumor gene detection, many tumor cell samples cannot obtain control tissue, resulting in the inability to identify the source of mutations from tumor somatic mutations, and a method is needed to identify germline mutations.
By comparing the sequencing data of the sample to be tested with the reference genome, the variation site information is obtained, and the genome is divided into multiple copy number intervals, clustering the variation site, performing a differential test, and finally inputting the interpretation model to distinguish germline variants and somatic variants.
It is possible to accurately distinguish germline variants and somatic variants in tumor samples without control, and improve the accuracy and sensitivity of tumor gene detection.
Smart Images

Figure BDA0005213965210000031 
Figure BDA0005213965210000032 
Figure BDA0005213965210000033
Abstract
Description
Technical Field
[0001] The present invention relates to a method and a device for detecting germline variation from a single tumor sample. Background Art
[0002] The mutation sites in the human genome can be divided into germline mutation and somatic mutation. Germline mutation refers to the mutation inherited by an organism from its parents, which is not necessarily harmful. Somatic mutation refers to the mutation that occurs in the normal cells of an organism. This mutation is acquired during the development of an individual and may be caused by changes in the genome sequence due to internal factors such as the accumulation of reactive oxygen species or external factors such as radiation and toxins. Somatic mutation may lead to the occurrence of cancer.
[0003] Using high-throughput sequencing methods, the genome sequence of sample cells can be obtained, and then mutations can be detected from the genome sequence using bioinformatics algorithms. In the field of tumor gene testing, detecting targeted mutations in tumor cells can provide medication guidance for tumor treatment. In order to ensure that the source of the mutation is tumor cells, a common approach is to use sample-paired normal tissue cells as controls to identify tumor genome-specific mutations that are not present in control cells during mutation detection. However, in clinical applications, control tissues cannot be obtained for many tumor cell samples, which can result in the inability to identify the source of mutations from tumor somatic cell mutations during mutation detection. Therefore, a method that can identify germline variations from tumor somatic cell genome sequencing results is necessary. Summary of the invention
[0004] The present disclosure is directed to methods and devices for distinguishing germline variation from somatic variation in a single tumor sample.
[0005] According to one aspect of the present disclosure, a method for distinguishing the types of gene mutations in a sample to be tested is provided, the method comprising the following steps:
[0006] Compare the sequencing data of the sample to be tested with the reference genome to obtain comparison data;
[0007] Based on the comparison data, one or more variant site information is obtained, and the variant frequency AF and population frequency Freq of each variant site are obtained. DB ;
[0008] Based on the comparison data, dividing the genome of the sample to be tested into multiple copy number intervals;
[0009] Clustering one or more variant sites in each copy number interval to obtain homozygous variant site classes and heterozygous variant site classes;
[0010] Performing a difference test on one or more variant sites of the homozygous variant site class and the heterozygous variant site class with the theoretical germline homozygous allele frequency and the theoretical germline heterozygous allele frequency, respectively, to obtain a test germP value for each variant site; and
[0011] The variation frequency AF and the population frequency Freq of the one or more variation sites are DB The test germP value is input into the interpretation model to obtain the type of the one or more variable sites.
[0012] In some embodiments, the one or more variant site information includes, but is not limited to, at least one of single nucleotide variation (SNV), insertion / deletion variation (InDel), copy number variation (CNV), and chromosome structural variation (SV).
[0013] In some embodiments, the one or more variant site information includes at least one of a single nucleotide variation (SNV) and an insertion / deletion variation (InDel).
[0014] In some embodiments, dividing the genome of the sample to be tested into multiple copy number intervals includes: using a cyclic binary segmentation algorithm (CBS) to divide the genome of the sample to be tested into multiple copy number intervals.
[0015] In some embodiments, the cyclic binary segmentation algorithm (CBS) divides the genome of the sample to be tested into multiple copy number intervals based on the sequencing depth obtained from the alignment data.
[0016] The sequencing depth can be obtained using software (e.g., CNVkit software) or means known to those skilled in the art. In an optional embodiment, the acquisition of the sequencing depth may include the following steps: constructing a normal person sequencing baseline, using the normal person sequencing baseline as a control, and obtaining the sequencing depth of the comparison data. In a preferred embodiment, the sequencing depth is corrected using the genome GC content.
[0017] In some embodiments, the cyclic binary segmentation algorithm (CBS) divides regions with similar sequencing depths into one copy number interval.
[0018] In some embodiments, the region with similar depth signals refers to a region where there is no significant difference in depth signals between adjacent sites based on the cyclic binary segmentation algorithm CBS.
[0019] In some embodiments, the population frequency Freq DB The method can be obtained by annotating the one or more variant sites using a population frequency database.
[0020] Those skilled in the art will appreciate that the population frequency database may include those databases conventionally used in the art. In some embodiments, the population frequency database may include, but is not limited to, gnomAD, dbSNP, 1000G, 1000G Asian, ExAC, ExAC Asian, Chinese population frequency database PAD, etc.
[0021] In some embodiments, the sequencing data of the sample to be tested can be obtained by, for example but not limited to, next-generation sequencing (NGS) technology.
[0022] In some embodiments, the step of clustering one or more variant sites within each copy number interval may include: obtaining a minor allele frequency MAF based on the variant frequency AF, and then clustering one or more variant sites within each copy number interval based on the minor allele frequency MAF.
[0023] In some embodiments, the minor allele frequency MAF is obtained according to the following formula 1:
[0024] MAF=min(AF,1-AF) Formula 1.
[0025] In some embodiments, the clustering may include: using a K-means clustering algorithm to divide one or more variant sites within each copy number interval into two categories according to the minor allele frequency MAF, comparing the average MAF values of the two categories of variant sites obtained by clustering, and judging the class with a larger average MAF as a heterozygous variant site class, and judging the class with a smaller average MAF as a homozygous variant site class.
[0026] In some embodiments, for the heterozygous variant site class, the variant site with 0.1≤AF≤0.5 and 0.1≤MAF≤0.5 is determined to be a heterozygous variant site.
[0027] In some embodiments, in the homozygous variant site class, a variant site with AF>0.5 is determined as a homozygous variant site.
[0028] In some embodiments, when the number of variant sites within a copy number interval is insufficient for clustering, variant sites with AF>0.8 are judged as homozygous sites, and variant sites with 0.2≤AF≤0.8 are judged as heterozygous sites.
[0029] In some embodiments, the theoretical germline heterozygous allele frequency can be obtained according to the following formula 2:
[0030]
[0031] Among them, W iRefers to the weight of the i-th variant site, MAF i Refers to the minor allele frequency of the ith variant site.
[0032] In some embodiments, the theoretical germline homozygous allele frequency can be obtained according to the following formula 3:
[0033]
[0034] Among them, W i Refers to the weight of the i-th variant site, AF i Refers to the mutation frequency of the ith mutation site.
[0035] In some embodiments, the weight of the i-th variant site can be obtained according to the following formula 4:
[0036]
[0037] Among them, Dep i refers to the sequencing depth of the i-th variant site, ΔMAF i It refers to the difference between the MAF of the ith variant site and the upper quartile of the variant frequency AF of one or more variant sites in the copy number interval to which the ith variant site belongs, and α is a training constant term.
[0038] In some embodiments, the training constant term α can be selected from, but not limited to, 0.01, 0.02, 0.03, 0.04, 0.05, etc.
[0039] In some embodiments, the difference test may include, but is not limited to, t-test, analysis of variance, chi-square test, etc. In some embodiments, the difference test may be a binomial difference test. In some embodiments, the difference test may be a two-tailed binomial difference test.
[0040] In some embodiments, the test germP value of a variant site in the one or more variant sites can be obtained by the following steps:
[0041] The variation frequency AF of the variation site is compared with the theoretical germline homozygous allele frequency THEOAF homo Perform a difference test and obtain the test germP homo value;
[0042] The minor allele frequency MAF of the variant site is compared with the theoretical germline heterozygous allele frequency THEOAF het Perform a difference test and obtain the test germP het value; and
[0043] The test germP value is obtained according to the following formula 5:
[0044] germP=max(germP homo , germP het ) Formula 5.
[0045] In some embodiments, the method may further include determining the gender of the individual from whom the sample to be tested originates.
[0046] In some embodiments, when the sample to be tested is from a male individual and the one variant site is located in a non-homologous region of a sex chromosome, germP=germP homo .
[0047] In some embodiments, the interpretation model can be constructed using a machine learning model.
[0048] In some embodiments, the algorithm of the machine learning model includes, but is not limited to at least one of K nearest neighbor, naive Bayes classifier, logistic regression, decision tree, random forest, support vector machine, neural network, and AdaBoost.
[0049] In some embodiments, the interpretation model is constructed using a decision tree model.
[0050] In some embodiments, the decision tree model can be constructed by inputting the variation frequency AF of one or more variation sites, the population frequency Freq DB In some embodiments, in the decision tree model, the weights of the features are ranked as follows: germP>Freq DB >AF.
[0051] In a specific embodiment, the interpretation model is constructed using the sklearn decision tree algorithm. In some embodiments, the sklearn decision tree algorithm can use the variation frequency AF of one or more variant sites, the population frequency Freq DB The germP value was tested and the model with the highest selectivity score was obtained through cross-validation method.
[0052] In some embodiments, the interpretation model can be constructed using DecisionTreeClassifier. In some embodiments, the cross-validation method can be used, such as a 3-fold cross-validation method. For example, 2 / 3 of the variant sites can be used for training, and the remaining 1 / 3 of the variant sites can be used for verification. In some embodiments, the GridSearchCV function can be used to perform a grid search on the parameters in DecisionTreeClassifier (e.g., max_depth, min_samples_leaf, and min_samples_split). This process is intended to prevent overfitting and select the model with the highest performance score.
[0053] In some embodiments, the method can determine whether the one or more variant sites are germline variants or somatic variants.
[0054] In some embodiments, the variant site may include any site in the sample to be tested that is inconsistent with the reference genome. For example, the variant site may include a site with one or more base insertions, deletions, or substitutions compared to the reference genome.
[0055] In some embodiments, the sample is derived from a diseased individual. It will be appreciated by those skilled in the art that the sample can be derived from an individual suffering from any disease. In an exemplary embodiment, the sample is a tumor sample or other disease sample.
[0056] In some embodiments, the other diseases include but are not limited to adverse reactions after organ transplantation; infection by pathogens such as bacteria, viruses, fungi, parasites, etc.; hepatitis; chronic gastroenteritis; diabetes; chronic pharyngitis; cardiovascular and cerebrovascular diseases; lung-related diseases, etc.
[0057] In some embodiments, the sample to be tested includes but is not limited to a body fluid sample or a tissue sample. In some embodiments, the body fluid sample includes but is not limited to plasma, urine, saliva, bronchoalveolar lavage fluid or cerebrospinal fluid from blood.
[0058] The method disclosed herein can distinguish somatic mutations and germline mutations in a sample based only on a tumor sample or other disease sample, without the need for comparative analysis with a control sample (eg, a blood cell sample, a reference genome, etc.).
[0059] In some embodiments, the reference genome includes, but is not limited to, GRCH37, b37, hs37d5 (b37+decoy), hg19, GRCH38 (hg38), etc. In some embodiments, the hg19 genome can be downloaded from UCSC (http: / / genome.ucsc.edu / ). In some embodiments, the GRCH38 genome can be downloaded from NCBI (https: / / www.ncbi.nlm.nih.gov / ). In some embodiments, the sample is taken from the human body, and the sequencing data is compared with the human reference genome. In some embodiments, the sample is taken from an animal, and the sequencing data is compared with the reference genome of the corresponding animal species.
[0060] According to another aspect of the present disclosure, a device for distinguishing the type of gene variation in a sample to be tested is provided, the device comprising:
[0061] A comparison data acquisition module is used to compare the sequencing data of the sample to be tested with the reference genome to obtain comparison data;
[0062] The variant site information analysis module is used to obtain one or more variant site information based on the comparison data, the variant frequency AF and the population frequency Freq of each variant site DB ;
[0063] A copy number interval division module, used for dividing the genome of the sample to be tested into multiple copy number intervals based on the comparison data;
[0064] A clustering module, used to cluster one or more variant sites in each copy number interval to obtain homozygous variant site classes and heterozygous variant site classes;
[0065] a difference test module, used to perform a difference test on one or more variant sites of the homozygous variant site class and the heterozygous variant site class respectively with the theoretical germline homozygous allele frequency and the theoretical germline heterozygous allele frequency to obtain a test germP value for each variant site; and
[0066] The variant site type judgment module is used to convert the variant frequency AF and the population frequency Freq of the one or more variant sites into DB The test germP value is input into the interpretation model to obtain the type of the one or more variable sites.
[0067] In some embodiments, the device may further include a sequencing data acquisition module for acquiring sequencing data of the sample to be tested.
[0068] In some embodiments, the device is used to perform the above method of the present disclosure.
[0069] In some embodiments, the copy number interval division module uses a cyclic binary segmentation algorithm (CBS) to divide the genome of the sample to be tested into multiple copy number intervals.
[0070] In some embodiments, the population frequency Freq DB The method can be obtained by annotating the one or more variant sites using a population frequency database.
[0071] In some embodiments, the clustering module can obtain a minor allele frequency MAF based on the variant frequency AF, and then cluster one or more variant sites in each copy number interval based on the minor allele frequency MAF.
[0072] In some embodiments, the minor allele frequency MAF is obtained according to the following formula 1:
[0073] MAF=min(AF,1-AF) Formula 1.
[0074] In some embodiments, the clustering may include: using a K-means clustering algorithm to divide one or more variant sites within each copy number interval into two categories according to the minor allele frequency MAF, comparing the average MAF values of the two categories of variant sites obtained by clustering, and judging the class with a larger average MAF as a heterozygous variant site class, and judging the class with a smaller average MAF as a homozygous variant site class.
[0075] In some embodiments, for the heterozygous variant site class, the variant site with 0.1≤AF≤0.5 and 0.1≤MAF≤0.5 is determined to be a heterozygous variant site.
[0076] In some embodiments, in the homozygous variant site class, a variant site with AF>0.5 is determined as a homozygous variant site.
[0077] In some embodiments, when the number of variant sites within a copy number interval is insufficient for clustering, variant sites with AF>0.8 are judged as homozygous sites, and variant sites with 0.2≤AF≤0.8 are judged as heterozygous sites.
[0078] In some embodiments, the device may further include a theoretical germline allele frequency acquisition module for acquiring a theoretical germline homozygous allele frequency and a theoretical germline heterozygous allele frequency.
[0079] In some embodiments, the theoretical germline heterozygous allele frequency can be obtained according to the following formula 2:
[0080]
[0081] Among them, W i Refers to the weight of the i-th variant site, MAF i Refers to the minor allele frequency of the ith variant site.
[0082] In some embodiments, the theoretical germline homozygous allele frequency can be obtained according to the following formula 3:
[0083]
[0084] Among them, W i Refers to the weight of the i-th variant site, AF i Refers to the mutation frequency of the ith mutation site.
[0085] In some embodiments, the weight of the i-th variant site can be obtained according to the following formula 4:
[0086]
[0087] Among them, Dep i refers to the sequencing depth of the i-th variant site, ΔMAF i It refers to the difference between the MAF of the ith variant site and the upper quartile of the variant frequency AF of one or more variant sites in the copy number interval to which the ith variant site belongs, and α is a training constant term.
[0088] In some embodiments, the training constant term α can be selected from, but not limited to, 0.01, 0.02, 0.03, 0.04, 0.05, etc.
[0089] In some embodiments, the difference test module can use, but is not limited to, t-test, analysis of variance, chi-square test, etc. to perform a difference test. In some embodiments, the difference test can be a binomial difference test. In some embodiments, the difference test can be a two-tailed binomial difference test.
[0090] In some embodiments, the test germP value of a variant site in the one or more variant sites can be obtained by the following steps:
[0091] The variation frequency AF of the variation site is compared with the theoretical germline homozygous allele frequency THEOAF homo Perform a difference test and obtain the test germP homo value;
[0092] The minor allele frequency MAF of the variant site is compared with the theoretical germline heterozygous allele frequency THEOAF het Perform a difference test and obtain the test germP hetvalue; and
[0093] The test germP value is obtained according to the following formula 5:
[0094] germP=max(germP homo , germP het ) Formula 5.
[0095] In some embodiments, the device may further include a gender determination module for determining the gender of the individual from which the sample to be tested originates.
[0096] In some embodiments, when the gender determination module determines that the sample to be tested is from a male individual based on the sex chromosome of the sample to be tested, and the one variant site is located in a non-homologous region of the sex chromosome, germP=germP homo .
[0097] In some embodiments, the device may further include a judgment model building module.
[0098] In some embodiments, the judgment model construction module can be constructed by machine learning. In a specific embodiment, the judgment model uses the sklearn decision tree algorithm to construct the judgment model. In some embodiments, the sklearn decision tree algorithm can use the variation frequency AF of one or more variant sites, the population frequency Freq DB The germP value was tested and the model with the highest selectivity score was obtained through cross-validation method.
[0099] In some embodiments, the interpretation model construction module can use DecisionTreeClassifier to construct the interpretation model. In some embodiments, the cross-validation method can be used, for example, a 3-fold cross-validation method. For example, 2 / 3 of the variant sites can be used for training, and the remaining 1 / 3 of the variant sites can be used for verification. In some embodiments, the GridSearchCV function can be used to perform a grid search on the parameters in DecisionTreeClassifier (e.g., max_depth, min_samples_leaf, and min_samples_split). This process is intended to prevent overfitting and select the model with the highest performance score.
[0100] In some embodiments, the method can determine whether the one or more variant sites are germline variants or somatic variants.
[0101] In some embodiments, the variant site may include any site in the sample to be tested that is inconsistent with the reference genome. For example, the variant site may include a site with one or more base insertions, deletions, or substitutions compared to the reference genome.
[0102] According to another aspect of the present disclosure, a storage medium is provided, on which a computer-executable program is stored, and the program is configured to execute the method of distinguishing the types of gene variations in a sample to be tested of the present disclosure when running.
[0103] According to another aspect of the present disclosure, there is provided a device, comprising:
[0104] Memory, used to store programs;
[0105] A processor is used to execute the program stored in the memory, wherein the program is configured to execute the method for distinguishing the types of gene variations in the sample to be tested disclosed herein when running. BRIEF DESCRIPTION OF THE DRAWINGS
[0106] Figure 1 The following is a flowchart for distinguishing gene variation types according to one embodiment of the present disclosure. DETAILED DESCRIPTION
[0107] The present invention is directed to a method for accurately distinguishing germline variation and somatic variation from a single sample (eg, a tumor sample).
[0108] In order to achieve the above objectives, the present disclosure provides a method and device for detecting germline variation in a single tumor sample.
[0109] Figure 1 The method for detecting germline variation in a single tumor sample disclosed in the present invention is exemplified.
[0110] In a specific embodiment, the method of detecting embryonic system variation using a single test sample disclosed herein comprises the following steps:
[0111] S1. Gene sequencing: Use the NGS platform to perform gene sequencing of a single sample;
[0112] S2. Sample data alignment: Align the sequencing data to the human reference genome;
[0113] S3. Single-sample interval copy number variation detection: Use alignment data to divide different copy number intervals based on different sequencing depths on the genome;
[0114] S4. Single sample variant annotation: Use variant detection software to detect variant signals from the alignment data, obtain variant frequency AF, and use the population frequency database to annotate the population frequency of the variant site;
[0115] S5. Define heterozygous variant sites and homozygous variant sites within the interval: According to the variant frequency AF, the variant sites in the same copy number interval are divided into heterozygous variant site class and homozygous variant site class;
[0116] S6. Define the interval theoretical germline allele frequency: calculate the theoretical germline heterozygous site allele frequency and the theoretical germline homozygous site allele frequency within the copy number interval determined in step S5 respectively;
[0117] S7. Site frequency difference test: Use the binomial test method to test the difference between each site and the theoretical allele frequency determined in step S6 to obtain the test germP value;
[0118] S8. Construct a model for determining the source of variation: Use the variation frequency AF and population frequency Freq obtained in step S4 DB , the test germP value obtained in step S7 is used to construct a decision tree model for determining the source of variation;
[0119] S9. Determination of the source of variation: Using the interpretation model constructed in step S8, determine whether the source of a single variation is germline or systemic.
[0120] In some embodiments, the sample to be tested is derived from a diseased individual. It will be appreciated by those skilled in the art that the sample to be tested can be derived from an individual suffering from any disease. In an exemplary embodiment, the sample to be tested is a tumor sample or other disease sample.
[0121] In some embodiments, the gene sequencing can be targeted sequencing.
[0122] In some embodiments, the sample data comparison step described in step S2 includes: using bwa-mem2 software to compare the NGS second-generation sequencing sequence data to the hg19 reference genome.
[0123] In some embodiments, the step of detecting interval copy number variation in step S3 comprises the following steps:
[0124] S3a. Statistical genome sequence depth signal from the alignment data obtained in step S2 using copy number variation detection software;
[0125] S3b. Correct the signal of step S3a using the sequence depth signal baseline constructed using the normal control sample;
[0126] S3b. The cyclic binary segmentation algorithm CBS is used to divide the regions with similar depth signals on the genome into a copy number interval.
[0127] In some embodiments, the region with similar depth signals refers to a region where there is no significant difference in depth signals between adjacent sites based on the cyclic binary segmentation algorithm CBS.
[0128] In some embodiments, the variant site annotation step described in step S4 comprises the following steps:
[0129] S4a. Using variation detection software such as the comparison data obtained from step S2 to detect sample variation, obtain variation frequency AF;
[0130] S4b. Use the population database to annotate the variant and obtain the population database frequency Freq DB .
[0131] In some embodiments, the variation detection software may include, for example, RealDcaller2 and Mutect2.
[0132] In some embodiments, the step of defining variant heterozygous sites and homozygous sites within a copy number interval in step S5 comprises the following steps:
[0133] S5a. Calculate the minor allele frequency (MAF) of the variant site according to formula 1
[0134] MAF = min(AF, 1-AF) Formula 1
[0135] S5b. Use machine learning methods to cluster all sites in the copy number interval described in step S3 according to the minor allele frequency of the site, use the K-means clustering algorithm, divide the variant sites into two categories according to the minor allele frequency MAF value, compare the average MAF values of the two types of variants obtained by clustering, the class with a large average MAF is the heterozygous class, and the class with a small average MAF is the homozygous class; in the heterozygous variant site class, the variant sites that meet 0.1≤AF≤0.5 are heterozygous variant sites; in the homozygous variant site class, the sites that meet AF>0.5 are homozygous sites. When there are too few sites to cluster, the variant sites that meet AF>0.8 are defined as homozygous sites, and the variant sites that meet 0.2≤AF≤0.8 are defined as heterozygous sites;
[0136] In the formula 1 described in step S5a, AF refers to the variation frequency obtained in S4a.
[0137] In some embodiments, the step of defining interval theoretical germline allele frequencies in step S6 comprises the following steps:
[0138] S6a. Using formula 2, calculate the theoretical germline heterozygous allele frequency (THEOAF) of the interval according to the heterozygous sites in the interval determined in step S5. het ), where W in Formula 2 iFrom formula 4
[0139]
[0140] Among them, W i Refers to the weight of the variant site calculated according to Formula 4, MAF i refers to the minor allele frequency of the variant site obtained according to Formula 1, and i refers to the serial number of the heterozygous site in the interval;
[0141] S6b. According to formula 3, the theoretical germline homozygous allele frequency (THEOAF) of the interval is calculated based on the homozygous sites in the interval determined in step S6. homo ), where W in Formula 3 i From formula 4
[0142]
[0143] Among them, W i Refers to the weight of the variant site calculated according to Formula 4, AF i refers to the variation frequency obtained in step S4a, i refers to the serial number of the homozygous site;
[0144] S6c. Formula 4 in steps S6a to S6b is:
[0145]
[0146] Among them, Dep i Refers to the depth of the variant site, ΔMAF i It refers to the difference between the MAF of the variant site calculated by formula 1 and the upper quartile of the frequency of the variant site in the interval determined in step S3b. α is a constant term for training;
[0147] In some embodiments, the site frequency difference test step described in step S7 includes the following steps:
[0148] S7a. For each variant site detected in step S4a, calculate the theoretical germline homozygous allele frequency (THEOAF) of the interval determined in step S6a. homo ) Perform a two-tailed binomial test and record the test p value as germP homo ;
[0149] S7b. For each variant site detected in step S4a, calculate the minimum allele frequency (MAF) of the variant site according to formula 1, and calculate the theoretical germline heterozygous allele frequency (THEOAF) of the interval determined in step S6b. het ) Perform a two-tailed binomial test and record the test p value as germP het ;
[0150] S7c. For all autosomes, confirm the values used for calculations according to Equation 5.
[0151] germP=max(germP homo , germP het ) Formula 5;
[0152] S7d. For sex chromosomes, when the sex is female, the variation on the X chromosome is determined according to Formula 5 to determine the value to be used for calculation; when the sex is male, the variation of the homologous segments on the X chromosome and the Y chromosome is determined according to Formula 5 to determine the value to be used for calculation; the non-homologous segments on the X chromosome and the Y chromosome germP = germP homo .
[0153] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with embodiments. The specific embodiments described herein are only used to explain the present invention and are not intended to constitute any limitation of the present invention. In addition, in the following description, the description of known structures and technologies is omitted to avoid unnecessary confusion of the concepts of the present disclosure. Such structures and technologies are also described in many publications.
[0154] definition
[0155] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly used in the field to which the present invention belongs. For the purpose of interpreting this specification, the following definitions will apply, and where appropriate, terms used in the singular will also include the plural form, and vice versa.
[0156] Unless the context clearly dictates otherwise, the expressions "a", "an" and "an" as used herein include plural references. For example, reference to "a cell" includes a plurality of such cells and equivalents thereof known to those skilled in the art, and so forth.
[0157] As used herein, the term "about" refers to a range of ±20% of the value that follows. In some embodiments, the term "about" refers to a range of ±10% of the value that follows. In some embodiments, the term "about" refers to a range of ±5% of the value that follows.
[0158] The term "allele" as used herein corresponds to one or more nucleotides (which can occur as a substitution or insertion) or the absence of one or more nucleotides. "Locus" corresponds to a position in a genome. For example, a locus can be a single base or a continuous series of bases. The term "genomic position" can refer to a specific nucleotide position in a genome or a continuous block of nucleotide positions. "Heterozygous loci" (also referred to as "het") are positions in a specific genome of a reference genome or an organism being located, where a copy of a chromosome does not have the same allele (e.g., a single nucleotide or a set of nucleotides). When a locus is a nucleotide with different alleles, "het" can be a single nucleotide polymorphism (SNP). "Het" can also be a position in which there is an insertion or deletion (collectively referred to as "insertion / deletion (indel)") of one or more nucleotides or one or more tandem repeats. A single nucleotide variation (SNV) corresponds to a genomic position with a nucleotide different from the reference genome of a specific person (person). If there is only one nucleotide at the position, then the SNV can be homozygous for a certain person, and if there are two alleles at the position, then it is heterozygous. A heterozygous SNV is a het. SNP and SNV are used interchangeably herein.
[0159] The terms "mutation" and "variation" used herein are used interchangeably and refer to changes in gene structure, which result in variant forms that can be passed on to offspring due to changes in base units in DNA or deletions, insertions, or rearrangements of larger portions of genes or chromosomes. Herein, somatic variation refers to mutations that occur in normal body cells, such as mutations that occur in the skin or organs. Mutations include, but are not limited to: single nucleotide polymorphisms (SNPs), whose mutant variants are referred to as single nucleotide variations (SNVs); insertions and deletions; and copy number variations (CNVs), chromosome structural variations (SVs). Some mutations are associated with diseases (e.g., tumors).
[0160] Single nucleotide polymorphisms (SNPs) are variations of a single nucleotide occurring at a specific position in the genome, where each variation is present to some appreciable degree (eg, >1%) in a population.
[0161] The term "allele frequency" as used herein is the frequency of an allele of a gene (or a variant of a gene) relative to all alleles of a gene, which can be expressed as a fraction or percentage. Allele frequency is often associated with a specific genomic locus, because genes are usually located at one or more loci.
[0162] As used herein, the term "variant allele frequency" refers to the frequency of a variant allele relative to all alleles.
[0163] As used herein, the term "machine learning" or "machine learning model" refers to a technique that predicts output base calls based on known outcomes (training data). The known outcome may be a hypothetical sequence, assuming that the hypothetical sequence is correct. Since the model attempts to predict the outcomes of the training data, the machine learning may be supervised learning, where the supervision comes from the training data.
[0164] Examples and drawings are provided below to help understand the present invention. However, it should be understood that these examples and drawings are only used to illustrate the present invention, but do not constitute any limitation. The actual scope of protection of the present invention is set forth in the claims. It should be understood that any modifications and changes can be made without departing from the spirit of the present invention. The reagents and / or kits used in the following examples are all commercially available or can be synthesized by known methods.
[0165] It should be noted that, if the specific conditions are not specified in the examples, the experimental conditions are carried out according to conventional conditions, manufacturer recommendations or publicly reported experimental conditions. If the manufacturer of the reagents or instruments used is not specified, they are all conventional products that can be purchased commercially. If the manufacturer of the reagents used is specified, similar products from other manufacturers are substitutable.
[0166] Example
[0167] Example 1: Verification of germline differentiation results
[0168] 1. Obtain tumor genome comparison results
[0169] 1.1 First, 2158 clinical tumor plasma samples and blood cell (i.e., white blood cell) samples, as well as 2264 clinical tumor tissue samples and blood cell (i.e., white blood cell) samples were collected, somatic cell DNA was extracted from the samples, randomly sheared into short DNA fragments, and adapter sequences were added to construct libraries using a human DNA library construction kit (Gene Plus);
[0170] 1.2 Further, the target region of the DNA library in step 1.1 was captured using a library capture kit (Gene Plus);
[0171] 1.3 Further, the somatic cell DNA sequence obtained in step 1.2 was subjected to second-generation sequencing by PE100 using the DNBSeq T7 platform to obtain sequencing data;
[0172] 1.4 Use fastp software to clean the sequencing results of step 1.3, filter out the adapter sequences, filter out the sequences with fragment length less than 35 nt, and filter out the sequences with N base content exceeding 10%, to obtain high-quality sequencing sequences;
[0173] 1.5 Use bwa-mem2 software to align the high-quality sequence obtained in 1.4 to the hg19 reference gene to obtain the alignment result.
[0174] 2. Obtaining genomic variation in tumor samples
[0175] 2.1 Obtaining genomic variation of tumor samples using tumor-blood cell pairing analysis
[0176] 2.1.1 Based on the comparison results between blood cells and the hg19 reference genome, HaplotypeCaller software was used for mutation detection to obtain germline mutations in blood cells.
[0177] 2.1.2 Based on the comparison results of tumor plasma samples or tissue samples and their corresponding blood cells, RealDcaller2 or Mutect2 software is used for joint mutation detection (paired analysis) to obtain somatic mutations in tumor samples.
[0178] 2158 clinical tumor plasma samples and their blood cells obtained a total of 10298 somatic mutations and 28632 germline mutations, which are used as plasma dataset 1. In addition, for 2264 clinical tumor tissue samples and their blood cell samples, a total of 31674 somatic mutations and 41738 germline mutations were obtained, which are used as tissue dataset 2. The results are shown in Table 1 below.
[0179] Table 1
[0180]
[0181]
[0182] 2.2 Obtaining genomic variation in a single tumor sample
[0183] 2.2.1 Further, the comparison results of the tumor plasma sample or tumor tissue sample obtained in step 1.5 with the hg19 reference genome were used to perform mutation detection using the mutation detection software RealDcaller2. The detection results were the mutation sites in the tumor single sample DNA sequence that were inconsistent with the reference genome sequence, and the mutation frequency AF corresponding to each mutation site was obtained at the same time;
[0184] 2.2.2 Further, the variant sites obtained in 2.2.1 were annotated using the annotation software SomVAS, and the population frequency Freq of the variants was annotated using the public population frequency databases gnomAD, dbSNP, 1000G, ExAC, and the Chinese population frequency database PAD (which was established by Genentech by performing genomic testing on normal blood cells of 89,767 cancer patients). DB ;
[0185] 3. Obtain copy number variation results for sample intervals
[0186] 3.1 Obtain the sequencing results using the leukocyte DNA of 50 normal subjects according to steps 1.1-1.4, and use CNVkit software to construct a normal subject baseline;
[0187] 3.2 Using CNVkit software, with the normal human baseline of 3.1 as the control, the genomic depth signal of the alignment results obtained in 1.5 was detected, and the depth signal was corrected using the genomic GC content, and the cyclic binary segmentation algorithm (CBS) was used to divide the genome into multiple copy number intervals according to the depth signal;
[0188] 4. Obtain the results of the variable site frequency test
[0189] 4.1 For the copy number intervals divided in 3.2, all mutations were detected in each interval, and homozygous and heterozygous classes were delineated according to the mutation frequency and minor allele frequency (MAF). Using the K-means clustering method, the variant sites were divided into two categories according to the minor allele frequency MAF value, and the average MAF values of the two types of variants obtained by clustering were compared. The class with a large average MAF was the heterozygous class, and the class with a small average MAF was the homozygous class; in the heterozygous variant site class, the variant sites that met 0.1≤AF≤0.5 were heterozygous variant sites; in the homozygous variant site class, the sites that met AF>0.5 were homozygous sites; when there were too few sites to be clustered, the variant sites that met AF>0.8 were defined as homozygous sites, and the variant sites that met 0.2≤AF≤0.8 were defined as heterozygous sites.
[0190] The minor allele frequency calculation formula 1 is:
[0191] MAF = min(AF, 1-AF) Formula 1
[0192] 4.2 Based on MAF, use Formula 2 to calculate the theoretical frequency of the heterozygous class. Formula 2 is
[0193]
[0194] Among them, W i Refers to the weight of the variant site calculated according to Formula 4, MAF i It refers to the minor allele frequency of the variant site obtained according to Formula 1, and i refers to the serial number of the SNP site in the interval.
[0195] 4.3 According to AF, the theoretical variation frequency of the homozygous class is calculated using formula 3, which is
[0196]
[0197] Among them, W i Refers to the weight of the variant site calculated according to Formula 4, AFi refers to the variation frequency obtained in step S4a, i refers to the serial number of the homozygous site;
[0198] 4.4 Weight W in Formula 2 and Formula 3 in Steps 4.2 and 4.3 i The calculation formula is:
[0199]
[0200] Dep i Refers to the depth of the variant site, ΔMAF i Refers to the difference between the MAF of the variant site calculated by formula 1 in step 4.1 and the upper quartile of the frequency of the variant site in the copy number interval. α is a training constant term, which can be 0.01, 0.05, etc.
[0201] 4.5 Perform a binomial test for each variant and obtain the significance p value of the binomial test for the heterozygous class germP het ;
[0202] 4.6 Perform a binomial test for each variant and obtain the significance p value of the binomial test for the homozygous class germP homo ;
[0203] 4.7 For all autosomes, confirm the final germP value according to Formula 5, which is:
[0204] germP=max(germP homo , germP het ) Formula 5
[0205] 4.8 Determine the germP value of the sex chromosome based on the input gender: When the gender is female, the variation on the X chromosome is determined according to Formula 5 to determine the value to be used for calculation; When the gender is male, the variation of the homologous segments on the X chromosome and the Y chromosome is determined according to Formula 5 to determine the value to be used for calculation; The non-homologous segments on the X chromosome and the Y chromosome germP = germP homo ;
[0206] 5. Decision tree interpretation test results: Use the constructed decision tree logic and the significance test germP and population frequency value Freq DB , and the mutation frequency AF determines whether the mutation source is germline. In the decision tree model, the weight order of the features is as follows: germP>Freq DB >AF.
[0207] Based on the method of detecting germline mutations in single tumor sample DNA, the results of germline mutations and somatic mutations of 2158 clinical tumor plasma samples and 2264 clinical tumor tissue samples were obtained as shown in Table 2 below:
[0208] Table 2
[0209]
[0210]
[0211] Note: TG stands for True Germline mutation, which refers to the number of mutations that are judged as germline mutations by the single-sample variation detection method disclosed herein, and are also judged as germline mutations by comparison of blood cells with the hg19 reference genome (Table 1);
[0212] FS is False Somatic mutation, which refers to the number of mutations that are judged as somatic mutations by the single-sample variation detection method disclosed in the present invention, but are judged as germline mutations by comparing blood cells with the hg19 reference genome (Table 1);
[0213] TS is True Somatic mutation, which refers to the number of mutations that are judged as somatic mutations by the single-sample variation detection method disclosed in the present invention and are also judged as somatic mutations by tumor-blood cell pairing analysis (Table 1);
[0214] FG is False Germline mutation, which refers to the number of mutations that are judged as germline mutations by the single-sample variation detection method disclosed herein, but are judged as somatic mutations by tumor-blood cell pairing analysis (Table 1);
[0215] G_PPA is the germline sensitivity, and the calculation formula is G_PPA = TG / (TG + FS),
[0216] S_PPA is the system sensitivity, and the calculation formula is S_PPA=TS / (TS+FG).
[0217] The paired sample variation detection results of the comparison samples showed that the sensitivity of the single sample method disclosed in the present invention for germline variation detection reached 97.4% in tissue samples and 97.3% in plasma samples; the sensitivity of system variation detection reached 99.0% in tissue samples and 97.9% in plasma samples. This result shows that the detection method disclosed in the present invention has high discrimination ability for germline variation detection of single tumor tissue samples and single tumor plasma samples.
[0218] The technical solution of the present invention is not limited to the above-mentioned specific embodiments. All technical variations made according to the technical solution of the present invention fall within the protection scope of the present invention.
Claims
1. A method for distinguishing the types of gene mutations in a sample to be tested, characterized in that: The method comprises the following steps: Compare the sequencing data of the sample to be tested with the reference genome to obtain comparison data; Based on the comparison data, one or more variant site information is obtained, and the variant frequency AF and population frequency Freq of each variant site are obtained. DB ; Based on the comparison data, dividing the genome of the sample to be tested into multiple copy number intervals; Clustering one or more variant sites in each copy number interval to obtain homozygous variant site classes and heterozygous variant site classes; Performing a difference test on the one or more variant sites and the theoretical germline homozygous allele frequency and the theoretical germline heterozygous allele frequency of the copy number interval in which the one or more variant sites are located, respectively, to obtain a test germP value for each variant site; and The variation frequency AF and the population frequency Freq of the one or more variation sites are DB The test germP value is input into the interpretation model to obtain the type of the one or more variable sites.
2. The method according to claim 1, characterized in that The one or more variant site information includes at least one of single nucleotide variation (SNV), insertion / deletion variation (InDel), copy number variation (CNV), and chromosome structural variation (SV).
3. The method according to claim 1 or 2, characterized in that: Dividing the genome of the sample to be tested into a plurality of copy number intervals comprises: using a cyclic binary segmentation algorithm (CBS) to divide the genome of the sample to be tested into a plurality of copy number intervals; Preferably, the cyclic binary segmentation algorithm (CBS) divides the genome of the sample to be tested into multiple copy number intervals based on the sequencing depth obtained from the comparison data; Preferably, the cyclic binary segmentation algorithm (CBS) divides regions with similar sequencing depth into one copy number interval.
4. The method according to any one of claims 1 to 3, characterized in that The crowd frequency Freq DB The method is obtained by annotating the one or more variant sites using a population frequency database.
5. The method according to any one of claims 1 to 4, characterized in that The step of clustering one or more variant sites in each copy number interval comprises: obtaining a minor allele frequency MAF based on the variant frequency AF, and then clustering one or more variant sites in each copy number interval based on the minor allele frequency MAF, Preferably, the minor allele frequency MAF is obtained according to the following formula 1: MAF=min(AF,1-AF) Formula 1 Preferably, the clustering comprises: using a K-means clustering algorithm to divide the multiple variant sites in each copy number interval into two categories according to the minor allele frequency MAF, comparing the average MAF values of the two categories of variant sites obtained by clustering, judging the category with a larger average MAF as a heterozygous variant site category, and judging the category with a smaller average MAF as a homozygous variant site category, Preferably, for the heterozygous variant site class, a variant site with 0.1≤AF≤0.5 is determined as a heterozygous variant site; Preferably, in the homozygous variant site class, the variant site with AF>0.5 is judged as a homozygous variant site. Preferably, when the number of variant sites within the copy number interval is insufficient for clustering, variant sites with AF>0.8 are judged as homozygous sites, and variant sites with 0.2≤AF≤0.8 are judged as heterozygous sites.
6. The method according to any one of claims 1 to 5, characterized in that The theoretical germline heterozygous allele frequency is obtained according to the following formula 2: Among them, W i Refers to the weight of the i-th variant site, MAF i refers to the minor allele frequency of the ith variant site, and / or The theoretical germline homozygous allele frequency is obtained according to the following formula 3: Among them, W i Refers to the weight of the i-th variant site, AF i refers to the mutation frequency of the ith mutation site, The weight of the i-th variant site is obtained according to the following formula 4: Among them, Dep i refers to the sequencing depth of the ith variant site, ΔMAF i It refers to the difference between the MAF of the ith variant site and the upper quartile of the variant frequency AF of one or more variant sites in the copy number interval to which the ith variant site belongs, α is a training constant term, Preferably, the training constant term α is selected from 0.01, 0.02, 0.03, 0.04, 0.05, and more preferably 0.
05.
7. The method according to any one of claims 1 to 6, characterized in that The difference test is selected from t-test, analysis of variance or chi-square test, Preferably, the difference test is a binomial difference test, More preferably, the difference test is a two-tailed binomial difference test.
8. The method according to any one of claims 1 to 7, characterized in that The test germP value of one of the one or more variant sites is obtained by the following steps: The variation frequency AF of the variation site is compared with the theoretical germline homozygous allele frequency THEOAF homo Perform a difference test and obtain the test germP homo value; The minor allele frequency MAF of the variant site is compared with the theoretical germline heterozygous allele frequency THEOAE het Perform a difference test and obtain the test germP het value; and The test germP value is obtained according to the following formula 5: germP = max(germP homo ,germP het ) Formula 5.
9. The method according to any one of claims 1 to 8, characterized in that The method further includes determining the gender of the individual from whom the sample to be tested originates, Preferably, when the sample to be tested is from a male individual and the one variant site is located in a non-homologous region of a sex chromosome, germP = germP homo .
10. The method according to any one of claims 1 to 9, characterized in that The interpretation model is constructed using a machine learning model. Preferably, the algorithm of the machine learning model includes at least one of K nearest neighbor, naive Bayes classifier, logistic regression, decision tree, random forest, support vector machine, neural network, and AdaBoost. Preferably, the interpretation model is constructed using a decision tree, more preferably using a sklearn decision tree algorithm. Preferably, the decision tree model is constructed by inputting the variation frequency AF of one or more variation sites, the population frequency Freq DB and test germP value to construct, Preferably, in the decision tree model, the weight order is germP>Freq DB >AF, Preferably, the method is used to determine whether the one or more variant sites belong to germline variants or somatic variants.
11. A device for distinguishing the types of gene mutations in a sample to be tested, characterized in that: The device comprises: A comparison data acquisition module is used to compare the sequencing data of the sample to be tested with the reference genome to obtain comparison data; The variant site information analysis module is used to obtain one or more variant site information based on the comparison data, the variant frequency AF and the population frequency Freq of each variant site DB ; A copy number interval division module, used for dividing the genome of the sample to be tested into multiple copy number intervals based on the comparison data; A clustering module, used to cluster one or more variant sites in each copy number interval to obtain homozygous variant site classes and heterozygous variant site classes; a difference test module, for performing a difference test on the one or more variant sites respectively with the theoretical germline homozygous allele frequency and the theoretical germline heterozygous allele frequency of the copy number interval in which the variant sites are located, to obtain a test germP value for each variant site; and The variant site type judgment module is used to convert the variant frequency AF and the population frequency Freq of the one or more variant sites into DB The test germP value is input into the interpretation model to obtain the type of the one or more variable sites.
12. The device according to claim 11, characterized in that The device also includes a sequencing data acquisition module, which is used to acquire sequencing data of the sample to be tested.
13. The device according to claim 11 or 12, characterized in that The copy number interval division module uses a cyclic binary segmentation algorithm (CBS) to divide the genome of the sample to be tested into multiple copy number intervals; Preferably, the clustering module obtains the minor allele frequency MAF based on the variant frequency AF, and then clusters one or more multiple variant sites in each copy number interval based on the minor allele frequency MAF. Preferably, the minor allele frequency MAF is obtained according to the following formula 1: MAF=min(AF,1-AF) Formula 1, Preferably, the clustering includes: using a K-means clustering algorithm to divide the multiple variant sites in each copy number interval into two categories according to the minor allele frequency MAF, comparing the average MAF values of the two categories of variant sites obtained by clustering, and judging the category with a larger average MAF as a heterozygous variant site category, and judging the category with a smaller average MAF as a homozygous variant site category; Preferably, for the heterozygous variant site class, a variant site with 0.1≤AF≤0.5 is determined as a heterozygous variant site; Preferably, in the homozygous variant site class, the variant site with AF>0.5 is judged as a homozygous variant site; Preferably, when the number of variant sites within the copy number interval is insufficient for clustering, variant sites with AF>0.8 are judged as homozygous sites, and variant sites with 0.2≤AF≤0.8 are judged as heterozygous sites.
14. The device according to any one of claims 11 to 13, characterized in that The device also includes a theoretical germline allele frequency acquisition module, which is used to acquire theoretical germline homozygous allele frequencies and theoretical germline heterozygous allele frequencies. Preferably, the theoretical germline heterozygous allele frequency is obtained according to the following formula 2: Among them, W i Refers to the weight of the i-th variant site, MAF i refers to the minor allele frequency of the ith variant site, Preferably, the theoretical germline homozygous allele frequency is obtained according to the following formula 3: Among them, W i Refers to the weight of the i-th variant site, AF i refers to the mutation frequency of the ith mutation site, Preferably, the weight of the i-th variant site is obtained according to the following formula 4: Among them, Dep i refers to the sequencing depth of the ith variant site, ΔMAF i It refers to the difference between the MAF of the ith variant site and the upper quartile of the variant frequency AF of one or more variant sites in the copy number interval to which the ith variant site belongs, α is a training constant term, Preferably, the training constant term α is selected from 0.01, 0.02, 0.03, 0.04 or 0.05, more preferably 0.
05.
15. The device according to any one of claims 11 to 14, characterized in that The difference test module uses t test, variance analysis, chi-square test, etc. to perform difference test. Preferably, the difference test is a binomial difference test, more preferably a two-tailed binomial difference test, Preferably, the test germP value of one of the one or more variant sites is obtained by the following steps: The variation frequency AF of the variation site is compared with the theoretical germline homozygous allele frequency THEOAF homo Perform a difference test and obtain the test germP homo value; The minor allele frequency MAF of the variant site is compared with the theoretical germline heterozygous allele frequency THEOAF het Perform a difference test and obtain the test germP het value; and The test germP value is obtained according to the following formula 5: germP = max(germP homo ,germP het ) Formula 5, Preferably, the device further comprises a gender determination module for determining the gender of the individual from which the sample to be tested originates. When the gender determination module determines that the sample to be tested is from a male individual according to the sex chromosome of the sample to be tested, and the one variant site is located in a non-homologous region of the sex chromosome, germP=germP homo , Preferably, the device further comprises a judgment model construction module for judging whether the one or more variation sites belong to germline variation or somatic variation.
16. The device according to any one of claims 11 to 15, characterized in that The device is used to perform the method according to any one of claims 1 to 10.
17. A storage medium having a computer-executable program stored thereon, wherein the program is configured to execute, when running, the method for distinguishing the type of gene variation in a sample to be tested according to any one of claims 1 to 10.
18. A device, comprising: Memory, used to store programs; A processor is used to execute the program stored in the memory, wherein the program is configured to execute the method for distinguishing the type of gene variation in a sample to be tested according to any one of claims 1 to 10 when running.
Citation Information
Patent Citations
Model construction method for identifying tumour purity sample and application thereof
CN110808081A
Sample quality evaluation method and device
CN111477277A
Single-sample tumor somatic mutation discrimination and TMB detection method based on NGS platform
CN114694750A
Non-contrast HRD detection method, system and device
CN117497056A
Computational modeling of loss of function based on allelic frequency
US20200273538A1