Method and apparatus for detecting variations in a single sample embryo system

CN120015109BActive Publication Date: 2026-08-18GENEPLUS-BEIJING CLINICAL LAB CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411945308.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2026-08-18
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

然而,在临床应用中,很多肿瘤细胞样本是无法获得对照组织的,这会导致突变检测时,无法从肿瘤体细胞突变中识别出突变来源

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_10
    Figure QLYQS_10
  • Figure QLYQS_43
    Figure QLYQS_43
  • Figure BDA0005213965210000031
    Figure BDA0005213965210000031
Patent Text Reader

Abstract

The present disclosure relates to methods and devices for detecting somatic variation in a single sample. The present disclosure specifically relates to methods for distinguishing between types of genetic variation in a sample under test, the methods of the present disclosure being capable of distinguishing between whether a mutation in a sample is a somatic mutation or a germline mutation based on the sample alone.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a method and apparatus for detecting embryological variations from a single tumor sample. Background Technology

[0002] Variation sites in the human genome can be divided into germline variations and somatic variations. Germline variations refer to variations inherited from parents, and these variations are not necessarily harmful. Somatic mutations refer to variations occurring in the normal cells of an organism. These mutations are acquired during individual development and may be caused by internal factors such as the accumulation of reactive oxygen species, or external factors such as radiation and toxins, resulting in changes to the genome sequence. Somatic mutations may lead to cancer.

[0003] High-throughput sequencing methods can be used to obtain the genomic sequence of sample cells, and then bioinformatics algorithms can be used to detect mutations from the genomic sequence. In the field of tumor gene detection, detecting targeted mutations in tumor cells can provide guidance for tumor treatment. To ensure that the source of mutation is tumor cells, a common approach is to use paired normal tissue cells as a control, identifying tumor-specific mutations not present in the control cells during mutation detection. However, in clinical applications, control tissues are often unavailable for many tumor cell samples, which makes it impossible to identify the source of mutation from tumor somatic cell mutations during mutation detection. Therefore, a method capable of identifying germline variations from tumor somatic cell genome sequencing results is necessary. Summary of the Invention

[0004] This disclosure aims to provide a method and apparatus for distinguishing germline variations and somatic variations in a single tumor sample.

[0005] According to one aspect of this disclosure, a method for distinguishing gene variant types in a sample to be tested is provided, the method comprising the following steps:

[0006] The sequencing data of the sample to be tested is compared with the reference genome to obtain the comparison data;

[0007] Based on the comparison data, information on one or more variant sites is obtained, including the variant frequency AF and the population frequency Freq for each variant site. DB ;

[0008] Based on the comparison data, the genome of the sample to be tested is divided into multiple copy number intervals;

[0009] Cluster one or more variant sites within each copy number interval to obtain homozygous variant site class and heterozygous variant site class;

[0010] One or more variant sites from the homozygous variant site class and the heterozygous variant site class are subjected to difference tests with the theoretical homozygous allele frequency and the theoretical heterozygous allele frequency in the germline, respectively, to obtain the GermP value for each variant site; and

[0011] The mutation frequency AF and population frequency Freq of the one or more mutation sites are used to determine the mutation frequency AF and population frequency Freq. DB The germP value is input into the interpretation model to obtain the type of the one or more variant sites.

[0012] In some embodiments, the one or more mutation site information includes, but is not limited to, at least one of single nucleotide variants (SNVs), insertion / deletion variants (InDel), copy number variants (CNVs), and chromosomal structural variants (SVs).

[0013] In some implementations, the one or more mutation site information includes at least one of single nucleotide variants (SNVs) and insertion / deletion variants (InDel).

[0014] In some implementations, dividing the genome of the sample to be tested into multiple copy number intervals includes: using a cyclic binary segmentation algorithm (CBS) to divide the genome of the sample to be tested into multiple copy number intervals.

[0015] In some implementations, the Cyclic Binary Segmentation (CBS) algorithm divides the genome of the sample to be tested into multiple copy number intervals based on the sequencing depth obtained from the alignment data.

[0016] The sequencing depth can be obtained using software (e.g., CNVkit software) or methods known to those skilled in the art. In an optional embodiment, obtaining the sequencing depth may include the following steps: constructing a normal human sequencing baseline, using the normal human sequencing baseline as a control, and obtaining the sequencing depth of the alignment data. In a preferred embodiment, the sequencing depth is corrected using genomic GC content.

[0017] In some implementations, the Cyclic Binary Partitioning (CBS) algorithm divides regions with similar sequencing depths into a single copy number interval.

[0018] In some implementations, the regions with similar depth signals refer to regions where there is no significant difference in depth signals between adjacent sites based on the Cyclic Binary Segmentation (CBS) algorithm.

[0019] In some implementations, the crowd frequency Freq DB This can be obtained by annotating one or more variant sites using a population frequency database.

[0020] Those skilled in the art will understand that the population frequency database may include those databases conventionally used in the art. In some embodiments, the population frequency database may include, but is not limited to, gnomAD, dbSNP, 1000G, 1000G Asian, ExAC, ExAC Asian, Chinese Population Frequency Database PAD, etc.

[0021] In some implementations, the sequencing data of the sample to be tested can be obtained by, for example, but not limited to, next-generation sequencing (NGS) technology.

[0022] In some implementations, the step of clustering one or more variant sites within each copy number interval may include: obtaining the minor allele frequency (MAF) based on the variant frequency (AF), and then clustering one or more variant sites within each copy number interval based on the minor allele frequency (MAF).

[0023] In some embodiments, the minor allele frequency (MAF) is obtained according to the following formula 1:

[0024] MAF=min(AF,1-AF) Formula 1.

[0025] In some implementations, the clustering may include: using the K-means clustering algorithm to divide one or more variant sites within each copy number interval into two classes based on the minor allele frequency (MAF), comparing the average MAF values ​​of the two classes of variant sites obtained from the clustering, and determining the class with the larger average MAF as the heterozygous variant site class and the class with the smaller average MAF as the homozygous variant site class.

[0026] In some implementations, for the heterozygous variant site class, variant sites with 0.1≤AF≤0.5 and 0.1≤MAF≤0.5 are identified as heterozygous variant sites.

[0027] In some implementations, within the class of homozygous variant sites, variant sites with AF > 0.5 are identified as homozygous variant sites.

[0028] In some implementations, when the number of variant sites within the copy number interval is insufficient for clustering, variant sites with AF > 0.8 are identified as homozygous sites, and variant sites with 0.2 ≤ AF ≤ 0.8 are identified as heterozygous sites.

[0029] In some embodiments, the theoretical germline heterozygous allele frequencies can be obtained according to the following formula 2:

[0030]

[0031] Among them, W iThe weight of the i-th mutation site, MAF i The frequency of the second allele at the i-th mutation site.

[0032] In some embodiments, the theoretical germline homozygous allele frequencies can be obtained according to the following formula 3:

[0033]

[0034] Among them, W i The weight of the i-th mutation site, AF i The mutation frequency refers to the mutation frequency of the i-th mutation site.

[0035] In some implementations, the weight of the i-th mutation site can be obtained according to the following formula 4:

[0036]

[0037] Among them, Dep i The sequencing depth of the i-th variant site, ΔMAF i The difference between the MAF of the i-th variant site and the upper quartile of the mutation frequency AF of one or more variant sites in the copy number interval to which the i-th variant site belongs, where α is a training constant.

[0038] In some implementations, the training constant α can be selected from, but is not limited to, 0.01, 0.02, 0.03, 0.04, 0.05, etc.

[0039] In some embodiments, the difference test may include, but is not limited to, t-tests, analysis of variance, chi-square tests, etc. In some embodiments, the difference test may be a binomial difference test. In some embodiments, the difference test may be a two-tailed binomial difference test.

[0040] In some implementations, the test germP value of one of the one or more variant sites can be obtained through the following steps:

[0041] The mutation frequency AF of the one mutation site is compared with the homozygous allele frequency THEOAF of the theoretical germline. homo Perform a difference test to obtain the test result (germP). homo value;

[0042] The secondary allele frequency MAF at the mutation site was compared with the heterozygous allele frequency THEOAF of the theoretical germline. het Perform a difference test to obtain the test result (germP). het Value; and

[0043] The test germP value is obtained according to the following formula 5:

[0044] germP = max(germP) homo ,germP het ) Formula 5.

[0045] In some implementations, the method may further include determining the gender of the individual from whom the sample to be tested originates.

[0046] In some implementations, when the sample to be tested comes from a male individual and the mutation site is located in a non-homologous region of a sex chromosome, germP = germP homo .

[0047] In some implementations, the interpretation model can be constructed using a machine learning model.

[0048] In some implementations, the algorithm of the machine learning model includes, but is not limited to, at least one of K-nearest neighbors, Naive Bayes classifier, logistic regression, decision tree, random forest, support vector machine, neural network, and AdaBoost.

[0049] In some implementations, the interpretation model is constructed using a decision tree model.

[0050] In some implementations, the decision tree model can be derived by inputting the mutation frequency AF of one or more mutation sites and the population frequency Freq. DB The germP value is used to construct the decision tree model. In some implementations, the feature weights are ordered as follows in the decision tree model: germP > Freq. DB >AF.

[0051] In a specific implementation, the interpretation model is constructed using the sklearn decision tree algorithm. In some implementations, the sklearn decision tree algorithm can utilize the mutation frequency AF of one or more mutation sites and the population frequency Freq. DB The germP value was tested, and the model with the highest selectivity score was obtained through cross-validation.

[0052] In some implementations, the interpretation model can be constructed using a DecisionTreeClassifier. In some implementations, the cross-validation method can be, for example, 3x cross-validation. For instance, 2 / 3 of the mutation sites can be used for training, and the remaining 1 / 3 for validation. In some implementations, the GridSearchCV function can be used to perform a grid search on the parameters (e.g., max_depth, min_samples_leaf, and min_samples_split) within the DecisionTreeClassifier. This process aims to prevent overfitting and select the model with the highest performance score.

[0053] In some implementations, the method can determine whether the one or more mutation sites belong to germline variation or somatic variation.

[0054] In some implementations, the variant site may include any site in the test sample that is inconsistent with the reference genome. For example, the variant site may include a site with one or more base insertions, deletions, or substitutions compared to the reference genome.

[0055] In some implementations, the sample is derived from an individual suffering from a disease. Those skilled in the art will understand that the sample can be derived from an individual suffering from any disease. In an exemplary implementation, the sample is a tumor sample or a sample from another disease.

[0056] In some embodiments, the other diseases include, but are not limited to, adverse reactions after organ transplantation; infections caused by pathogens such as bacteria, viruses, fungi, and parasites; hepatitis; chronic gastroenteritis; diabetes; chronic pharyngitis; cardiovascular and cerebrovascular diseases; and lung-related diseases.

[0057] In some embodiments, the sample to be tested includes, but is not limited to, body fluid samples or tissue samples. In some embodiments, the body fluid sample includes, but is not limited to, plasma, urine, saliva, bronchoalveolar lavage fluid, or cerebrospinal fluid from blood.

[0058] The method disclosed herein can distinguish between somatic mutations and germline mutations in tumor samples or other disease samples, without the need for comparative analysis with control samples (such as blood cell samples, reference genomes, etc.).

[0059] In some embodiments, the reference genome includes, but is not limited to, GRCH37, b37, hs37d5 (b37+decoy), hg19, GRCH38 (hg38), etc. In some embodiments, the hg19 genome can be downloaded from UCSC (http: / / genome.ucsc.edu / ). In some embodiments, the GRCH38 genome can be downloaded from NCBI (https: / / www.ncbi.nlm.nih.gov / ). In some embodiments, if the sample is taken from a human body, the sequencing data is compared with the human reference genome. In some embodiments, if the sample is taken from an animal, the sequencing data is compared with the reference genome of the corresponding animal species.

[0060] According to another aspect of this disclosure, an apparatus for distinguishing gene variant types in a sample to be tested is provided, the apparatus comprising:

[0061] The alignment data acquisition module is used to align the sequencing data of the sample to be tested with the reference genome to obtain alignment data;

[0062] The variant site information analysis module is used to obtain information on one or more variant sites based on the comparison data, including the variant frequency AF and population frequency Freq for each variant site. DB ;

[0063] The copy number interval division module is used to divide the genome of the sample to be tested into multiple copy number intervals based on the alignment data.

[0064] The clustering module is used to cluster one or more variant sites within each copy number interval to obtain homozygous variant site classes and heterozygous variant site classes.

[0065] The difference testing module is used to perform difference tests on one or more variant sites of the homozygous variant site class and the heterozygous variant site class with the theoretical homozygous allele frequency and the theoretical heterozygous allele frequency of the germline, respectively, to obtain the test germP value for each variant site; and

[0066] The variant site type determination module is used to determine the variant frequency AF and population frequency Freq of one or more variant sites. DB The germP value is input into the interpretation model to obtain the type of the one or more variant sites.

[0067] In some embodiments, the apparatus may further include a sequencing data acquisition module for acquiring sequencing data of the sample to be tested.

[0068] In some embodiments, the apparatus is used to perform the methods described above in this disclosure.

[0069] In some implementations, the copy number interval partitioning module uses a cyclic binary segmentation algorithm (CBS) to divide the genome of the sample to be tested into multiple copy number intervals.

[0070] In some implementations, the crowd frequency Freq DB This can be obtained by annotating one or more variant sites using a population frequency database.

[0071] In some implementations, the clustering module may obtain the minor allele frequency (MAF) based on the mutation frequency (AF), and then cluster one or more mutation sites within each copy number interval based on the minor allele frequency (MAF).

[0072] In some embodiments, the minor allele frequency (MAF) is obtained according to the following formula 1:

[0073] MAF=min(AF,1-AF) Formula 1.

[0074] In some implementations, the clustering may include: using the K-means clustering algorithm to divide one or more variant sites within each copy number interval into two classes based on the minor allele frequency (MAF), comparing the average MAF values ​​of the two classes of variant sites obtained from the clustering, and determining the class with the larger average MAF as the heterozygous variant site class and the class with the smaller average MAF as the homozygous variant site class.

[0075] In some implementations, for the heterozygous variant site class, variant sites with 0.1≤AF≤0.5 and 0.1≤MAF≤0.5 are identified as heterozygous variant sites.

[0076] In some implementations, within the class of homozygous variant sites, variant sites with AF > 0.5 are identified as homozygous variant sites.

[0077] In some implementations, when the number of variant sites within the copy number interval is insufficient for clustering, variant sites with AF > 0.8 are identified as homozygous sites, and variant sites with 0.2 ≤ AF ≤ 0.8 are identified as heterozygous sites.

[0078] In some embodiments, the apparatus may further include a theoretical germline allele frequency acquisition module for acquiring the theoretical germline homozygous allele frequency and the theoretical germline heterozygous allele frequency.

[0079] In some embodiments, the theoretical germline heterozygous allele frequencies can be obtained according to the following formula 2:

[0080]

[0081] Among them, W i The weight of the i-th mutation site, MAF i The frequency of the second allele at the i-th mutation site.

[0082] In some embodiments, the theoretical germline homozygous allele frequencies can be obtained according to the following formula 3:

[0083]

[0084] Among them, W i The weight of the i-th mutation site, AF i The mutation frequency refers to the mutation frequency of the i-th mutation site.

[0085] In some implementations, the weight of the i-th mutation site can be obtained according to the following formula 4:

[0086]

[0087] Among them, Dep i The sequencing depth of the i-th variant site, ΔMAF i The difference between the MAF of the i-th variant site and the upper quartile of the mutation frequency AF of one or more variant sites in the copy number interval to which the i-th variant site belongs, where α is a training constant.

[0088] In some implementations, the training constant α can be selected from, but is not limited to, 0.01, 0.02, 0.03, 0.04, 0.05, etc.

[0089] In some implementations, the difference testing module may employ, but is not limited to, t-tests, analysis of variance, chi-square tests, etc., to perform difference testing. In some implementations, the difference testing may be a binomial difference test. In some implementations, the difference testing may be a two-tailed binomial difference test.

[0090] In some implementations, the test germP value of one of the one or more variant sites can be obtained through the following steps:

[0091] The mutation frequency AF of the one mutation site is compared with the homozygous allele frequency THEOAF of the theoretical germline. homo Perform a difference test to obtain the test result (germP). homo value;

[0092] The secondary allele frequency MAF at the mutation site was compared with the heterozygous allele frequency THEOAF of the theoretical germline. het Perform a difference test to obtain the test result (germP). hetValue; and

[0093] The test germP value is obtained according to the following formula 5:

[0094] germP = max(germP) homo ,germP het ) Formula 5.

[0095] In some embodiments, the device may further include a gender determination module for determining the gender of the individual from which the sample to be tested originates.

[0096] In some implementations, when the sex determination module determines that the test sample originates from a male individual based on the sex chromosome of the test sample, and the mutation site is located in a non-homologous region of the sex chromosome, then germP = germP. homo .

[0097] In some embodiments, the apparatus may further include a judgment model construction module.

[0098] In some implementations, the interpretation model building module can be constructed using machine learning. In a specific implementation, the interpretation model is constructed using the sklearn decision tree algorithm. In some implementations, the sklearn decision tree algorithm can utilize the mutation frequency AF of one or more mutation sites and the population frequency Freq. DB The germP value was tested, and the model with the highest selectivity score was obtained through cross-validation.

[0099] In some implementations, the interpretation model building module may use a DecisionTreeClassifier to construct the interpretation model. In some implementations, the cross-validation method may be, for example, 3-fold cross-validation. For example, 2 / 3 of the mutation sites may be used for training, and the remaining 1 / 3 for validation. In some implementations, the GridSearchCV function may be used to perform a grid search on the parameters (e.g., max_depth, min_samples_leaf, and min_samples_split) within the DecisionTreeClassifier. This process aims to prevent overfitting and select the model with the highest performance score.

[0100] In some implementations, the method can determine whether the one or more mutation sites belong to germline variation or somatic variation.

[0101] In some implementations, the variant site may include any site in the test sample that is inconsistent with the reference genome. For example, the variant site may include a site with one or more base insertions, deletions, or substitutions compared to the reference genome.

[0102] According to another aspect of this disclosure, a storage medium is provided that stores a computer-executable program configured to execute the method of this disclosure for distinguishing gene variant types in a sample to be tested.

[0103] According to another aspect of this disclosure, an apparatus is provided, the apparatus comprising:

[0104] Memory, used to store programs;

[0105] A processor for executing a program stored in the memory, the program being configured to run and perform the method of the present disclosure for distinguishing gene variant types in a sample to be tested. Attached Figure Description

[0106] Figure 1 An exemplary flowchart illustrating the differentiation of gene variant types according to one embodiment of this disclosure is shown. Detailed Implementation

[0107] This invention aims to provide a method for accurately distinguishing embryonic variation from somatic variation in a single sample (e.g., a tumor sample).

[0108] To achieve the above objectives, this disclosure provides a method and apparatus for detecting embryonic system variations in a single tumor sample.

[0109] Figure 1 An exemplary method for detecting embryological variations in a single tumor sample according to this disclosure is shown.

[0110] In a specific implementation, the method for detecting embryonic system variation using a single test sample of this disclosure includes the following steps:

[0111] S1. Gene sequencing: Gene sequencing of a single sample was performed using an NGS platform;

[0112] S2. Sample data alignment: Aligning sequencing data to the human reference genome;

[0113] S3. Single-sample interval copy number variation detection: Using alignment data, different copy number intervals are divided based on different sequencing depths on the genome;

[0114] S4. Single-sample variant annotation: Variant signals are detected from the alignment data using variant detection software, the variant frequency AF is obtained, and the population frequency of the variant site is annotated using a population frequency database.

[0115] S5. Define heterozygous and homozygous variant sites within the interval: Based on the mutation frequency AF, variant sites in the same copy number interval are divided into heterozygous variant site class and homozygous variant site class;

[0116] S6. Define the theoretical germline allele frequencies: Calculate the theoretical germline heterozygous allele frequencies and theoretical germline homozygous allele frequencies within the copy number intervals determined in step S5, respectively.

[0117] S7. Locus frequency difference test: Using the binomial test, the difference between each locus and the theoretical allele frequency determined in step S6 is tested to obtain the test germP value;

[0118] S8. Construct a mutation source determination model: Use the mutation frequency AF and population frequency Freq obtained in step S4. DB The test germP values ​​obtained in step S7 are used to construct a decision tree model for determining the source of variation.

[0119] S9. Determining the source of variation: Using the determination model constructed in step S8, determine whether a single variation originates from the germline or the phylogenetic line.

[0120] In some implementations, the sample to be tested is derived from an individual with a disease. Those skilled in the art will understand that the sample to be tested can originate from an individual suffering from any disease. In an exemplary implementation, the sample to be tested is a tumor sample or a sample from another disease.

[0121] In some implementations, the gene sequencing may be targeted sequencing.

[0122] In some implementations, the sample data alignment step S2 includes: aligning the NGS next-generation sequencing data to the hg19 reference genome using bwa-mem2 software.

[0123] In some implementations, the interval copy number variation detection step in step S3 includes the following steps:

[0124] S3a. Use copy number variation detection software to statistically analyze the genome sequence depth signal from the alignment data obtained in step S2;

[0125] S3b. Correct the signal from step S3a using a sequence depth signal baseline constructed from normal control samples;

[0126] S3b. The Cyclic Binary Segmentation (CBS) algorithm is used to divide regions of similar depth signals on the genome into a copy number interval.

[0127] In some implementations, the regions with similar depth signals refer to regions where there is no significant difference in depth signals between adjacent sites based on the Cyclic Binary Segmentation (CBS) algorithm.

[0128] In some implementations, the variant site annotation step in step S4 includes the following steps:

[0129] S4a. Use mutation detection software, such as from the alignment data obtained in step S2, to detect sample mutations and obtain the mutation frequency AF;

[0130] S4b. Use the population database to annotate variants and obtain the population database frequency Freq for the variants. DB .

[0131] In some implementations, the mutation detection software may include, for example, RealDcaller2 and Mutect2.

[0132] In some implementations, step S5, which defines the heterozygous and homozygous mutation sites within the copy number interval, includes the following steps:

[0133] S5a. Calculate the minor allele frequency (MAF) of the variant site according to Formula 1.

[0134] MAF = min(AF, 1-AF) Formula 1

[0135] S5b. Using machine learning methods, cluster all sites within the copy number interval described in step S3 based on the second allele frequency. Use the K-means clustering algorithm to divide the variant sites into two classes according to the MAF value of the second allele frequency. Compare the average MAF values ​​of the two classes obtained from clustering; the class with the larger average MAF is heterozygous, and the class with the smaller average MAF is homozygous. In the heterozygous variant site class, variant sites satisfying 0.1 ≤ AF ≤ 0.5 are considered heterozygous variant sites; in the homozygous variant site class, variant sites satisfying AF > 0.5 are considered homozygous sites. When there are too few sites to cluster, variant sites satisfying AF > 0.8 are defined as homozygous sites, and variant sites satisfying 0.2 ≤ AF ≤ 0.8 are defined as heterozygous sites.

[0136] In step S5a, in formula 1, AF refers to the mutation frequency obtained in S4a.

[0137] In some implementations, the step of defining the theoretical germline allele frequency of the defined interval described in step S6 includes the following steps:

[0138] S6a. Using Formula 2, calculate the theoretical germline heterozygous allele frequency (THEOAF) of the interval based on the interval heterozygous sites determined in step S5. het ), where W in Formula 2 iFrom Formula 4

[0139]

[0140] Among them, W i MAF refers to the weight of the variant site calculated according to Formula 4. i The frequency of the second allele at the variant site obtained according to Formula 1, where i refers to the sequence number of the heterozygous site within the interval;

[0141] S6b. Based on Formula 3, calculate the theoretical germline homozygous allele frequency (THEOAF) of the interval according to the homozygous loci determined in step S6. homo ), where W in Formula 3 i From Formula 4

[0142]

[0143] Among them, W i The AF refers to the weight of the variant site calculated according to Formula 4. i The mutation frequency obtained in step S4a refers to i, where i refers to the sequence number of the homozygous site.

[0144] S6c. Formula 4 mentioned in steps S6a to S6b is:

[0145]

[0146] Among them, Dep i The depth of the mutation site, ΔMAF i The difference between the MAF of the variant site calculated by Formula 1 and the quartile of the frequency of the variant site within the interval determined in step S3b. α is a constant term for training.

[0147] In some implementations, the site frequency difference test step in step S7 includes the following steps:

[0148] S7a. For the frequency of each variant site detected in step S4a, calculate its frequency relative to the theoretical germline homozygous allele frequency (THEOAF) determined in step S6a. homo Perform a two-tailed binomial test and record the test p-value as GermP. homo ;

[0149] S7b. For the frequency of each variant site detected in step S4a, calculate the minimum allele frequency (MAF) of the variant site according to Formula 1, and calculate its relationship with the theoretical germline heterozygous allele frequency (THEOAF) determined in step S6b. het Perform a two-tailed binomial test and record the test p-value as GermP. het ;

[0150] S7c. For all autosomes, confirm the values ​​to be applied in the calculation according to Formula 5.

[0151] germP = max(germP) homo ,germP het ) Formula 5;

[0152] S7d. For sex chromosomes, when the sex is female, the variation on the X chromosome is determined according to Formula 5, and the value to be applied to the calculation is confirmed; when the sex is male, the variation of the homologous segments on the X and Y chromosomes is determined according to Formula 5, and the value to be applied to the calculation is: germP = germP (the variation of the non-homologous segments on the X and Y chromosomes). homo .

[0153] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. The specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention in any way. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of this disclosure. Such structures and techniques have also been described in many publications.

[0154] definition

[0155] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly used in the field to which this invention pertains. For the purposes of interpreting this specification, the following definitions will apply, and where appropriate, terms used in the singular will also include the plural forms, and vice versa.

[0156] Unless the context clearly indicates otherwise, the terms “a” and “an” as used herein include plural references. For example, reference to “a cell” includes multiple such cells and equivalents known to those skilled in the art, etc.

[0157] As used herein, the term "about" indicates a range of ±20% of the following value. In some embodiments, the term "about" indicates a range of ±10% of the following value. In some embodiments, the term "about" indicates a range of ±5% of the following value.

[0158] As used herein, the term “allele” corresponds to one or more nucleotides (which can occur as substitutions or insertions) or the deletion of one or more nucleotides. A “locus” corresponds to a location in the genome. For example, a locus can be a single base or a continuous sequence of bases. The term “genomic location” can refer to a specific nucleotide location or a continuous block of nucleotide locations in the genome. A “heterozygous locus” (also known as a “het”) is a location in a reference genome or a specific genome of a located organism where copies of the chromosome do not share the same allele (e.g., a single nucleotide or a set of nucleotides). When a locus is a single nucleotide with a different allele, a “het” can be a single nucleotide polymorphism (SNP). A “het” can also be a location with one or more nucleotides or one or more tandem repeats of insertion or deletion (collectively referred to as an “insertion / deletion”). A single nucleotide variation (SNV) corresponds to a genomic location having a nucleotide that differs from a reference genome of a specific person. If there is only one nucleotide at the location, the SNV can be homozygous for a person, and if there are two alleles at the location, it is heterozygous. A heterozygous SNV is a het. SNP and SNV can be used interchangeably in this article.

[0159] As used in this article, the terms "mutation" and "variation" are used interchangeably and refer to changes in gene structure that result in a variant form that can be passed on to offspring, caused by alterations in base units of DNA or deletions, insertions, or major rearrangements of genes or chromosomes. In this article, somatic variants refer to mutations occurring in normal cells of a organism, such as those occurring in the skin or organs. Mutations include, but are not limited to: single nucleotide polymorphisms (SNPs), the mutated variants of which are called single nucleotide variations (SNVs); insertions and deletions; and copy number variations (CNVs) and chromosomal structural variations (SVs). Some mutations are associated with diseases such as cancer.

[0160] A single nucleotide polymorphism (SNP) is a variation of a single nucleotide that occurs at a specific location in the genome, where each variation is present in the population at some assessable level (e.g., >1%).

[0161] As used in this article, “allele frequency” is the frequency of an allele of a gene (or a variant of a gene) relative to all its alleles, and can be expressed as a fraction or percentage. Allele frequency is often associated with a specific genomic locus, as genes typically reside at one or more loci.

[0162] The term “variant allele frequency” as used in this article refers to the frequency of a variant allele relative to all alleles.

[0163] As used in this paper, the term "machine learning" or "machine learning model" refers to the technique of predicting output base determinations based on known results (training data). The known results can be hypothetical sequences, which are assumed to be correct. Supervised learning of machine learning can be performed as the model attempts to predict results from the training data, where supervision comes from the training data.

[0164] The following embodiments and accompanying drawings are provided to aid in understanding the present invention. However, it should be understood that these embodiments and drawings are for illustrative purposes only and do not constitute any limitation. The actual scope of protection of the present invention is set forth in the claims. It should be understood that any modifications and changes can be made without departing from the spirit of the invention. The reagents and / or kits used in the following embodiments are commercially available or can be synthesized by known methods.

[0165] It should be noted that, unless specific conditions are specified in the examples, experimental conditions should be performed according to standard conditions, manufacturer recommendations, or publicly reported experimental conditions. Reagents or instruments whose manufacturers are not specified are all commercially available, standard products. For reagents whose manufacturers are specified, similar products from other manufacturers are substitutes.

[0166] Example

[0167] Example 1: Verification of embryologization results

[0168] 1. Obtain tumor genome alignment results

[0169] 1.1 First, 2158 clinical tumor plasma samples and blood cell (i.e., white blood cell) samples, and 2264 clinical tumor tissue samples and blood cell (i.e., white blood cell) samples were collected. Somatic cell DNA was extracted from the samples, randomly fragmented into short DNA fragments, adapter sequences were added, and libraries were constructed using a human DNA library construction kit (GenePlus).

[0170] 1.2 Further, the target region of the DNA library from step 1.1 was captured using a library capture kit (GenePlus);

[0171] 1.3 Further, the somatic cell DNA sequence obtained in step 1.2 was subjected to next-generation sequencing using the DNBSeq T7 platform and PE100 to obtain sequencing data;

[0172] 1.4 The sequencing results from step 1.3 were cleaned using the FASTP software to filter out adapter sequences, sequences with a fragment length of less than 35 nt, and sequences with an N base content of more than 10%, in order to obtain high-quality sequencing sequences.

[0173] 1.5 The high-quality sequence obtained in 1.4 was aligned to the hg19 reference gene using the bwa-mem2 software to obtain the alignment results.

[0174] 2. Obtain genomic variations from tumor samples

[0175] 2.1 Genomic variations in tumor samples were obtained using tumor-blood cell pairing analysis.

[0176] 2.1.1 Based on the alignment results of blood cells with the hg19 reference genome, mutation detection was performed using HaplotypeCaller software to obtain germline mutations in blood cells.

[0177] 2.1.2 Based on the comparison results of tumor plasma or tissue samples and their corresponding blood cell comparison results, RealDcaller2 or Mutect2 software is used to perform joint mutation detection (paired analysis) to obtain somatic mutations in tumor samples.

[0178] From 2158 clinical tumor plasma samples and blood cell samples, a total of 10298 somatic mutations and 28632 germline mutations were obtained; these data constitute the first plasma dataset. Additionally, from 2264 clinical tumor tissue samples and blood cell samples, a total of 31674 somatic mutations and 41738 germline mutations were obtained; these data constitute the second tissue dataset. The results are shown in Table 1 below.

[0179] Table 1

[0180]

[0181]

[0182] 2.2 Obtaining genomic variations from a single tumor sample

[0183] 2.2.1 Further, the comparison results of the tumor plasma sample or tumor tissue sample obtained in step 1.5 with the hg19 reference genome are used to perform variant detection using the variant detection software RealDcaller2. The detection results are the variant sites in the single tumor sample DNA sequence that are inconsistent with the reference genome sequence, and the variant frequency AF corresponding to each variant site is obtained at the same time.

[0184] 2.2.2 Further, the variant sites obtained in 2.2.1 were annotated using the annotation software SomVAS. The population frequencies of the variants were annotated using public population frequency databases gnomAD, dbSNP, 1000G, ExAC, and the Chinese population frequency database PAD (established by Geneplus through genomic testing of normal blood cells from 89,767 cancer patients). DB ;

[0185] 3. Obtain the copy number variation results for the sample interval.

[0186] 3.1 Following steps 1.1-1.4, sequencing results were obtained using leukocyte DNA from 50 normal individuals. A baseline of normal individuals was constructed using CNVkit software.

[0187] 3.2 Using CNVkit software, with the normal human baseline in 3.1 as a control, the genome depth signal of the alignment results obtained in 1.5 was detected, and the genome GC content was used to correct the depth signal. The cyclic binary segmentation algorithm (CBS) was used to divide the genome into multiple copy number intervals according to the depth signal.

[0188] 4. Obtain the results of the variant site frequency test.

[0189] 4.1 For the copy number intervals defined in 3.2, all variants were detected within each interval. Homozygous and heterozygous groups were identified based on variant frequency and minor allele frequency (MAF). K-means clustering was used to divide variant sites into two classes based on their MAF values. The average MAF values ​​of the two classes were compared; the class with the larger average MAF was heterozygous, and the class with the smaller average MAF was homozygous. Within the heterozygous variant group, variant sites satisfying 0.1 ≤ AF ≤ 0.5 were considered heterozygous. Within the homozygous variant group, variant sites satisfying AF > 0.5 were considered homozygous. When there were too few sites to cluster, variant sites satisfying AF > 0.8 were defined as homozygous, and variant sites satisfying 0.2 ≤ AF ≤ 0.8 were defined as heterozygous.

[0190] Formula 1 for calculating secondary allele frequencies is:

[0191] MAF = min(AF, 1-AF) Formula 1

[0192] 4.2 Based on MAF, the theoretical frequency of the heterozygous class is calculated using Equation 2. Equation 2 is:

[0193]

[0194] Among them, W i MAF refers to the weight of the variant site calculated according to Formula 4. i The frequency of the second allele at the variant site is obtained according to Formula 1, and i refers to the sequence number of the SNP site within the interval.

[0195] 4.3 Based on AF, the theoretical mutation frequency of homozygous classes is calculated using Formula 3. Formula 3 is:

[0196]

[0197] Among them, W i The AF refers to the weight of the variant site calculated according to Formula 4.i The mutation frequency obtained in step S4a refers to i, where i refers to the sequence number of the homozygous site.

[0198] 4.4 The weights W in formulas 2 and 3 mentioned in steps 4.2 and 4.3 i The calculation formula is:

[0199]

[0200] Dep i The depth of the mutation site, ΔMAF i The difference between the MAF of the variant site calculated by Formula 1 in step 4.1 and the quartile of the frequency of the variant site within the copy number interval, where α is a constant term for training and can be a value such as 0.01 or 0.05.

[0201] 4.5 Perform a binomial test for heterozygosity on each variant and obtain the significance p-value (germP) of the binomial test for heterozygosity. het ;

[0202] 4.6 Perform a binomial test for homozygous variants for each variant and obtain the significance p-value (germP) for the homozygous class binomial test. homo ;

[0203] 4.7 For all autosomes, the final germP value is determined according to Formula 5, which is:

[0204] germP = max(germP) homo ,germP het ) Formula 5

[0205] 4.8 Determine the germP value of the sex chromosome based on the input sex: When the sex is female, the variation on the X chromosome is determined according to Formula 5, and the value applied to the calculation is confirmed; when the sex is male, the variation of the homologous segments on the X and Y chromosomes is determined according to Formula 5, and the germP value of the non-homologous segments on the X and Y chromosomes is equal to the germP value. homo ;

[0206] 5. Decision Tree Interpretation and Test Results: Using the constructed decision tree logic, the significance test GermP and the population frequency value Freq are applied. DB The mutation frequency (AF) is used to determine whether the mutation originates from the germline. In the decision tree model, the feature weights are ordered as follows: germP > Freq. DB >AF.

[0207] Based on the method of detecting germline mutations using single-sample tumor DNA, the results of germline and somatic mutations in 2158 clinical tumor plasma samples and 2264 clinical tumor tissue samples are shown in Table 2 below:

[0208] Table 2

[0209]

[0210]

[0211] Note: TG stands for True Germline mutation, which refers to the number of mutations that are identified as germline mutations by the single-sample variation detection method disclosed in this publication, and which are also identified as germline mutations by blood cell alignment with the hg19 reference genome (Table 1).

[0212] FS stands for False Somatic mutation, which refers to the number of mutations that are identified as somatic mutations by the single-sample variant detection method disclosed herein, but are identified as germline mutations by blood cell comparison with the hg19 reference genome (Table 1).

[0213] TS stands for True Somatic mutation, which refers to the number of mutations that are identified as somatic mutations by the single-sample mutation detection method disclosed herein, and are also identified as somatic mutations by tumor-blood cell pairing analysis (Table 1).

[0214] FG stands for False Germline mutation, which refers to the number of mutations that are identified as germline mutations by the single-sample variant detection method disclosed herein, but are identified as somatic mutations by tumor-blood cell pairing analysis (Table 1).

[0215] G_PPA is the embryonic sensitivity, calculated using the formula G_PPA = TG / (TG + FS).

[0216] S_PPA is the system sensitivity, and the calculation formula is S_PPA=TS / (TS+FG).

[0217] Comparing the paired sample variation detection results, the single-sample method of this disclosure achieved a sensitivity of 97.4% for germline variation detection in tissue samples and 97.3% in plasma samples; the sensitivity for systemic variation detection reached 99.0% in tissue samples and 97.9% in plasma samples. These results demonstrate that the detection method of this disclosure possesses high discriminative power for germline variation detection in both single tumor tissue samples and single tumor plasma samples.

[0218] The technical solutions of the present invention are not limited to the specific embodiments described above. Any technical modifications made in accordance with the technical solutions of the present invention fall within the protection scope of the present invention.

Claims

1. A method for distinguishing gene variant types in a sample to be tested, characterized in that, The method includes the following steps: The sequencing data of the sample to be tested is compared with the reference genome to obtain the comparison data; Based on the comparison data, information on one or more variant sites is obtained, along with the variant frequency (AF) and population frequency for each variant site. ; Based on the comparison data, the genome of the sample to be tested is divided into multiple copy number intervals; Cluster one or more variant sites within each copy number interval to obtain homozygous variant site class and heterozygous variant site class; The one or more variant sites are subjected to difference tests with the theoretical germline homozygous allele frequency and the theoretical germline heterozygous allele frequency of their respective copy number intervals to obtain the test results for each variant site. Value; and The variation frequency AF of the one or more variant sites and the population frequency and inspection The input value is used to interpret the model to obtain the type of the one or more variant sites; The theoretical germline heterozygous allele frequencies were obtained according to the following formula 2: Official 2, Among them, W i The weight of the i-th mutation site. The frequency of the second allele at the i-th mutation site; The theoretical homozygous allele frequencies of the germline were obtained according to the following formula 3: Official 3, in, The weight of the i-th mutation site. The mutation frequency of the i-th mutation site; The weight of the i-th mutation site is obtained according to the following formula 4: , The sequencing depth refers to the i-th variant site. Refers to the i-th mutation site The difference between the quartile of the mutation frequency AF of one or more mutation sites within the copy number interval to which the i-th mutation site belongs, For training constant terms; Testing of one of the one or more variant sites The value is obtained through the following steps: The mutation frequency AF of the aforementioned mutation site is compared with the homozygous allele frequency of the theoretical germline. Perform a difference test to obtain the test result. value; The secondary allele frequency (MAF) of the mutation site was compared with the heterozygous allele frequency of the theoretical germline. Perform a difference test to obtain the test result. Value; and The test is obtained according to the following formula 5. value: Official 5.

2. The method according to claim 1, characterized in that, The information on one or more mutation sites includes at least one of single nucleotide variants, insertion / deletion variants, copy number variants, and chromosomal structural variants.

3. The method according to claim 1 or 2, characterized in that, Dividing the genome of the sample to be tested into multiple copy number intervals includes: using a cyclic binary partitioning algorithm to divide the genome of the sample to be tested into multiple copy number intervals.

4. The method according to claim 3, characterized in that, The cyclic binary segmentation algorithm divides the genome of the sample to be tested into multiple copy number intervals based on the sequencing depth obtained from the alignment data.

5. The method according to claim 3, characterized in that, The cyclic binary segmentation algorithm divides regions with similar sequencing depths into a single copy number interval.

6. The method according to claim 1, characterized in that, The frequency of the crowd This is obtained by annotating one or more variant sites using a population frequency database.

7. The method according to claim 1, characterized in that, The step of clustering one or more variant sites within each copy number interval includes: obtaining the minor allele frequency (MAF) based on the variant frequency (AF), and then clustering one or more variant sites within each copy number interval based on the minor allele frequency (MAF).

8. The method according to claim 7, characterized in that, The suballele frequency (MAF) is obtained according to the following formula 1: Official 1.

9. The method according to claim 7, characterized in that, The clustering includes: using the K-means clustering algorithm, dividing multiple variant sites within each copy number interval into two categories based on the minor allele frequency (MAF), comparing the average MAF values ​​of the two categories of variant sites obtained from the clustering, identifying the category with the larger average MAF as the heterozygous variant site category, and identifying the category with the smaller average MAF as the homozygous variant site category.

10. The method according to claim 9, characterized in that, For the aforementioned heterozygous variant site class, The variant site was identified as a heterozygous variant site.

11. The method according to claim 9, characterized in that, In the class of homozygous variant sites, The variant site was determined to be a homozygous variant site.

12. The method according to claim 9, characterized in that, When the number of variant sites within a copy number interval is insufficient for clustering, The variant site was determined to be a homozygous site, and the... The variant site was identified as a heterozygous site.

13. The method according to claim 1, characterized in that, The training constant term Selected from 0.01, 0.02, 0.03, 0.04, and 0.

05.

14. The method according to claim 1, characterized in that, The training constant term It is 0.

05.

15. The method according to claim 1, characterized in that, The difference test is selected from t-test, ANOVA or chi-square test.

16. The method according to claim 1, characterized in that, The difference test is a binary difference test.

17. The method according to claim 1, characterized in that, The difference test is a two-tailed binomial difference test.

18. The method according to claim 1, characterized in that, The method also includes determining the gender of the individual from whom the sample to be tested originates.

19. The method according to claim 1, characterized in that, When the sample to be tested comes from a male individual, and the mutation site is located in a non-homologous region of a sex chromosome, .

20. The method according to any one of claims 1 to 2, 4 to 19, characterized in that, The interpretation model is constructed using a machine learning model.

21. The method according to claim 20, characterized in that, The algorithm of the machine learning model includes at least one of K-nearest neighbors, Naive Bayes classifier, logistic regression, decision tree, random forest, support vector machine, neural network, and AdaBoost.

22. The method according to claim 20, characterized in that, The interpretation model is constructed using a decision tree.

23. The method according to claim 22, characterized in that, The interpretation model is constructed using the sklearn decision tree algorithm.

24. The method according to claim 22, characterized in that, The decision tree model takes as input the mutation frequency AF of one or more mutation sites and the population frequency. and inspection Values ​​are used to construct.

25. The method according to claim 22, characterized in that, In the decision tree model, the weights are ordered as follows: .

26. The method according to any one of claims 1 to 2, 4 to 19, and 21-25, characterized in that, The method is used to determine whether one or more mutation sites belong to germline variation or somatic variation.

27. A device for distinguishing gene variant types in a sample to be tested, characterized in that, The device includes: The alignment data acquisition module is used to align the sequencing data of the sample to be tested with the reference genome to obtain alignment data; The variant site information analysis module is used to obtain information on one or more variant sites based on the comparison data, including the variant frequency (AF) and population frequency for each variant site. ; The copy number interval division module is used to divide the genome of the sample to be tested into multiple copy number intervals based on the alignment data. The clustering module is used to cluster one or more variant sites within each copy number interval to obtain homozygous variant site classes and heterozygous variant site classes. The difference detection module is used to perform difference detection on the one or more variant sites with the theoretical germline homozygous allele frequency and the theoretical germline heterozygous allele frequency of their respective copy number intervals, and obtain the detection result for each variant site. Value; and The variant site type identification module is used to determine the variant frequency (AF) and population frequency of the one or more variant sites. and inspection The input value is used to interpret the model to obtain the type of the one or more variant sites; The device further includes a theoretical germline allele frequency acquisition module, used to acquire the theoretical germline homozygous allele frequency and the theoretical germline heterozygous allele frequency. The theoretical germline heterozygous allele frequencies were obtained according to the following formula 2: Official 2, Among them, W i The weight of the i-th mutation site. The frequency of the second allele at the i-th mutation site; The theoretical homozygous allele frequencies of the germline were obtained according to the following formula 3: Official 3, in, The weight of the i-th mutation site. The mutation frequency of the i-th mutation site; The weight of the i-th mutation site is obtained according to the following formula 4: , The sequencing depth refers to the i-th variant site. Refers to the i-th mutation site The difference between the quartile of the mutation frequency AF of one or more mutation sites within the copy number interval to which the i-th mutation site belongs, For training constant terms; Testing of one of the one or more variant sites The value is obtained through the following steps: The mutation frequency AF of the aforementioned mutation site is compared with the homozygous allele frequency of the theoretical germline. Perform a difference test to obtain the test result. value; The secondary allele frequency (MAF) of the mutation site was compared with the heterozygous allele frequency of the theoretical germline. Perform a difference test to obtain the test result. Value; and The test is obtained according to the following formula 5. value: Official 5.

28. The apparatus according to claim 27, characterized in that, The device also includes a sequencing data acquisition module for acquiring sequencing data of the sample to be tested.

29. The apparatus according to claim 27 or 28, characterized in that, The copy number interval partitioning module uses a cyclic binary partitioning algorithm to divide the genome of the sample to be tested into multiple copy number intervals.

30. The apparatus according to claim 27, characterized in that, The clustering module obtains the minor allele frequency (MAF) based on the mutation frequency (AF), and then clusters one or more mutation sites within each copy number interval based on the minor allele frequency (MAF).

31. The apparatus according to claim 30, characterized in that, The suballele frequency (MAF) is obtained according to the following formula 1: Official 1.

32. The apparatus according to claim 30, characterized in that, The clustering includes: using the K-means clustering algorithm, dividing multiple variant sites within each copy number interval into two categories based on the minor allele frequency (MAF), comparing the average MAF values ​​of the two categories of variant sites obtained from the clustering, identifying the category with the larger average MAF as the heterozygous variant site category, and identifying the category with the smaller average MAF as the homozygous variant site category.

33. The apparatus according to claim 32, characterized in that, For the aforementioned heterozygous variant site class, The variant site was identified as a heterozygous variant site.

34. The apparatus according to claim 32, characterized in that, For the aforementioned homozygous variant site class, The variant site was determined to be a homozygous variant site.

35. The apparatus according to claim 32, characterized in that, When the number of variant sites within a copy number interval is insufficient for clustering, The variant site was determined to be a homozygous site, and the... The variant site was identified as a heterozygous site.

36. The apparatus according to claim 27, characterized in that, The training constant α is selected from 0.01, 0.02, 0.03, 0.04 or 0.

05.

37. The apparatus according to claim 27, characterized in that, The training constant α is 0.

05.

38. The apparatus according to claim 27, characterized in that, The difference detection module uses t-test, ANOVA, and chi-square test to detect differences.

39. The apparatus according to claim 38, characterized in that, The difference test is a binary difference test.

40. The apparatus according to claim 39, characterized in that, The difference test is a two-tailed binomial difference test.

41. The apparatus according to any one of claims 27 to 28, 30 to 40, characterized in that, The device also includes a gender determination module for determining the gender of the individual from whom the sample to be tested originates. When the sex determination module determines that the test sample comes from a male individual based on the sex chromosome of the test sample, and the mutation site is located in a non-homologous region of the sex chromosome, .

42. The apparatus according to claim 41, characterized in that, The device also includes a judgment model construction module for judging whether one or more mutation sites belong to germline mutations or somatic cell mutations.

43. The apparatus according to any one of claims 27 to 28, 30 to 40, and 42, characterized in that, The apparatus is used to perform the method according to any one of claims 1 to 26.

44. A storage medium storing a computer-executable program configured to perform, when run, the method for distinguishing gene variant types in a sample to be tested, as described in any one of claims 1 to 26.

45. An apparatus comprising: Memory, used to store programs; A processor for executing a program stored in the memory, the program being configured to run to perform the method for distinguishing gene variant types in a sample to be tested, as described in any one of claims 1 to 26.