Ami and mds co-test analysis method, applications, systems, devices, and media

CN117198403BActive Publication Date: 2026-09-08GUANGZHOU KINGMED CENTER FOR CLINICAL LABORATORY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202311180192.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-13
Publication Date
2026-09-08
Estimated Expiration
2043-09-13

AI Technical Summary

Technical Problem

[0005]但是,目前的二代测序产品针对AML和MDS的检测,同时存在靶点不全面和靶点过多的缺点

Benefits of technology

[0039] The technical solution provided by this invention has the following advantages and effects: By obtaining the normal genome sequence of a normal person and aligning the normal genome sequence with the genome reference sequence, a second variation frequency table is obtained. Variation frequencies less than 1% in the second variation frequency table are statistically analyzed to obtain the 95th percentile upper limit of the variation frequency of single nucleotide variants and the 95th percentile upper limit of the variation frequency of small fragment insertions or deletions. Combining the data from these two variations, the first 95th percentile upper limit is used as the first preset threshold, and the second 95th percentile upper limit is used as the second preset threshold. That is, sites with a first variation frequency less than the first preset threshold and sites with a second variation frequency less than the second preset threshold are considered false positives and filtered in the bioinformatics workflow to improve the detection accuracy of single nucleotide variants and small fragment insertions or deletions. Then, by calculating the difference between "duplication" and "deletion" at the first and second sites and comparing this difference with a third threshold set by establishing a population baseline, the authenticity of the first and second site variations is determined, thereby improving the sensitivity and accuracy of interpreting single nucleotide tandem repeat variations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117198403B_ABST
    Figure CN117198403B_ABST
Patent Text Reader

Abstract

The application discloses an AML and MDS co-detection analysis method, application, system, device and medium, and technical scheme points are: obtaining the genomic sequencing sequence of the DNA library to be detected, the genomic sequencing sequence includes: AML and MDS related gene combination;The genomic sequencing sequence is compared with the genomic reference sequence, and the first variation frequency table containing the first variation frequency of the single nucleotide variation and the second variation frequency of the small fragment insertion or deletion corresponding to each site of the genomic sequencing sequence is obtained;After filtering the first variation frequency less than the first preset threshold value, the second variation frequency less than the second preset threshold value and the third variation frequency corresponding to the preset site in the first variation frequency table, the filtered first variation frequency table is obtained;The application can realize AML and MDS co-detection, can avoid the problems of large data volume and high cost caused by too large gene combination, can save detection time, and is suitable for large-scale rapid diagnosis of AML and MDS.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention pertains to gene detection technology, specifically relating to a method, application, system, equipment, and medium for co-detection analysis of AML and MDS. Background Technology

[0002] Hematologic malignancies are a highly heterogeneous group of diseases, and their diagnosis and treatment require a comprehensive analysis combining morphology, immunology, genetics, and molecular biology. In recent years, an increasing number of molecular alterations have been discovered. On the one hand, these molecular-level changes have redefined the classification of hematologic malignancies; on the other hand, molecular biology-based classification has also provided a basis for the prognostic grouping of hematologic malignancies, making precision diagnosis and precision treatment possible.

[0003] Acute myeloid leukemia (AML) is a malignant disease of hematopoietic stem cells characterized by clonal expansion of abnormally differentiated myeloid blasts. The consequences of immature myeloid cell proliferation include the accumulation of immature progenitor cells, impaired normal hematopoiesis, leading to severe infections, anemia, and bleeding. Some patients may also develop extramedullary diseases, including central nervous system involvement. Myelodysplastic syndromes (MDS) are also a heterogeneous group of myeloid clonal diseases originating from hematopoietic stem cells, characterized by abnormal hematopoietic cell development, manifested as ineffective hematopoiesis, refractory cytopenia, and a high risk of transformation to acute myeloid leukemia (AML). The probability of MDS transforming into AML is high. It is estimated that 30%–40% of MDS cases eventually transform into AML. Once transformed into AML, treatment becomes much more difficult, and the patient's condition deteriorates rapidly. For both AML and MDS, timely diagnosis and initiation of treatment are crucial, requiring rapid and accurate diagnostic methods.

[0004] Current gene mutation detection technologies mainly include real-time quantitative PCR (qPCR), Sanger sequencing, and next-generation sequencing (NGS). qPCR is a method that uses fluorescent chemicals to detect the total amount of product after each polymerase chain reaction cycle during DNA amplification, thereby quantifying a specific DNA sequence in the sample. It has the advantages of being simple and inexpensive, but it can only detect known sites, and the number of sites and samples that can be detected simultaneously is relatively small. Sanger sequencing starts at a fixed point and terminates randomly at a specific base, with each base fluorescently labeled to generate a series of DNA strands of different lengths ending in A, T, C, and G. The lengths of these differently labeled DNA strands are then detected by electrophoresis to obtain the DNA base sequence. Sanger sequencing has high accuracy and long sequencing fragments, but it also suffers from low throughput and low sensitivity. Next-generation sequencing is a short-read sequencing method that detects a large number of small DNA fragments and then uses specific bioinformatics algorithms to compare the results with a reference gene sequence to identify mutations. As a novel molecular biology technique, NGS has advantages such as high throughput, high sensitivity, and low cost, making it an important tool for exploring the molecular pathogenesis of hematological malignancies and guiding clinical diagnosis and treatment.

[0005] However, current next-generation sequencing products for AML and MDS detection suffer from both incomplete and excessive target coverage. Incomplete target coverage can lead to missed detection of important gene variants, thus failing to accurately guide medication. On the other hand, excessive target coverage requires a massive amount of data to maintain accuracy at the same sequencing depth, increasing sequencing costs and bioinformatics analysis time, hindering widespread adoption and efficiency. Furthermore, an overly broad detection range can generate many unexplained and meaningless variants, which offer no guidance for disease treatment. Summary of the Invention

[0006] The purpose of this invention is to provide a method, application, system, device and medium for co-detection analysis of AML and MDS, which can realize co-detection of AML and MDS, avoid the problems of large data volume and high cost caused by excessive gene combination, and save detection time. It is suitable for large-scale rapid diagnosis of AML and MDS.

[0007] The first aspect of this invention discloses a method for co-detection and analysis of AML and MDS, comprising:

[0008] Obtain the genome sequencing sequence of the DNA library to be tested, wherein the genome sequencing sequence includes: gene combinations related to AML and MDS;

[0009] The genome sequencing sequence is compared with the genome reference sequence to obtain a first variation frequency table containing the first variation frequency of single nucleotide variants and the second variation frequency of small fragment insertions or deletions corresponding to each site of the genome sequencing sequence.

[0010] After filtering the first mutation frequency that is less than the first preset threshold, the second mutation frequency that is less than the second preset threshold, and the third mutation frequency corresponding to the preset site in the first mutation frequency table, a filtered first mutation frequency table is obtained.

[0011] Optionally, the method for determining the first preset threshold and the second preset threshold includes:

[0012] Obtain the normal genome sequence of DNA from a normal person;

[0013] The normal genome sequence was compared with the genome reference sequence to obtain a second variation frequency table containing the variation frequency of single nucleotide variants at each site and the variation frequency of small fragment insertions or deletions.

[0014] The mutation frequencies less than 1% in the second mutation frequency table are statistically analyzed to obtain the first 95th percentile upper limit for single nucleotide variants and the second 95th percentile upper limit for small fragment insertions or deletions.

[0015] The first 95th percentile upper limit is used as the first preset threshold, and the second 95th percentile upper limit is used as the second preset threshold.

[0016] Optionally, the method for selecting the preset site includes:

[0017] Obtain the normal genomic sequences of DNA from multiple healthy individuals;

[0018] Each of the normal genome sequences is compared with the genome reference sequence to obtain the variation frequency of each corresponding site;

[0019] Sites that exhibit variations in more than 20% of normal individuals were selected to obtain a set of false positive variant sites;

[0020] The first and second sites were selected from the set of false positive variant sites, wherein the first site is chr20_31022441_31022441_-_G and the second site is chr20_31022442_31022442_G_-.

[0021] Calculate the difference in variation frequency between the first and second sites in the genome sequencing sequence. If the difference is not less than a third preset threshold, the variation at the first or second site is considered to be real, and all sites in the set of false positive variation sites after removing the first or second site are taken as preset sites. If the difference is less than the third preset threshold, the variation at the first and second sites is considered to be false, and all sites in the set of false positive variation sites are taken as preset sites.

[0022] Optionally, the method for determining the third preset threshold includes:

[0023] The difference in variation frequency between the first and second sites in each of the normal genomic sequences is calculated, and the mean, maximum and upper limit of the third 95th percentile of the first difference are obtained through statistical analysis.

[0024] The difference between the mutation frequencies of the first and second sites in the genome sample sequence is calculated, and the mean, minimum and 5th percentile lower limit of the second difference are obtained by statistics. The genome sample sequence includes a first sample with a mutation at the first site and a second sample without mutation. The mutation frequency of the first site after mixing the first sample and the second sample is within a preset frequency range.

[0025] The third preset threshold is obtained based on the first average value, the maximum value, the third 95th percentile upper limit, the second average value, the minimum value, and the 5th percentile lower limit.

[0026] Optionally, the AML and MDS-related gene combinations include the genes shown in Table 1:

[0027] Table 1 Panels related to AML and MDS

[0028]

[0029]

[0030] RefSeqID (RefSeq Accession Number) refers to the number in the database of biologically non-redundant gene and protein fragment sequences provided by the National Center for Biotechnology Information Technology (NCBI).

[0031] The second aspect of the present invention discloses the application of a gene combination related to AML and MDS in constructing a DNA library for co-detection of AML and MDS, wherein the gene combination includes the genes shown in Table 1.

[0032] A third aspect of this invention discloses the application of a probe combination capable of co-detecting AML and MDS-related gene combinations in constructing DNA libraries for co-detection of AML and MDS, wherein the gene combination includes the genes shown in Table 1:

[0033] The fourth aspect of this invention discloses a co-detection and analysis system for ML and MDS, comprising:

[0034] The sequence acquisition module is used to acquire the genome sequencing sequence of the DNA library to be tested, wherein the genome sequencing sequence includes: gene combinations related to AML and MDS;

[0035] The sequence alignment module is used to align the genome sequencing sequence with the genome reference sequence to obtain a first variation frequency table containing the first variation frequency of single nucleotide variants and the second variation frequency of small fragment insertions or deletions corresponding to each site of the genome sequencing sequence.

[0036] The site filtering module is used to filter the first mutation frequency that is less than a first preset threshold, the second mutation frequency that is less than a second preset threshold, and the third mutation frequency corresponding to a preset site in the first mutation frequency table to obtain the filtered first mutation frequency table.

[0037] The fifth aspect of the present invention discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0038] The sixth aspect of the present invention discloses a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the above-described method.

[0039] The technical solution provided by this invention has the following advantages and effects: By obtaining the normal genome sequence of a normal person and aligning the normal genome sequence with the genome reference sequence, a second variation frequency table is obtained. Variation frequencies less than 1% in the second variation frequency table are statistically analyzed to obtain the 95th percentile upper limit of the variation frequency of single nucleotide variants and the 95th percentile upper limit of the variation frequency of small fragment insertions or deletions. Combining the data from these two variations, the first 95th percentile upper limit is used as the first preset threshold, and the second 95th percentile upper limit is used as the second preset threshold. That is, sites with a first variation frequency less than the first preset threshold and sites with a second variation frequency less than the second preset threshold are considered false positives and filtered in the bioinformatics workflow to improve the detection accuracy of single nucleotide variants and small fragment insertions or deletions. Then, by calculating the difference between "duplication" and "deletion" at the first and second sites and comparing this difference with a third threshold set by establishing a population baseline, the authenticity of the first and second site variations is determined, thereby improving the sensitivity and accuracy of interpreting single nucleotide tandem repeat variations. Attached Figure Description

[0040] Figure 1This is a flowchart illustrating the method provided by the present invention;

[0041] Figure 2 This is a low-frequency variation statistical chart of SNVs in the baseline sample set provided by the present invention;

[0042] Figure 3 This is a low-frequency variation statistical graph of indel in the baseline sample set provided by the present invention;

[0043] Figure 4 This is a statistical chart showing the frequency difference between chr20:g.31022441dup and chr20:g.31022442del in baseline and negative samples provided by this invention.

[0044] Figure 5 This is a statistical chart of the frequency difference between chr20:g.31022441dup and chr20:g.31022442del in the first and second samples provided by this invention;

[0045] Figure 6 This is the SNV precision statistics chart provided by the present invention;

[0046] Figure 7 This is the indel precision statistical chart provided by the present invention;

[0047] Figure 8 This is an operation flowchart of Embodiment 1 provided by the present invention;

[0048] Figure 9 This is an example flowchart of bioinformatics analysis according to Embodiment 1 of the present invention;

[0049] Figure 10 This is a structural block diagram of the AML and MDS co-detection analysis model provided by the present invention;

[0050] Figure 11 This is an internal structural diagram of the computer device in an embodiment of the present invention. Detailed Implementation

[0051] To facilitate understanding of the present invention, specific embodiments of the present invention will be described in more detail below with reference to the accompanying drawings.

[0052] Unless otherwise specified or defined, the terms "first," "second," etc., used in this invention are merely for distinguishing names and do not represent a specific number or order.

[0053] Unless otherwise stated or defined, the term "and / or" as used in this invention includes any and all combinations of one or more of the associated listed items.

[0054] Acute myeloid leukemia (AML) is a malignant disease of hematopoietic stem cells characterized by clonal expansion of abnormally differentiated myeloid blasts. The consequences of immature myeloid cell proliferation include the accumulation of immature progenitor cells, impaired normal hematopoiesis, and severe infections, anemia, and bleeding. Myelodysplastic syndromes (MDS), also a heterogeneous group of myeloid clonal diseases originating from hematopoietic stem cells, are characterized by abnormal hematopoietic cell development, manifesting as ineffective hematopoiesis, refractory cytopenia, and a high risk of transformation to acute myeloid leukemia (AML).

[0055] This invention references the COSMIC database, TCGA database, WHO blood-lymphatic classification, Chinese / European / American clinical practice guidelines / expert consensus, and research progress on AML and MDS from other laboratories and clinical institutions. The selected genes are based on the detailed classification and association of gene variations introduced in the guidelines with existing AML and MDS risk prognostic stratification systems. Furthermore, it utilizes important genes and their variations as molecular biological indicators for accurate prognostic assessment, playing a crucial auxiliary role in the diagnosis, prognostic stratification, and treatment decision support of AML and MDS. Finally, the number of genes to be detected in the co-examined AML and MDS-related gene panel was determined to be 46, including 5 sex loci, with a panel size of 93.8 kb. The panels cover important exon coding regions of the genes, with variation types including point mutations, small insertions / deletions, and large insertions / deletions. The AML and MDS-related panels include the genes shown in Table 1.

[0056] like Figure 1 As shown, this invention provides a method for co-detection and analysis of AML and MDS, comprising:

[0057] Step 1: Obtain the genome sequencing sequence of the DNA library to be tested, wherein the genome sequencing sequence includes: gene combinations related to AML and MDS;

[0058] Specifically, the DNA corresponding to the genome sequencing sequence to be detected is extracted from peripheral blood / bone marrow. After extraction, the DNA undergoes library construction, probe hybridization, library capture and elution, product amplification and purification, and library quality control before sequencing to obtain the genome sequencing sequence of the DNA library to be detected. This invention utilizes the sequencing-by-synthesis technology of the Illumina sequencing platform, enabling massively parallel sequencing of millions of fragments simultaneously through reversible termination of chemical reactions. The Illumina sequencing platform, through continuous updates to sequencing reagents and improvements to its algorithm software system, effectively reduces sequencing errors from homopolymers and repetitive sequences, achieving high base coverage and accuracy.

[0059] Step 2: Align the genome sequencing sequence with the genome reference sequence to obtain a first variant frequency table containing the first variant frequency of single nucleotide variants (SNVs) and the second variant frequency of small fragment insertions or deletions (indels) corresponding to each site in the genome sequencing sequence; specifically, if the genome sequencing sequence can be aligned to the genome reference sequence using STAR software and / or bwa software, in this embodiment, the genome reference sequence is the hg19 reference genome, and the first variant frequency table includes each site in the genome sequencing sequence and its corresponding variant frequency.

[0060] Step 3: Filter the first mutation frequency that is less than the first preset threshold, the second mutation frequency that is less than the second preset threshold, and the third mutation frequency corresponding to the preset site in the first mutation frequency table to obtain the filtered first mutation frequency table.

[0061] In practical applications, since the Illumina sequencing platform relies on PCR (Polymerase Chain Reaction) amplification, the PCR process is prone to introducing base mismatches. Therefore, it is necessary to first remove low-frequency false positive variants caused by base mismatches. This means filtering out the first false positive sites with a first variant frequency less than a first preset threshold and their first variant frequencies, and the second false positive sites with a second variant frequency less than a second preset threshold and their second variant frequencies. Then, false positive sites generated in special DNA sequences, such as sequencing errors caused by high GC fragments, are removed by filtering out preset sites. After filtering out the first false positive sites and their corresponding variant frequencies, the second false positive sites and their corresponding variant frequencies, and the preset sites and their corresponding variant frequencies, a filtered table of first variant frequencies is obtained. This invention screens panels related to AML and MDS for detection, achieving co-detection of both diseases while ensuring that the panels are not too large, thus balancing detection quality and cost. In addition, after filtering, it does not generate a large amount of data that would increase the difficulty of analysis reports. At the same time, the variant sites output by the bioinformatics workflow (that is, the sites included in the first variant frequency table after filtering) can encompass both important and real variant sites, ensuring the accuracy of the analysis. In other words, using the AML and MDS co-detection analysis method in this invention can improve the accuracy of AML and MDS co-detection.

[0062] Furthermore, the method for determining the first preset threshold and the second preset threshold includes:

[0063] Obtain the normal genome sequence of DNA from a normal person;

[0064] The normal genome sequence was compared with the genome reference sequence to obtain a second variation frequency table containing the variation frequency of single nucleotide variants at each site and the variation frequency of small fragment insertions or deletions.

[0065] The mutation frequencies less than 1% in the second mutation frequency table are statistically analyzed to obtain the first 95th percentile upper limit for single nucleotide variants and the second 95th percentile upper limit for small fragment insertions or deletions.

[0066] The first 95th percentile upper limit is used as the first preset threshold, and the second 95th percentile upper limit is used as the second preset threshold.

[0067] In practical applications, bone marrow samples from 30 healthy individuals were selected to establish a normal population SNV / indel database, which was used to filter subsequent patient test results. Each bone marrow sample was sequenced using the Illumina sequencing platform to obtain the corresponding normal genomic sequence, forming a baseline sample. Since the baseline sample originated from healthy individuals, it should theoretically not contain low-frequency variations. Therefore, after aligning the normal genomic sequence with the genomic reference sequence, a second variation frequency table was obtained. Variation frequencies less than 1% in the second variation frequency table were statistically analyzed, such as... Figure 2 and Figure 3 As shown, the 95th percentile upper limit of the SNV mutation frequency is 0.87%, and the 95th percentile upper limit of the indel mutation frequency is 0.93%. Combining the data of these two types of mutations, the first preset threshold is set to 0.87%, and the second preset threshold is set to 0.93%. That is, sites with a first mutation frequency of less than 0.87% (i.e., the first false positive site) and sites with a second mutation frequency of less than 0.93% (i.e., the second false positive site) are considered false positives. In the bioinformatics workflow, these sites are filtered out to identify false positives caused by base mismatches, thereby improving the detection accuracy of SNV and indel and thus improving the accuracy of AML and MDS co-detection analysis. The 95th percentile upper limit means that in a sample set consisting of a set of data, if the number of samples with a value less than a certain value accounts for 95% of the entire sample set, then the value of that sample is the 95th percentile upper limit corresponding to the 95th percentile.

[0068] Furthermore, the method for selecting the preset site includes:

[0069] Obtain the normal genomic sequences of DNA from multiple healthy individuals;

[0070] Each of the normal genome sequences is compared with the genome reference sequence to obtain the variation frequency of each corresponding site;

[0071] Sites that exhibit variations in more than 20% of normal individuals were selected to obtain a set of false positive variant sites;

[0072] The first and second sites were selected from the set of false positive variant sites, wherein the first site is chr20_31022441_31022441_-_G and the second site is chr20_31022442_31022442_G_-.

[0073] Calculate the difference in variation frequency between the first and second sites in the genome sequencing sequence. If the difference is not less than a third preset threshold, the variation at the first or second site is considered to be real, and all sites in the set of false positive variation sites after removing the first or second site are taken as preset sites. If the difference is less than the third preset threshold, the variation at the first and second sites is considered to be false, and all sites in the set of false positive variation sites are taken as preset sites.

[0074] In practical applications, besides false positives caused by base mismatches, specific DNA sequences can also produce false positives, such as sequencing errors caused by high GC fragments. Such sequencing errors are often widespread and recur in multiple samples. Therefore, after filtering out the first and second false positive sites, the frequency of variants in the baseline sample set is statistically analyzed. Variants appearing in more than 20% of the normal genome sequence are considered non-specific false positive variants. All non-specific false positive variant sites are screened out to obtain the false positive variant site set, including the variant sites shown in Table 2. These variant sites are compared with the hg19 reference genome.

[0075] Table 2 False Positive Variants

[0076]

[0077]

[0078]

[0079] Here, ref_batch represents the number of samples selected from 30 samples that produced mutations at the corresponding sites. For example, [15|30] in the second row and second column of Table 2 indicates that 15 out of 30 samples tested positive for the mutation at the chr10_112337259_112337259_-_T site. Then, dividing 15 by 30 gives a 50% proportion. In the set of false positive mutation sites, there are two clinically significant sites (i.e., the first site and the second site). Filtering the first site and the second site will lead to missed detection of important sites. If the first site and the second site are output, the number of low-frequency false positive sites will increase the difficulty of interpretation. For single nucleotide tandem repeat regions like chr20:g.31022441dup (i.e., the first site mutation) and chr20:g.31022442del (i.e., the second site mutation), which have a high mutation frequency (>1%) in all samples, dup refers to the duplication of DNA sequences in certain segments of the chromosome, resulting in rearrangement of coding gene sequences due to duplication of some structural gene sequences; del refers to the loss of DNA sequences in certain segments of the chromosome, resulting in rearrangement of coding gene sequences due to loss of some structural gene sequences. If the third preset threshold is too low, it will cause more false positives, while if the third preset threshold is too high, it will reduce the sensitivity of the site. Therefore, by calculating the difference between "duplication" and "deletion" to determine the third preset threshold, the authenticity of the mutation is judged, which improves the sensitivity and accuracy of the interpretation of single nucleotide tandem repeat region mutations. If the difference between the mutation frequency of the first site and the mutation frequency of the second site is greater than the third preset threshold, the mutation of the first site is considered to be true. If the difference between the mutation frequency of the second site and the mutation frequency of the first site is greater than the third preset threshold, the mutation of the second site is considered to be true. If the mutation of the first site or the second site is considered to be true, all sites in the set of false positive mutation sites after removing the first site or the second site are taken as preset sites. If the mutation of the first site and the second site is considered to be false, all sites in the set of false positive mutation sites are taken as preset sites.

[0080] Furthermore, the method for determining the third preset threshold includes:

[0081] The difference in variation frequency between the first and second sites in each of the normal genomic sequences is calculated, and the mean, maximum and upper limit of the third 95th percentile of the first difference are obtained through statistical analysis.

[0082] The difference between the mutation frequencies of the first and second sites in the genome sample sequence is calculated, and the mean, minimum and 5th percentile lower limit of the second difference are obtained by statistics. The genome sample sequence includes a first sample with a mutation at the first site and a second sample without mutation. The mutation frequency of the first site after mixing the first sample and the second sample is within a preset frequency range.

[0083] The third preset threshold is obtained based on the first average value, the maximum value, the third 95th percentile upper limit, the second average value, the minimum value, and the 5th percentile lower limit.

[0084] In practical applications, by examining the first and second loci using IGV (Integrative Genomics Viewer), it was found that the first and second loci are located in the same region, which is an 8-guanine nucleotide (G) tandem repeat region. 31022441_31022441_-_G indicates that the sequencing result for this region is 9 G tandem repeats, while 31022442_31022442_G_- indicates that the sequencing result for this region is 7 G tandem repeats. Since the baseline samples are from healthy individuals, it can be assumed that all variants at the first and second loci in the baseline samples are spurious. Negative clinical samples where ASXL1 variants at the first and second loci have been verified not to be detected using other methods can also be considered spurious. The ASXL1 gene is located on chromosome 20q11. If the variants originate from random errors in the Illumina sequencing platform, then the probabilities of deletion and duplication are similar, so the frequency difference between these two variants in the baseline samples and negative clinical samples should be within a certain range. Therefore, the frequency difference between the first and second locus variants in the baseline sample and the negative clinical sample was statistically analyzed, yielding a mean difference of 0.84%, a maximum of 2.48%, and a 95th percentile upper limit of 2.08%. Figure 4 As shown.

[0085] The first sample containing the chr20:g.31022441dup variant was diluted with the second sample without the variant, so that the variant frequency of chr20:g.31022441dup was between 5% and 10% (i.e., the preset frequency range), and the difference between the first and second site variants was calculated. Figure 5 As shown, the mean difference was 6.46%, the minimum was 3.08%, and the lower limit of the 5th percentile was 3.08%. Figure 4 and Figure 5Based on the statistical results, the frequency difference is preferably set as the third preset threshold of 3%. The 5th percentile lower limit means that if the number of samples with a value less than this value accounts for 5% of the entire sample set, then the value of this sample is the 5th percentile lower limit corresponding to the 5th percentile.

[0086] In this invention, to verify the accuracy of the co-detection analysis method for AML and MDS, bone marrow samples from normal individuals (30 cases), bone marrow DNA standards (1 case), standards from the general population (1 case), and clinical samples tested by other methods (18 cases) were analyzed. The test results were compared with the expected results for the corresponding samples to explore the accuracy of the detection method. The test data for each sample were compared with the expected results after step 3 to perform accuracy statistics, and the statistical data are shown in Tables 3 and 4.

[0087] Table 3 SNV Accuracy Analysis

[0088]

[0089] The SNV statistics are as follows:

[0090] Positive predictive value (PPV) = 134 / (134+4)*100% = 97.10%;

[0091] Positive compliance rate (PPA) = 134 / (134+2)*100% = 98.53%;

[0092] Negative predictive value NPV = 2853 / (2853+2)*100% = 99.93%;

[0093] The negative compliance rate (NPA) = 2853 / (2853+4)*100% = 99.86%;

[0094] The overall compliance rate P = (134 + 2853) / (134 + 4 + 2853 + 2) * 100% = 99.80%;

[0095] Table 4. Indel Accuracy Analysis

[0096]

[0097]

[0098] The indel statistics are as follows:

[0099] Positive predictive value (PPV) = 86 / (86+0)*100% = 100.00%;

[0100] Positive compliance rate (PPA) = 86 / (86+0)*100% = 100.00%;

[0101] Negative predictive value NPV = 2250 / (2250+0)*100% = 100.00%;

[0102] Negative compliance rate (NPA) = 2250 / (2250 + 0) * 100% = 100.00%;

[0103] Overall compliance rate P = (86 + 2250) / (86 + 0 + 2250 + 0) * 100% = 100.00%;

[0104] The accuracy of this AML and MDS co-detection analysis method meets the requirements, demonstrating the rationality of the experimental and bioinformatics analysis procedures. To verify the detection precision of the AML and MDS co-detection analysis method, this invention also set up inter-batch and intra-batch replicates for the same sample. Statistical analysis of the locus variation frequencies obtained from different replicates revealed that the detected values ​​and theoretical values ​​of the variation frequencies exhibit a linear distribution, R0. 2 >0.95, meets the testing requirements, such as Figure 6 and Figure 7 As shown.

[0105] The present invention also provides the application of a gene combination related to AML and MDS in constructing a DNA library for co-detection of AML and MDS, wherein the gene combination includes the genes shown in Table 1.

[0106] In practical applications, when using AML and MDS-related panels to construct DNA libraries for co-detection of AML and MDS, probes that can specifically detect the exon regions or hotspot regions of genes contained in the AML and MDS-related panels can be designed, and the probes can be used to construct DNA libraries for co-detection of AML and MDS.

[0107] The present invention also provides the application of a probe combination capable of co-detecting AML and MDS-related gene combinations in constructing a DNA library for co-detection of AML and MDS, wherein the gene combination includes the genes shown in Table 1.

[0108] The technical solution of the present invention will be further described in detail below through embodiments and in conjunction with the accompanying drawings.

[0109] Example 1

[0110] The target fragment was enriched using a probe capture method. The probes were of conventional design, and next-generation sequencing was performed using the Illumina platform. Figure 8 As shown, the specific steps include:

[0111] (1) DNA extraction

[0112] DNA extracted from peripheral blood / bone marrow.

[0113] (2) Library Construction

[0114] The construction stage of the AML and MDS co-detection analysis library can utilize commercially available universal DNA library construction kits, including but not limited to the following: Qiagen's QIAseq FX DNA Library Kit, Roche's KAPA Hyperplus Kits, and Guangzhou Qikai's Modular DNA Library Kit. The following example uses the QIAseq FX DNA Library Kit:

[0115] a. Enzyme digestion, end repair, and A base addition: The DNA was digested into DNA fragments with a major band of 200–300 bp. End repair and A base addition were then performed on these DNA fragments. The reaction was carried out according to the reaction system in Table 5 and the following reaction procedure: 4℃, 1 min; 32℃, 18 min; 65℃, 30 min; 4℃, ∞.

[0116] Table 5. Reaction system for enzyme digestion, end repair, and A base.

[0117] 10X Fragmentation Buffer 2.5 5X WGS Fragmentation Mix 5 Enhancer Buffer 1 DNA (100 ng) from step (1) 16.5

[0118] b. Connect and purify the reaction product from step a above, and carry out the reaction according to the reaction system in Table 6 and the following reaction procedure: 20℃, 15min; 4℃, ∞.

[0119] Table 6 Connector Connection and Purification Reaction System

[0120] PCR-grade water 8 5X Rapid Ligation Buffer 10 WGS ligase 5 The reaction products of step a 25 Specific connector 2

[0121] c. Library fragment screening: Add 35 μL (0.7*) Agencourt AMPure XP beads to 50 μL of library, mix well, transfer 80 μL of supernatant to a new 96-well plate, add 10 μL (0.9*) Agencourt AMPure XP beads, mix well and discard the supernatant; wash the magnetic beads twice with 80% ethanol, add 22 μL of esuspension buffer to resuspend the magnetic beads, let stand at room temperature for 2 minutes, centrifuge slightly and place on a magnetic rack for 2 minutes, transfer 20 μL of supernatant for the next PCR reaction.

[0122] d. Conduct the reaction according to the reaction system in Table 7 and the following reaction procedure: 98℃, 2 min; [98℃, 20 s; 60℃, 30 s; 72℃, 30 s] 8 cycles; 10℃, ∞. The product can be purified using magnetic beads to obtain an AML and MDS co-detection analytical library.

[0123] Table 7 Amplification System of Purified Products

[0124] PCR Master Mix 25 Connector primers 1.5 The product purified in step c 23.5

[0125] (3) Probe hybridization

[0126] a. Due to the small target area for detection, the original 8-sample hybridization was adjusted to 16-sample hybridization. 100ng of each sample was taken and mixed in the same PCR tube, for a total of 1600ng, for hybridization.

[0127] b. Take a PCR tube and mix the components according to Table 8.

[0128] Table 8 Library hybridization system

[0129] Blocker Solution 6.0 Hybrid Library <50 Universal Blockers 9.0 probe 4.0 Total <69

[0130] c. Place the mixed components into a vacuum filtration system (60°C) and dry them into a dry powder.

[0131] d. Add the components listed in Table 9 to the dried PCR tubes, incubate at room temperature for 5 min, and then place them on a PCR instrument to run the denaturation hybridization program: [95℃, 10 min]; [60℃, ∞].

[0132] Table 9 Library hybridization systems

[0133] Hybridization Buffer 20 Hybridization Buffer Enhancer 30 Total 50.0

[0134] e. Hybridize the probe with the library overnight (approximately 12-18 hours).

[0135] (4) Library capture and elution

[0136] a. Take 100 μl of Capture Beads and wash twice, using 200 μl of 1X Bead Wash Buffer each time. Add the hybridization sample to the Capture Beads, mix well, and react at 65°C for 45 min.

[0137] b. After capture, add 100 μL of preheated 1X Wash Buffer I to the magnetic beads, mix well, and discard the supernatant.

[0138] c. Add 200 μL of preheated 1X Stringent Wash Buffer, vortex to mix, incubate at 65°C for 5 min, then discard the supernatant. Repeat this step twice.

[0139] d. Clean the magnetic beads with 200 μl of 1X Wash Buffer I, 200 μl of 1X Wash Buffer II, and 200 μl of 1X Wash Buffer III respectively.

[0140] e. Add 20 μl ddH2O to resuspend the Capture Beads.

[0141] (5) Product amplification and purification

[0142] Add the Mix according to Table 10, mix well, and then place on a PCR instrument for reaction. The reaction program is as follows: [98℃, 45s]; [98℃, 15s; 65℃, 30s; 72℃, 30s; 8 cycles]; [72℃, 1min]; [10℃, ∞]. Purify the amplification product with magnetic beads.

[0143] Table 10 Product amplification system

[0144] Previous step: DNA with magnetic beads 20 KAPA HiFi HotStart ReadyMix 25 10×Primer premix(for illumina) 5 Total 50

[0145] (6) Document Quality Control

[0146] Library concentration was determined using Qubit, ranging from 20 to 50 ng / μL; library fragment size was determined using an Agilent 2100 bioanalyzer (DNA 1000 Kit), with the main fragment peak at 300-500 bp.

[0147] (7) Sequencing

[0148] Sequencing can be performed using different models of Illumina platform sequencers, including but not limited to Miniseq / Nextseq / Miseq / Hiseq / Novaseq, with a sequencing length of 2*150bp.

[0149] (8) Bioinformatics Analysis

[0150] like Figure 9As shown, the genome sequencing sequences (i.e., raw sequencing data) obtained from sequencing are converted into FASTQ files. The genome sequencing sequences include gene combinations related to AML and MDS, which include the genes shown in Table 1. Using BWA software, the genome sequencing sequences in the FASTQ files are aligned to the genome reference sequences to obtain rawbam files. The genome sequencing sequences in the rawbam files are then sorted, and duplicates caused by PCR amplification are removed, resulting in rmdupbam files. Subsequently, the rmdupbam files are analyzed for target region coverage, sequencing depth, and sequencing uniformity. Snv callable and indel variant sites can be detected, and these sites are annotated. A first variant frequency table is obtained, showing the first variant frequency of single nucleotide variants and the second variant frequency of small fragment insertions or deletions corresponding to each site in the genome sequencing sequence. The first variant frequencies below a first preset threshold, the second variant frequencies below a second preset threshold, and the third variant frequencies corresponding to preset sites in the first variant frequency table are filtered to obtain a filtered first variant frequency table, which is then output.

[0151] Furthermore, the method for determining the first preset threshold and the second preset threshold includes:

[0152] Obtain the normal genome sequence of DNA from a normal person;

[0153] The normal genome sequence was compared with the genome reference sequence to obtain a second variation frequency table containing the variation frequency of single nucleotide variants at each site and the variation frequency of small fragment insertions or deletions.

[0154] The mutation frequencies less than 1% in the second mutation frequency table are statistically analyzed to obtain the first 95th percentile upper limit for single nucleotide variants and the second 95th percentile upper limit for small fragment insertions or deletions.

[0155] The first 95th percentile upper limit is used as the first preset threshold, and the second 95th percentile upper limit is used as the second preset threshold.

[0156] Furthermore, the method for selecting the preset site includes:

[0157] Obtain the normal genomic sequences of DNA from multiple healthy individuals;

[0158] Each of the normal genome sequences is compared with the genome reference sequence to obtain the variation frequency of each corresponding site;

[0159] Sites that exhibit variations in more than 20% of normal individuals were selected to obtain a set of false positive variant sites;

[0160] The first and second sites were selected from the set of false positive variant sites, wherein the first site is chr20_31022441_31022441_-_G and the second site is chr20_31022442_31022442_G_-.

[0161] Calculate the difference in variation frequency between the first and second sites in the genome sequencing sequence. If the difference is not less than a third preset threshold, the variation at the first or second site is considered to be real, and all sites in the set of false positive variation sites after removing the first or second site are taken as preset sites. If the difference is less than the third preset threshold, the variation at the first and second sites is considered to be false, and all sites in the set of false positive variation sites are taken as preset sites.

[0162] Furthermore, the method for determining the third preset threshold includes:

[0163] The difference in variation frequency between the first and second sites in each of the normal genomic sequences is calculated, and the mean, maximum and upper limit of the third 95th percentile of the first difference are obtained through statistical analysis.

[0164] The difference between the mutation frequencies of the first and second sites in the genome sample sequence is calculated, and the mean, minimum and 5th percentile lower limit of the second difference are obtained by statistics. The genome sample sequence includes a first sample with a mutation at the first site and a second sample without mutation. The mutation frequency of the first site after mixing the first sample and the second sample is within a preset frequency range.

[0165] The third preset threshold is obtained based on the first average value, the maximum value, the third 95th percentile upper limit, the second average value, the minimum value, and the 5th percentile lower limit.

[0166] like Figure 10 As shown, the present invention also provides a co-detection and analysis system for AML and MDS, comprising:

[0167] The sequence acquisition module 10 is used to acquire the genome sequencing sequence of the DNA library to be tested, wherein the genome sequencing sequence includes: gene combinations related to AML and MDS;

[0168] The sequence alignment module 20 is used to align the genome sequencing sequence with the genome reference sequence to obtain a first variation frequency table containing the first variation frequency of single nucleotide variants and the second variation frequency of small fragment insertions or deletions corresponding to each site of the genome sequencing sequence.

[0169] The site filtering module 30 is used to filter the first variation frequency that is less than a first preset threshold, the second variation frequency that is less than a second preset threshold, and the third variation frequency corresponding to a preset site in the first variation frequency table to obtain the filtered first variation frequency table.

[0170] For a detailed description of the AML and MDS co-inspection and analysis system, please refer to the section on the composition of the AML and MDS co-inspection and analysis method above; it will not be repeated here. Each module of the aforementioned AML and MDS co-inspection and analysis system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0171] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 11 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements an AML and MDS co-detection analysis method.

[0172] Those skilled in the art will understand that Figure 11 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0173] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the AML and MDS co-detection analysis methods described in the above embodiments.

[0174] This invention also discloses a computer-readable storage medium storing a computer program that causes a computer to execute the AML and MDS co-detection analysis methods described in the above embodiments.

[0175] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0176] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A method for co-detection and analysis of AML and MDS, characterized in that, include: Obtain the genome sequencing sequence of the DNA library to be tested, wherein the genome sequencing sequence includes: gene combinations related to AML and MDS; The genome sequencing sequence is compared with the genome reference sequence to obtain a first variation frequency table containing the first variation frequency of single nucleotide variants and the second variation frequency of small fragment insertions or deletions corresponding to each site of the genome sequencing sequence. After filtering the first mutation frequency that is less than the first preset threshold, the second mutation frequency that is less than the second preset threshold, and the third mutation frequency corresponding to the preset site in the first mutation frequency table, the filtered first mutation frequency table is obtained. The method for determining the first preset threshold and the second preset threshold includes: Obtain the normal genome sequence of DNA from a normal person; The normal genome sequence was compared with the genome reference sequence to obtain a second variation frequency table containing the variation frequency of single nucleotide variants at each site and the variation frequency of small fragment insertions or deletions. The mutation frequencies less than 1% in the second mutation frequency table are statistically analyzed to obtain the first 95th percentile upper limit for single nucleotide variants and the second 95th percentile upper limit for small fragment insertions or deletions. The first 95th percentile upper limit is used as the first preset threshold, and the second 95th percentile upper limit is used as the second preset threshold. The method for selecting the preset site includes: Obtain the normal genomic sequences of DNA from multiple healthy individuals; Each of the normal genome sequences is compared with the genome reference sequence to obtain the variation frequency of each corresponding site; Sites that exhibit mutations in more than 20% of normal individuals are selected to obtain a set of false positive variant sites; The first and second sites were selected from the set of false positive variant sites, wherein the first site is chr20_31022441_31022441_-_G and the second site is chr20_31022442_31022442_G_-. Calculate the difference in variation frequency between the first and second sites in the genome sequencing sequence. If the difference is not less than a third preset threshold, the variation at the first or second site is considered to be real, and all sites in the set of false positive variation sites after removing the first or second site are taken as preset sites. If the difference is less than the third preset threshold, the variation at the first and second sites is considered to be false, and all sites in the set of false positive variation sites are taken as preset sites. The method for determining the third preset threshold includes: The difference in variation frequency between the first and second sites in each of the normal genomic sequences is calculated, and the mean, maximum and upper limit of the third 95th percentile of the first difference are obtained through statistical analysis. The difference between the mutation frequencies of the first and second sites in the genome sample sequence is calculated, and the mean, minimum and 5th percentile lower limit of the second difference are obtained by statistics. The genome sample sequence includes a first sample with a mutation at the first site and a second sample without mutation. The mutation frequency of the first site after mixing the first sample and the second sample is within a preset frequency range. The third preset threshold is obtained by rounding down to the 5th percentile lower limit.

2. The method for co-detection and analysis of AML and MDS according to claim 1, characterized in that, The gene combinations associated with AML and MDS include the genes shown in the table below: 。 3. An AML and MDS co-detection analysis system, used to perform the AML and MDS co-detection analysis method as described in claim 1 or 2, characterized in that, include: The sequence acquisition module is used to acquire the genome sequencing sequence of the DNA library to be tested, wherein the genome sequencing sequence includes: gene combinations related to AML and MDS; The sequence alignment module is used to align the genome sequencing sequence with the genome reference sequence to obtain a first variation frequency table containing the first variation frequency of single nucleotide variants and the second variation frequency of small fragment insertions or deletions corresponding to each site of the genome sequencing sequence. The site filtering module is used to filter the first mutation frequency that is less than a first preset threshold, the second mutation frequency that is less than a second preset threshold, and the third mutation frequency corresponding to a preset site in the first mutation frequency table to obtain the filtered first mutation frequency table.

4. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method of claim 1 or 2.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 1 or 2.

Citation Information

Patent Citations

  • Detection kit for detecting AML related gene group

    CN106381332A

  • High-flux sequencing data analysis method and device

    CN109767810A

  • Method for detecting somatic cell mutation of paraffin section samples based on next-generation sequencing and device thereof

    CN110729025A

  • Targeted high-throughput sequencing MDS detection kit based on multiple PCR and preparation method

    CN111304331A

  • Method for constructing myelodysplastic syndrome progression gene prediction model

    CN113764044A