A monitoring model for minimal residual disease in human hematological malignancies
A simplified MRD detection method for blood cancers uses a single PCR step and advanced clustering to enhance sensitivity and accuracy in monitoring residual cancer cells, addressing the limitations of existing methods by improving library construction and tracking low-abundance ctDNA.
Patent Information
- Application Number
- CN202510412727.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing MRD detection methods have limitations in terms of sensitivity, library construction complexity, sample demand and cost, and it is difficult to effectively monitor the tiny residual lesions of hematologic tumors, especially the low abundance of ctDNA.
The one-step library building technology is used to synchronize the detection of immune receptor rearrangement and key fusion gene events through a single pool amplification strategy, and combine high-specific primers and significant cloning screening algorithms to optimize the calculation methods of clonotype and fusion type to achieve high sensitivity monitoring of IGH, IGK, IGL and IGH-BCL fusion.
High sensitivity monitoring of tiny residual lesions of blood tumors has been achieved, and the detection performance reaches 10-7 levels, reducing the library construction time and sample demand, improving the detection accuracy and comprehensiveness, and providing more in-depth tumor biological information.
Smart Images

Figure CN119943149B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and particularly to a model for monitoring minimal residual disease of human hematological malignancies. Background Art
[0002] Hematological malignancies are common malignant tumors, with high incidence and mortality rates, seriously threatening human health. With the advent of the era of individualized precision medicine for hematological malignancies, the detection of minimal residual disease (MRD) in hematological malignancies has played an increasingly important role in the precise monitoring of tumor cells or clones, early intervention, and evaluation of immune reconstitution.
[0003] MRD (minimal residual disease) refers to a small number of cancer cells remaining in the body after cancer treatment, which may lead to disease recurrence. As a routine detection method for acute lymphoblastic leukemia (ALL), MRD is currently the most important prognostic indicator and can also be used as a biomarker for monitoring chemotherapy efficacy and drug resistance. Existing MRD detection methods include flow cytometry (MFC), PCR detection, and next-generation sequencing (NGS). MFC detection tracks abnormal immunophenotypes through multi-color flow cytometry, with a short turnaround time and wide applicability. However, its sensitivity is unstable, relying on professional interpretation, and the antigen phenotype may change after chemotherapy, limiting its early prediction ability. NGS detection is an emerging high-sensitivity method that uses consensus primers to amplify immunoglobulin and T cell receptor gene rearrangement fragments, without the need for patient-specific customization, can detect clonal evolution, and track MRD after treatment. Through efficient bioinformatics analysis, NGS can accurately identify specific clonal sequences, with a sensitivity of 10 -5 ~10 -6 . However, NGS technology also has some challenges and limitations, including the complexity of the library construction process, the limitation of ctDNA sensitivity, tumor spatio-temporal heterogeneity, limited independent prognostic value of IGK and IGL-MRD, and high costs. Traditional NGS techniques for MRD detection usually employ multi-step amplification and pooling-based library construction, which also have some defects in practical applications, such as cumbersome library construction procedures, the need for two rounds of PCR amplification, large sample requirements, easy sample cross-contamination, and high library construction costs. Summary of the Invention
[0004] To solve the above problems, the present invention provides a model for monitoring minimal residual disease of human hematological malignancies.
[0005] To achieve the above object, the technical solutions adopted by the present invention are as follows:
[0006] The present invention provides a minimal residual disease (MRD) monitoring model for human hematological malignancies, comprising: a data acquisition module for receiving DNA sequencing data before and after treatment of the same detection subject; a nucleated cell count acquisition module for obtaining the nucleated cell count by aligning the sequencing data of the reference gene to the reference genome; a clonal map and sequence acquisition module for aligning the DNA sequencing data to an immune database including IGH, IGK, and IGL genes through an alignment algorithm, and then obtaining the clonal map and the corresponding sequences through a clustering algorithm; a fusion genotype and sequence acquisition module for aligning the DNA sequencing data to the reference gene sequence to obtain the fusion genotype and the corresponding sequences; a calculation module for calculating the clonal frequency, clonal cell content, fusion frequency, and fusion cell content based on the nucleated cell count, the clonal map and the corresponding sequences, the fusion genotype and the corresponding sequences, and determining significant clones and significant fusions based on the clonal frequency, clonal cell content, fusion frequency, and fusion cell content; an MRD calculation and determination module for screening effective clones and fusions based on the clones and fusions obtained before and after treatment of the same subject, calculating the MRD value based on the clonal cell content and fusion cell content corresponding to the effective clones and fusions, and determining the residual status of human hematological malignancies based on the MRD value.
[0007] Further, the clustering algorithm includes: calculating the edit distance between sequences based on the DNA sequencing data; performing clustering analysis on the sequenced sequences based on the edit distance and its relationship with the V, D, J regions of the IGH gene, the V, J regions of the IGK gene, and the V, J regions of the IGL gene to obtain the clonal map.
[0008] Further, the calculation method of the edit distance is:
[0009] ;
[0010] ;
[0011] ;
[0012] Among them, seq1 and seq2 represent two sequenced sequences. i represents the i-th base in seq1, and j represents the j-th base in seq2. Both i and j start from 1 and go to the last position, with initial values of 0. ms represents the matching score. When the i-th base of seq1 is the same as the j-th base of seq2, the value of m(i, j) is positive. ss represents the mismatch score. When the i-th base of seq1 is different from the j-th base of seq2, the value of m(i, j) is negative. os represents the score for the start of a gap. Gap opening represents the starting situation of a gap in sequence matching. es represents the score for the extension of a gap. Gap extending represents the continuation situation of a gap in sequence matching. S represents the clip score. When the bases of seq1 and seq2 do not match highly and the negative value is too high, truncation will be selected at this time.
[0013] Further, the cluster analysis includes: giving the best matching method according to the alignment result and the edit distance, clustering the sequencing sequences with the same V regions and J regions of IGH, IGK, and IGL genes. When IGH-DJ rearrangement occurs, clustering the sequencing sequences with the same D regions and J regions of the IGH gene. When IGK-KDE rearrangement occurs, clustering the sequencing sequences with the same V regions of the KDE gene and the IGK gene; in each VJ class, dividing according to the CDR3 region consistency; in each DJ class, dividing according to the 50bp-length region consistency centered on the junction position of the D region and J region of the IGH gene; in each KDE-V class, dividing according to the 50bp-length region consistency centered on the junction position of the V region of the IGK gene and the KDE gene; obtaining the number of sequencing sequences corresponding to each subclass; using the hierarchical clustering method to cluster each subclass. In each cluster, the sequences of the child nodes and the parent nodes are highly similar, and the parent node has a significantly higher number of sequencing sequences than the child nodes; summing the number of sequencing sequences of all nodes in each cluster and reporting the sequence of the root node as a single clone type; combining each reported single clone type to form a clone map.
[0014] Further, the calculation methods of the clone frequency and the clone cell content are as follows:
[0015] ;
[0016] ;
[0017] ;
[0018] ;
[0019] Among them, cfl represents the clonal frequency of a single IGL clone type, count1 is the number of sequences of a single IGL clone type, count IGL is the number of sequences of all IGL clone types, cfh represents the clonal frequency of a single IGH clone type, count2 is the number of sequences of a single IGH clone type, count IGH is the number of sequences of all IGH clone types except IGH-DJ clone types, count HDJ is the number of sequences of all IGH-DJ clone types; cfk represents the clonal frequency of a single IGK clone type, count3 is the number of sequences of a single IGK clone type, count IGK is the number of sequences of all IGK clone types except IGK-KDE clone types, count KDE is the number of sequences of all IGK-KDE clone types; the number of clonal cells is the number of sequences of a single clone type.
[0020] Furthermore, the calculation methods for the fusion frequency and the content of fusion cells are as follows:
[0021] ;
[0022] ;
[0023] Among them, fc is the number of sequences of a single fusion type, hc is the total number of IGH sequences, bc is the total number of BCL1 or BCL2 sequences, and the number of fusion cells is the number of sequences of a single fusion type.
[0024] Furthermore, the screening criteria for significant clone types are as follows: (1) clonal frequency ≥ 0.03; (2) clonal cell content ≥ 0.002; (3) the distribution of clonal frequencies is discontinuous, and the judgment criteria are as follows: (a) Take the clone types with frequencies less than 0.03 and arrange them in descending order; (b) Use 10 times the clonal frequency of the 5th clone type in the sequential arrangement as the threshold; (c) Compare the clonal frequencies of all clone types with the threshold, and the clone types with frequencies greater than or equal to the threshold are considered to satisfy the discontinuous distribution of clonal frequencies.
[0025] Furthermore, the screening criteria for significant fusion types are as follows: (1) fusion frequency ≥ 0.1%; (2) fusion cell content ≥ 0.02%.
[0026] Furthermore, the calculation method for MRD is as follows:
[0027] MRD 克隆 = max(sum (IGH clonal cell content), sum(IGK clonal cell content) + sum(IGL clonal cell content));
[0028] MRD 融合= max(sum of BCL1-IGH fusion cell content), sum(BCL2-IGH fusion cell content));
[0029] ;
[0030] or
[0031] 。
[0032] Furthermore, the model further includes: obtaining the clonotypes of healthy individuals and / or patients with non-homogeneous diseases, and calculating the clonal cell content; based on the clonotypes of the healthy individuals and / or patients with non-homogeneous diseases, screening out the clonotypes with non-rare presence and clonal cell content ≥ 10 -7 from the significant clonotypes obtained from the pre-treatment DNA sequencing data, where the non-rare presence is the clonotype present in ≥ 5% of the population, which forms the basis for subsequent MRD calculation.
[0033] The beneficial effects brought by the technical solution provided by the embodiments of the present invention include:
[0034] The present invention proposes a simple library construction method. By using the one-step method for library construction, the library can be obtained only through one round of PCR amplification, significantly shortening the library construction time. At the same time, through the single-pool amplification strategy, the primers of genes such as IGH, IGK, IGL, KDE, etc. and the IGH-BCL1 and IGH-BCL2 fusion sites are mixed in one pool, realizing the synchronous detection of immune receptor rearrangement and key fusion gene events in a single PCR reaction, greatly reducing the sample requirement, and being particularly beneficial for tracking the rearrangement types of samples after treatment. By introducing the IGH-BCL fusion detection, the capture ability of specific molecular markers for blood tumors such as lymphoma is further enhanced. In terms of the construction of the detection system, this scheme uses hundreds of highly specific primers to comprehensively cover all immune cell receptor chains (including IGH, IGL, IGK, IGH-DJ, IGK-KDE, etc.) and the IGH-BCL fusion hot spot region, breaking through the detection blind spots of traditional customized Panels. Through the optimized significant clone screening algorithm combined with the uniqueness index evaluation model, malignant clones and physiological rearrangements can be accurately distinguished, effectively solving the clinical pain point of the limited prognostic value of IGK / IGL detection. Experimental verification shows that the minimum detection sensitivity of this technical solution reaches 10 -7Level, and can still maintain stable detection performance in trace samples. In MRD detection, the sensitivity of ctDNA is a key challenge. The proportion of ctDNA in cfDNA is only 0.01%, and it is affected by various factors such as tumor type and burden. The technical solution proposed by the present invention has excellent detection performance, can effectively solve the problem of low abundance of ctDNA, and provide accurate MRD detection. This method not only improves the detection accuracy, but also provides more comprehensive and in-depth tumor biological information for clinical practice. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0036] Figure 1 It is a relationship diagram of the number of sample cases and the ratio of internal references provided by the embodiment of the present invention. The abscissa is the number of sample cases, and the ordinate is the ratio of the number of sequences after correction of internal reference 1 and internal reference 2;
[0037] Figure 2 It is a schematic diagram of the quantitative result of purified DNA provided by the embodiment of the present invention;
[0038] Figure 3 It is a schematic diagram of the best matching mode of seq1 and seq2 provided by the embodiment of the present invention;
[0039] Figure 4 It is a schematic diagram of hierarchical clustering provided by the embodiment of the present invention;
[0040] Figure 5 It is a schematic diagram of the quantitative result of purified DNA in the linear range of ALL cell line with a loading amount of 20 μg provided by Embodiment 3 of the present invention;
[0041] Figure 6 It is a schematic diagram of the quantitative result of purified DNA in the linear range of CLL cell line with a loading amount of 20 μg provided by Embodiment 3 of the present invention;
[0042] Figure 7 It is a schematic diagram of the quantitative result of purified DNA in the linear range of MM cell line with a loading amount of 20 μg provided by Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The present invention will be further described in detail below through specific embodiments. However, those skilled in the art will understand that the following embodiments are only used to illustrate the present invention and should not be construed as limiting the scope of the present invention. For those without specific techniques or conditions noted in the embodiments, the techniques or conditions described in the literature in this field or according to the product specifications shall be followed. For reagents or instruments without the manufacturer noted, they are all conventional products that can be obtained through commercial purchase.
[0044] As used herein, the words "comprising", "including", "having" or any other variant thereof are intended to cover non-exclusive inclusion. For example, a process, method, article or apparatus that comprises a listed element(s) is not necessarily limited to those element(s), but may include other elements not expressly listed or inherent to such process, method, article or apparatus. Unless the context clearly dictates otherwise, the singular forms "a / an" and "the" include plural referents.
[0045] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the art. In addition to the specific methods, devices, and materials used in the embodiments, any methods, devices, and materials of the prior art similar or equivalent to those described in the embodiments of the present invention can also be used to implement the present invention according to the knowledge of those skilled in the art in the prior art and the description of the present invention.
[0046] Unless otherwise specified, the experimental methods, detection methods, and preparation methods disclosed in the present invention all adopt the conventional techniques in the fields of molecular biology, immunology, laboratory medicine, gene sequencing technology, bioinformatics technology, and related fields in this technical field.
[0047] The embodiments of the present invention provide a monitoring model for minimal residual disease in human blood tumors, including:
[0048] (1) A data acquisition module, configured to receive DNA sequencing data before and after treatment of the same detection object;
[0049] (2) A nucleated cell count acquisition module, configured to align the sequencing data of the reference gene to the reference genome to obtain the nucleated cell count;
[0050] (3) A clone map and sequence acquisition module, configured to align the DNA sequencing data to an immune database including IGH, IGK, and IGL genes through an alignment algorithm, and then obtain a clone map and the corresponding sequence through a clustering algorithm;
[0051] (4) A fusion genotype and sequence acquisition module, configured to align the DNA sequencing data to the reference gene sequence to obtain the fusion genotype and the corresponding sequence;
[0052] (5) A calculation module, configured to calculate the clone frequency, clone cell content, fusion frequency, and fusion cell content based on the number of nucleated cells, clone maps and corresponding sequences, fusion genotypes, and corresponding sequences; and determine significant clone types and significant fusion types based on the clone frequency, clone cell content, fusion frequency, and fusion cell content;
[0053] (6) An MRD calculation and determination module, configured to screen effective clone types and fusion types according to the clone types and fusion types obtained before and after the treatment of the same subject, calculate the MRD value based on the clone cell content and fusion cell content corresponding to the effective clone types and fusion types, and determine the residual status of human hematological malignancies based on the MRD value.
[0054] Specifically, the screening of effective clone types and fusion types includes:
[0055] (1) Clone types and fusion types that exist after the treatment of the same detection subject and are the same as the significant clone types and significant fusion types before the treatment.
[0056] (2) Clone types that exist in the subsequent detections of the same detection subject and are the same as the significant clone types newly discovered after the treatment.
[0057] The significant clone types newly discovered after the treatment are significant clone types that exist after the treatment and are different from the significant clone types before the treatment, and the subsequent detections are carried out for tracking the significant clone types newly discovered after the treatment.
[0058] In MRD detection, the sensitivity of ctDNA is a key challenge. The proportion of ctDNA in cfDNA is only 0.01%, and it is affected by various factors such as tumor type and burden. Especially in early cancer patients, low-abundance ctDNA may lead to false negative MRD results. Based on the BCR database, through significant clone screening and unique index evaluation, it is possible to determine whether it is a cancer cell clone, solving the problem of limited prognostic value of IGK and IGL-MRD. This method not only improves the detection accuracy, but also provides more comprehensive and in-depth tumor biological information for clinical practice. The detection performance is excellent, and the lowest detection limit can reach 10 -7 。
[0059] Specifically, the data acquisition module realizes the DNA sequencing data of a plurality of the same detection subject before and after the treatment through the cooperation of a detection kit and a monitoring device. The same clone types and fusion types refer to the same gene sequences.
[0060] Specifically, the detection kit includes: an upstream outer primer F1, an upstream inner primer F2, and a downstream outer primer R; the upstream outer primer F1 consists of a sequencing adapter sequence 1, a Barcode sequence, and an upstream universal sequence 1, the adapter sequence 1 is as shown in SEQ ID No.5, the Barcode sequence is as shown in SEQ ID No.6-15, and the upstream universal sequence 1 is as shown in SEQ ID No.16;
[0061] The upstream inner primer F2 consists of an upstream universal sequence 2 and an upstream specific primer sequence, the upstream universal sequence 2 is as shown in SEQ ID No.17, and the upstream specific primer sequence is as shown in SEQ ID No.18-163;
[0062] The downstream outer primer R consists of a downstream specific primer sequence, a downstream adapter sequence, and a downstream sequencing adapter sequence, and the primer 5' is phosphorylated. The downstream specific primer sequence is as shown in SEQ ID NO.164-183, the downstream adapter sequence is as shown in SEQ ID No.184, and the downstream sequencing adapter sequence is as shown in SEQ ID No.185.
[0063] Optionally, the kit includes 2 internal reference genes. The upstream specific primer sequence of the first internal reference gene is as shown in SEQ ID NO.1, and the downstream specific primer sequence of the first internal reference gene is as shown in SEQ ID NO.2; the upstream specific primer sequence of the second internal reference gene is as shown in SEQ ID NO.3, and the downstream specific primer sequence of the second internal reference gene is as shown in SEQ ID NO.4.
[0064] I. Selection of internal reference genes
[0065] Specifically, considering the diverse sample types, 2 amplicons are selected as internal references to ensure the stability of the total internal reference sequence number. In the selection of internal references, regions with good sequence specificity are preferentially selected for design. In addition, since the time interval between pre-treatment samples and post-treatment samples is relatively long, in order to ensure the traceability and comparison of the same sample before and after, 3 SNP sites are specifically designed on the internal reference amplicon. Each site has 3 genotypes. Samples collected from the same individual at different time periods should have the same genotypes at the 3 SNP sites, significantly reducing the risk of sample confusion caused by sampling or experimental operations.
[0066] Two genes are selected as internal references, with hg19 as the reference. The specific information is as follows:
[0067] (1) Internal reference 1
[0068] Position information: chr20:15124894 - 15125033; upstream specific primer sequence: GGCCTACAATTCAAATTAATGTAAAAACTGC (SEQ ID NO:1), upstream primer position: chr20:15124894 - 15124924; downstream primer position: chr20:15125006 - 15125033, downstream specific primer sequence: TCTGATTCAAACTTTTCTTTTGTGGCTG (SEQ ID NO:2).
[0069] (2) Internal reference 2
[0070] Position information: chr5:17374840 - 17374979; upstream specific primer sequence: GCACTTTCAATCAACTGTGTTAGATTGA (SEQ ID NO:3), upstream primer position: chr5:17374840 - 17374867; downstream primer position: chr5:17374951 - 17374979, downstream specific primer sequence: TATGATCCACATTGTATGGTTTTTAGGCA (SEQ ID NO:4).
[0071] The primer sequences of the internal reference genes have good sequence specificity on the genome, and there is no amplification of pseudogene sequences. After performing whole - genome alignment, there is a 140 - bp fragment that can be 100% aligned for internal reference 1 and internal reference 2 respectively, and the positions are as expected. The other two alignable fragments are only partial regions of the amplicons, and their sizes do not exceed 30 bp. It is considered that the probability of non - amplification is relatively high. See Table 1 for details.
[0072] Table 1 Results of internal reference alignment
[0073]
[0074] There are 3 SNP sites on the internal reference genes. As shown in Table 2, they stably exist in normal and tumor cells, are not related to the cell cycle and cell activation, and are not affected by any endogenous or exogenous factors (such as not being affected by any experimental treatment measures). At the same time, they can be used as paired information for samples before and after treatment to prevent sample contamination.
[0075] Table 2 Information on SNP sites of internal reference
[0076]
[0077] Select 136 clinical samples of hematological tumors and perform library construction, purification, quantification, pooling and dilution, sequencing, and data analysis according to the following steps, and calculate the ratio of the sequence number of internal reference 1 to the sequence number of internal reference 2, as Figure 1As shown, the results indicate that the proportion of the two internal reference-corrected sequence numbers is relatively stable.
[0078] II. Library Construction
[0079] Different samples use different specific adapters; the gDNA of the same sample is mixed into a pool and uses the same specific adapter (upstream outer primer); the samples are divided into pre-treatment samples and post-treatment samples, and their starting amounts are different.
[0080] For each clinical sample, an equal amount of gDNA is used to construct a library in a single tube. Take 0.2 mL PCR tubes with the number of clinical samples + 1 (control product), and prepare the reaction system according to Table 3. Among them, the ratio of upstream outer primer: upstream inner primer mixture: downstream primer mixture = 5:3:20.
[0081] Table 3 Amplification System
[0082]
[0083] Preparation of primer mixture:
[0084] The primer sequences consist of upstream outer primer F1, upstream inner primer F2, and downstream outer primer R, as shown in Tables 4 and 5.
[0085] Among them, upstream outer primer F1 consists of sequencing adapter sequence 1 (CAAGCAGAAGACGGCATACGAGAT SEQ ID NO:5), Barcode sequence (SEQ ID NO:6 - 15), and upstream universal sequence 1 (GGCACCCGAGAATTCCA SEQ ID NO:16).
[0086] Upstream inner primer F2 (i.e., internal reference upstream primer, IGH-V, IGH-D upstream primer, IGK-V, IGKED upstream primer, IGL-V upstream primer, fusion upstream primer) consists of upstream universal sequence 2 (CGAACCGCTTTCGGAGG SEQ ID NO:17) and upstream specific primer sequence (SEQ ID NO:18 - 163).
[0087] Downstream outer primer R (i.e., internal reference downstream primer, IGH-J downstream primer, IGK-J, IGKDE downstream primer, and IGL-J downstream primer) consists of downstream specific primer sequence (SEQ ID NO:164 - 183), downstream adapter sequence (TCAGAGTTCTACAGTCCGACGATC SEQ ID NO:184), and downstream sequencing adapter sequence (AATGATACGGCGACCACCGAGATCTACACGT SEQ ID NO:185), and the 5' end of the primer is phosphorylated.
[0088] Dilute the upstream outer primer F1 step by step to 100 μM;
[0089] Dilute the upstream inner primer F2 step by step to 100 μM, and mix them in equimolar amounts to obtain the upstream inner primer MIX1 (i.e., the internal reference and IG, fusion upstream primer mixture);
[0090] Dilute the downstream outer primer R step by step to 100 μM, and mix them in equimolar amounts to obtain the downstream primer MIX2 (i.e., the internal reference and IG downstream primer mixture).
[0091] The PCR amplification reaction conditions are shown in Table 6.
[0092] Table 4 Barcode sequence information
[0093]
[0094] Table 5 Specific primer sequence information
[0095]
[0096] Table 6 PCR amplification reaction conditions
[0097]
[0098] III. Library purification
[0099] Use the MPure XP-PCR product purification kit (Beckman Coulter, A63882) to purify the amplification product obtained in step II to obtain the purified DNA library. The purified DNA library solution can be stored at 4°C for no more than 16 hours. If it needs to be stored for a longer time, it should be stored below -18°C.
[0100] IV. Library quantification
[0101] Use TapeStation2200 / 4200 to quantify the concentration and band size of the purified DNA library obtained above. The results of TapeStation2200 quantification are as Figure 2 .
[0102] Agilent2200 / 4200 quality control: After magnetic bead purification, for the library to be qualified, the target fragment should be between 220 - 400 bp, and the proportion of dimer bands (less than 200 bp) should not be greater than 20%, and it can be directly loaded onto the machine. If the proportion of dimers after purification is greater than 20%, it is unqualified and the library needs to be rebuilt.
[0103] V. Library Pooling and Dilution
[0104] (1) Mix all the DNA libraries selected in Step 4 together according to the mass ratio, and label it as the library mixture for loading.
[0105] (2) The total concentration of the target fragments in the library mixture for loading within the range of 220 - 400 bp is the library concentration. It is recommended that the data volume for each sample loaded is 5G.
[0106] Note: The mixed library needs to be vortexed for 1 min and centrifuged for 2 seconds, and this step should be repeated at least 3 times until the mixed library is evenly mixed. It is necessary to ensure that the library is evenly mixed before quantitative detection.
[0107] VI. Sequencing on the Machine
[0108] After the prepared library mixture for loading is quantified by Qubit, it is sequenced on the GENETRONS2000 platform according to the established data volume requirements. The data downloaded from the machine is a PE150 fastq file, which is the original sequencing data.
[0109] The original sequencing data is obtained by using the data acquisition module provided by the present invention for the human leukemia minimal residual monitoring model, that is, receiving the DNA sequencing data before and after the treatment of the same detection object.
[0110] VII. Data Analysis
[0111] It should be noted that the module for obtaining the number of nucleated cells, which is used to align the sequencing data (sequencing sequences) of the reference genes (chr5:17374840 - 17374979, chr20:15124894 - 15125033, human hg19 genome position) to the reference genome (hg19, human hg38) to obtain the number of nucleated cells, is the prior art and will not be elaborated here.
[0112] The module for obtaining the clone map and sequence is used to align the DNA sequencing data to the immune database including IGH, IGK, and IGL genes through a comparison algorithm, and then obtain the clone map and the corresponding sequence through a clustering algorithm.
[0113] Specifically, the sequencing data is aligned to the immune database including IGH, IGK, and IGL genes (source: IMGT http: / / www.imgt.org) through a comparison algorithm (reference: https: / / academic.oup.com / nar / article / 41 / 10 / e108 / 1075719), and then the clone map is obtained through a clustering algorithm. The clone map includes, but is not limited to, clones of the types IGH, IGK, IGL, IGH - DJ, IGK - KDE or subsets of these types.
[0114] The clustering algorithm first calculates the edit distance between sequences as follows:
[0115] ;
[0116] Among them, editdist represents the edit distance between sequences. The edit distance is a function related to the sequences. seq1 and seq2 represent two sequences obtained by sequencing.
[0117] The edit distance can be calculated by the following formula:
[0118] ;
[0119] Among them, i represents the i-th base in seq1, j represents the j-th base in seq2. Both i and j start from 1 until the last position, and their initial values are both 0.
[0120] ;
[0121] Among them, ms represents the match score (positive number). When the i-th base of seq1 is the same as the j-th base of seq2, the value of m(i, j) is positive. ss represents the mismatch score (negative number). When the i-th base of seq1 is different from the j-th base of seq2, the value of m(i, j) is negative.
[0122] ;
[0123] Among them, os represents the score for the start of a gap (negative number), gap opening represents the start situation of a gap in sequence matching, es represents the score for the extension of a gap (negative number), gap extending represents the extension situation of a gap in sequence matching. Generally, os is greater than es.
[0124] S represents the clip score (negative number). When the bases of seq1 and seq2 do not match highly and the negative value is too high, truncation will be selected at this time.
[0125] To obtain the best matching method, see Figure 3 .
[0126] In order to better cluster the subsequent data, the values of ms, ss, os, es, and S satisfy the following formula:
[0127] ;
[0128] Wherein, n is an integer greater than or equal to 0 and less than the sequencing read length. The sixth inequality in the above formula indicates that there must exist a certain boundary n such that the absolute value of the soft truncation (S) is definitely less than the gap opening (os) plus n gap extensions (es) plus 2.
[0129] Preferably, ms can take the value of 1, and other values can be integer multiples of ms.
[0130] Through the above limitations, weight distribution can be carried out for various situations, which is beneficial to the subsequent clustering of data and avoids unclear boundaries and chaotic parent nodes during the clustering process. For example, when ms takes the value of 1, ss takes the value of -4, os takes the value of -5, es takes the value of -1, and S takes the value of -6.
[0131] Then, according to the edit distance and its relationship with the V, D, J regions of the IGH gene, the V, J regions of the IGK gene, and the V, J regions of the IGL gene, the sequenced sequences are subjected to clustering analysis to obtain a clone map. The specific process is as follows:
[0132] (1) According to the alignment result of the alignment algorithm and the edit distance, the best matching method is given, that is, the edit distance score is the largest. The sequencing sequences with the same V and J regions of the IGH, IGK, and IGL genes are clustered. When IGH-DJ rearrangement occurs, the sequencing sequences with the same D and J regions of the IGH gene are clustered. When IGK-KDE rearrangement occurs, the sequencing sequences with the same V regions of the KDE gene and the IGK gene are clustered.
[0133] (2) In each VJ class, it is divided according to the CDR3 region consistency; in each DJ class, it is divided according to the 50bp-length region consistency centered on the junction position of the D and J regions of the IGH gene; in each KDE-V class, it is divided according to the 50bp-length region consistency centered on the junction position of the V region of the IGK gene and the KDE gene; the number of sequencing sequences corresponding to each subclass is obtained.
[0134] (3) The hierarchical clustering method is used to cluster each subclass. In each cluster, the child node and the parent node sequences are highly similar, and the parent node has a significantly higher number of sequencing sequences than the child node.
[0135] Specifically, it includes: obtaining the average error rate of sequencing bases. This coefficient is generally an empirical value constant related to the sequencing platform, etc. The subclass coefficient is evaluated based on the error rate (hereinafter referred to as factor, and this coefficient is greater than 1). In each category (VJ class, DJ class, and KDE-V class), they are arranged in descending order according to the number of sequencing sequences. According to the sorting result and the edit distance algorithm, the category with a higher number of sequences is compared with other subclasses in turn, and hierarchical clustering method is used for clustering. In each cluster, the child node and the parent node sequences are highly similar, and the parent node has a significantly higher number of sequencing sequences than the child node, that is, the number of parent node sequences divided by the number of child node sequences is greater than factor.
[0136] (4) Sum up the number of sequencing sequences of all nodes in each cluster, and report the sequence of the root node (the top parent node) as a separate clonotype.
[0137] (5) Combine each reported separate clonotype to form a clonotype map.
[0138] Figure 4 is a schematic diagram of a hierarchical clustering tree. Clonetype1 is the parent node of clonetype2 and clonetype3, clonetype4 and clonetype5 are the child nodes of clonetype2, and clonetype6 is the child node of clonetype3.
[0139] The fusion genotype and sequence acquisition module is used to align DNA sequencing data to the reference gene sequence to obtain the fusion genotype and the corresponding sequence.
[0140] Specifically, it includes:
[0141] (1) Obtain the number of sequences of IGH, BCL1, and BCL2 genes respectively through the bwa alignment software.
[0142] (2) Construct the FMDIndex data structure (https: / / pmc.ncbi.nlm.nih.gov / articles / PMC3389770 / ) according to the reference sequences of IGH, BCL1, and BCL2 genes;
[0143] Use FMDIndex to find the maximum exact match set between each sequencing sequence and the reference sequence;
[0144] Use the dynamic programming algorithm to find a set of optimal maximum exact match subsets between the reference sequence and the sequencing sequence;
[0145] Use the edit distance algorithm to align the sequencing sequence (seq1) to the reference gene sequence (seq2) corresponding to the optimal exact match subset to obtain the fusion type and the corresponding number of sequences.
[0146] Using the partial - order alignment algorithm, based on the sequencing sequences corresponding to the fusion type, the fusion sequences are obtained.
[0147] A calculation module, configured to calculate the clone frequency, clone cell content, fusion frequency, and fusion cell content based on the number of nucleated cells, the clone map and the corresponding sequences, the fusion genotype and the corresponding sequences; and determine the significant clone types and significant fusion types based on the clone frequency, clone cell content, fusion frequency, and fusion cell content.
[0148] Specifically, the calculation methods of the clone frequency and clone cell content are as follows:
[0149] ;
[0150] ;
[0151] ;
[0152] ;
[0153] Among them, cfl represents the clone frequency of a single IGL clone type, count1 is the number of sequences of a single IGL clone type, count IGL is the number of sequences of all IGL clone types, cfh represents the clone frequency of a single IGH clone type, count2 is the number of sequences of a single IGH clone type, count IGH is the number of sequences of all IGH clone types except IGH - DJ clone types, count HDJ is the number of sequences of all IGH - DJ clone types; cfk represents the clone frequency of a single IGK clone type, count3 is the number of sequences of a single IGK clone type, count IGK is the number of sequences of all IGK clone types except IGK - KDE clone types, count KDE is the number of sequences of all IGK - KDE clone types; the number of clone cells is the number of sequences of a single clone type.
[0154] The calculation methods of the fusion frequency and fusion cell content are as follows:
[0155] ;
[0156] ;
[0157] Among them, fc is the number of sequences of a single fusion type, hc is the total number of IGH sequences, bc is the total number of BCL1 or BCL2 sequences, and the number of fusion cells is the number of sequences of a single fusion type.
[0158] Preferably, the calculation methods for the cloning frequency and the content of cloned cells are as follows:
[0159] ;
[0160] ;
[0161] ;
[0162] ;
[0163] clonetype represents the clone type, wf (clonetype) is a correction function.
[0164] The calculation methods for the fusion frequency and the content of fused cells are as follows:
[0165] ;
[0166] ;
[0167] fusion represents the fusion type, wf(fusion) is a correction function.
[0168] It should be noted that wf (clonetype) is a function related to the amplification efficiency and the amplification length of the clone type; wf(fusion) is a function related to the amplification efficiency and the amplification length of the fusion type.
[0169] The calculation method for the amplification efficiency of the clone type is as follows:
[0170] First, a series of plasmids with permutations and combinations of the V regions and J regions of IGH, IGK, and IGL genes (i.e., clone types) are artificially synthesized. When IGH-DJ rearrangement occurs, plasmids with permutations and combinations of the D region and J region of the IGH gene are synthesized. When IGK-KDE rearrangement occurs, plasmids with permutations and combinations of the KDE gene and the J region of the IGK gene are synthesized. After dissolving these plasmids, they are quantified, and then randomly mixed in equal mass to form 10 different plasmid mixture libraries. The 10 plasmid mixture libraries are respectively constructed and sequenced on the machine, and the obtained data are analyzed to obtain a cloning map. For each plasmid, the number of reads is divided by the average number of reads of the plasmids in the plasmid library, denoted as value a, and the amplification coefficient is calculated according to a b(clonetype) .
[0171] ;
[0172] The calculation method for the amplification efficiency of the fusion type is as follows:
[0173] First, a series of plasmids with different combinations of IGH genes and BCL1 or BCL2 genes (i.e., fusion types) were artificially synthesized. After dissolving these plasmids, they were quantified, and then randomly mixed in equal amounts to form 10 different plasmid mixture libraries. Each of the 10 plasmid mixture libraries was used to construct a library and sequenced on a machine. The obtained data were analyzed to obtain the fusion genotypes. For each plasmid, the number of reads was divided by the average number of reads of plasmids in the plasmid library, denoted as value a, and the amplification factor was calculated based on a. b (fusion).
[0174] ;
[0175] After obtaining the amplification factor, we further calculated the correction function based on the amplification length as follows:
[0176] ;
[0177] Among them, λ1 and λ2 take fixed values. Preferably, λ1 takes the value of 1.03 and λ2 takes the value of 1.08. lengh(plasmid) is the length of the synthetic clone / fusion (plasmid) sequence, length(clonetype) is the length of this clone sequence in the test sample, length(fusion) is the length of this fusion sequence in the test sample.
[0178] After obtaining the above data, significant clones and significant fusions were screened. The screening criteria for significant clones are as follows:
[0179] (1) Cloning frequency ≥ 0.03;
[0180] (2) Cloning cell content ≥ 0.002;
[0181] (3) The cloning frequency distribution is discontinuous, and the judgment criteria are as follows:
[0182] (a) Take the clones with a frequency less than 0.03 and arrange them in descending order;
[0183] (b) Take 10 times the cloning frequency of the 5th clone in the arranged order as the threshold;
[0184] (c) Compare the cloning frequencies of all clones with the threshold. Clones with a cloning frequency greater than or equal to the threshold are considered to satisfy the discontinuous cloning frequency distribution.
[0185] Significant clones can be any one or more combinations and subsets of IGH, IGK, IGL, IGH-DJ, IGK-KDE.
[0186] The screening criteria for significant fusions are as follows:
[0187] (1)The fusion frequency ≥ 0.1%;
[0188] (2)The content of fused cells ≥ 0.02%.
[0189] For the same detection object, sample DNA is extracted before and after treatment respectively, and the significant clone types and significant fusion types before treatment and the clone types and fusion types after treatment are obtained respectively. The treatment includes but is not limited to surgical treatment, immunotherapy, targeted therapy, chemotherapy or radiotherapy.
[0190] Effective clone types are screened out according to the clone types obtained after treatment of the same detection object, and the clone cell content of the effective clone types is recorded, and MRD is calculated therefrom. 克隆 。
[0191] Specifically, the screening methods for effective clone types include:
[0192] (1)Clone types that exist after treatment of the same detection object and are the same as the significant clone types before treatment.
[0193] (2)Clone types that exist in subsequent detections of the same detection object and are the same as the newly discovered significant clone types after treatment.
[0194] The newly discovered significant clone types after treatment are significant clone types that exist after treatment and are different from the significant clone types before treatment, and subsequent detections are carried out for the newly discovered significant clone types after treatment.
[0195] MRD 克隆 Value calculation method, tracking the heavy chain (IGH) and light chains (IGK, IGL).
[0196] a. When only the heavy chain (IGH) clone is tracked;
[0197] MRD 克隆 = sum(IGH clone cell content);
[0198] b. When only the light chain (IGK, IGL) clone is tracked;
[0199] MRD 克隆 = sum(IGK clone cell content) + sum(IGL clone cell content);
[0200] c. When both the heavy chain and light chain clones are tracked;
[0201] MRD 克隆 = max(sum(IGH clone cell content), sum(IGK clone cell content) + sum(IGL clone cell content)), that is, the maximum value between the heavy chain and the light chain is selected to represent MRD. 克隆 。
[0202] Select effective fusion types from the fusion types obtained after treatment of the same detection object, record the fusion cell content of the effective fusion types, and calculate MRD from this. 融合 The effective fusion types are the fusion types that exist after treatment of the same detection object and are the same as the significant fusion types before treatment.
[0203] a. When only BCL1-IGH fusion is traced;
[0204] MRD 融合 = sum(BCL1-IGH fusion cell content);
[0205] b. When only BCL2-IGH fusion is traced;
[0206] MRD 融合 = sum(BCL2-IGH fusion cell content);
[0207] c. When both BCL1-IGH and BCL2-IGH fusions are traced;
[0208] MRD 融合 = max(sum (BCL1-IGH fusion cell content), sum(BCL2-IGH fusion cell content));
[0209] The calculation method of MRD is:
[0210] MRD 克隆 = max(sum(IGH clonal cell content), sum(IGK clonal cell content) + sum(IGL clonal cell content));
[0211] MRD 融合 = max(sum (BCL1-IGH fusion cell content), sum(BCL2-IGH fusion cell content));
[0212] ;
[0213] Or
[0214] 。
[0215] When MRD > 0, it is judged that there is residue in human blood tumor. When MRD = 0, it is judged that there is no residue in human blood tumor.
[0216] For example, for blood cancer patient 1, significant clonotypes IGH-A (clonal cell content 3.64E-03) and IGH-B (clonal cell content 8.25E-03) were detected before treatment. After 3 months of treatment, clonotype IGH-A (clonal cell content 2.58E-06) was detected. Then, clonotype IGH-A is the effective clonotype selected by this detection, and MRD = MRD 克隆 = 2.58E-06 > 0, so the result of the detection after 3 months of treatment indicates that there is residual disease.
[0217] For example, for blood cancer patient 2, significant fusion type BCL1-IGHJ (fusion cell content 5.79E-04) was detected before treatment. After 3 months of treatment, fusion type BCL1-IGHJ (fusion cell content 6.38E-06) was detected. Then, fusion type BCL1-IGHJ is the effective fusion type selected by this detection, and MRD = MRD 融合 = 6.38E-06 > 0, so the result of the detection after 3 months of treatment indicates that there is residual disease.
[0218] For example, for blood cancer patient 3, significant clonotypes IGH-A (clonal cell content 6.42E-02) and IGH-B (clonal cell content 4.12E-03) were detected before treatment. After 3 months of treatment, no clonotypes identical to significant clonotypes IGH-A and IGH-B were detected, but a newly discovered significant clonotype IGH-C (clonal cell content 5.35E-03) was detected. MRD = 0, so the result of the detection after 3 months of treatment indicates that there is no residual disease. This patient adopted a new treatment plan. After 3 months of the new plan treatment, only clonotype IGH-C (clonal cell content 4.38E-04) was detected. Then, clonotype IGH-C is the effective clonotype selected by this detection, and MRD = MRD 克隆 = 4.38E-04 > 0, so the result of the detection after 3 months of the new plan treatment indicates that there is residual disease.
[0219] The model further includes: obtaining the clonotypes of healthy people and / or patients with non-homogeneous diseases, and calculating their clonal cell contents; based on the clonotypes of the healthy people and / or patients with non-homogeneous diseases, screening out the clonotypes that are not rarely present and have a clonal cell content ≥ 10 in the significant clonotypes obtained from the pre-treatment DNA sequencing data, which constitutes the basis for subsequent MRD calculation. -7 of the clonotypes to form the basis for subsequent MRD calculation.
[0220] The above method can run in the calculation module or in the MRD calculation and determination module. By obtaining the clonotypes of healthy people and / or patients with non-homogeneous diseases, for those that are not rarely present and have a clonal cell content ≥ 10 -7Clone types are screened to obtain the best set of clone types. In addition, for the same detection object, new significant clone types may be generated after treatment, and these significant clone types are different from the significant clone types before treatment, but the new significant clone types need to be included in the subsequent MRD calculation.
[0221] An embodiment of the present invention also provides a computer storage medium, which stores a computer program. When the computer program is executed by a processor, the following method is implemented. The method includes:
[0222] Receiving DNA sequencing data of the same detection object before and after treatment;
[0223] Aligning the sequencing data of the reference gene to the reference genome to obtain the number of nucleated cells;
[0224] Aligning the DNA sequencing data to an immune database including IGH, IGK, and IGL genes through an alignment algorithm, and then obtaining a clone map and corresponding sequences through a clustering algorithm;
[0225] Aligning the DNA sequencing data to the reference gene sequence to obtain the fusion genotype and corresponding sequences;
[0226] Based on the number of nucleated cells, the clone map and corresponding sequences, the fusion genotype and corresponding sequences, calculate the clone frequency, clone cell content, fusion frequency, and fusion cell content; and determine significant clone types and significant fusion types based on the clone frequency, clone cell content, fusion frequency, and fusion cell content;
[0227] According to the clone types and fusion types obtained before and after treatment of the same object, screen effective clone types and fusion types, and calculate the MRD value based on the clone cell content and fusion cell content corresponding to the effective clone types and fusion types.
[0228] Further, the clustering algorithm includes: calculating the edit distance between sequences based on the DNA sequencing data; performing clustering analysis on the sequenced sequences based on the edit distance and its relationship with the V, D, J regions of the IGH gene, the V, J regions of the IGK gene, and the V, J regions of the IGL gene to obtain a clone map.
[0229] Further, the calculation method of the edit distance is:
[0230] ;
[0231] ;
[0232] ;
[0233] Among them, seq1 and seq2 represent two sequenced sequences. i represents the i-th base in seq1, and j represents the j-th base in seq2. Both i and j start from 1 and go to the last position, with initial values of 0. ms represents the matching score. When the i-th base of seq1 is the same as the j-th base of seq2, the value of m(i, j) is positive. ss represents the mismatch score. When the i-th base of seq1 is different from the j-th base of seq2, the value of m(i, j) is negative. os represents the score for the start of a gap. Gap opening represents the starting situation of a gap in sequence matching. es represents the score for the continuation of a gap. Gap extending represents the continuation situation of a gap in sequence matching. S represents the clip score. When the bases of seq1 and seq2 do not match highly and the negative value is too high, truncation will be selected at this time.
[0234] Furthermore, the clustering analysis includes: giving the best matching method according to the alignment result and the edit distance, clustering the sequencing sequences with the same V regions and J regions of IGH, IGK, and IGL genes. When IGH-DJ rearrangement occurs, clustering the sequencing sequences with the same D regions and J regions of the IGH gene. When IGK-KDE rearrangement occurs, clustering the sequencing sequences with the same J regions of the KDE gene and the IGK gene. In each VJ class, it is divided according to the CDR3 region consistency. In each DJ class, it is divided according to the consistency of a 50bp-length region centered on the junction position of the D region and J region of the IGH gene. In each KDE-V class, it is divided according to the consistency of a 50bp-length region centered on the junction position of the V region of the IGK gene and the KDE gene. Obtaining the number of sequencing sequences corresponding to each subclass. Using the hierarchical clustering method to cluster each subclass. In each cluster, the sequences of the child nodes and the parent nodes are highly similar, and the parent node has a significantly higher number of sequencing sequences than the child nodes. Summing up the number of sequencing sequences of all nodes in each cluster and reporting the sequence of the root node as a single clonotype. Combining each reported single clonotype to form a clone map.
[0235] Furthermore, specifically, the calculation methods of the clone frequency and the clone cell content are as follows:
[0236] ;
[0237] ;
[0238] ;
[0239] ;
[0240] Among them, cfl represents the clonal frequency of a single IGL clone type, count1 is the number of sequences of a single IGL clone type, count IGL is the number of sequences of all IGL clone types, cfh represents the clonal frequency of a single IGH clone type, count2 is the number of sequences of a single IGH clone type, count IGH is the number of sequences of all IGH clone types except IGH-DJ clone types, count HDJ is the number of sequences of all IGH-DJ clone types; cfk represents the clonal frequency of a single IGK clone type, count3 is the number of sequences of a single IGK clone type, count IGK is the number of sequences of all IGK clone types except IGK-KDE clone types, count KDE is the number of sequences of all IGH-KDE clone types; the number of clonal cells is the number of sequences of a single clone type.
[0241] The calculation methods of the fusion frequency and the fusion cell content are as follows:
[0242] ;
[0243] ;
[0244] Among them, fc is the number of sequences of a single fusion type, hc is the total number of IGH sequences, bc is the total number of BCL1 or BCL2 sequences, and the number of fusion cells is the number of sequences of a single fusion type.
[0245] Preferably, the calculation methods of the clonal frequency and the clonal cell content are as follows:
[0246] ;
[0247] ;
[0248] ;
[0249] ;
[0250] clonetype represents the clone type, wf (clonetype) is the correction function.
[0251] The calculation methods of the fusion frequency and the fusion cell content are as follows:
[0252] ;
[0253] ;
[0254] fusion represents the fusion type, wf(fusion)It is a correction function.
[0255] Furthermore, the screening criteria for significant clones are as follows:
[0256] (1) Cloning frequency ≥ 0.03;
[0257] (2) Cloned cell content ≥ 0.002;
[0258] (3) The cloning frequency distribution is discontinuous, and the judgment criteria are as follows:
[0259] (a) Take the clones with a frequency less than 0.03 and arrange them in descending order;
[0260] (b) Take 10 times the cloning frequency of the 5th clone in the sequential arrangement as the threshold;
[0261] (c) Compare the cloning frequencies of all clones with the threshold. Clones with a frequency greater than or equal to the threshold are considered to satisfy the discontinuous cloning frequency distribution.
[0262] Furthermore, the screening criteria for significant fusions are as follows:
[0263] (1) Fusion frequency ≥ 0.1%;
[0264] (2) Fusion cell content ≥ 0.02%.
[0265] Furthermore, the calculation method of MRD is as follows:
[0266] MRD 克隆 = max(sum(IGH cloned cell content), sum(IGK cloned cell content) + sum(IGL cloned cell content));
[0267] MRD 融合 = max(sum (BCL1-IGH fusion cell content), sum(BCL2-IGH fusion cell content));
[0268] ;
[0269] Or
[0270] .
[0271] Furthermore, the model further includes: obtaining the clones of healthy people and / or patients with non-homogeneous diseases, and calculating their cloned cell content; based on the clones of healthy people and / or patients with non-homogeneous diseases, screening the significant clones obtained from the pre-treatment DNA sequencing data for those that are not rarely present in healthy people and / or patients with non-homogeneous diseases and have a cloned cell content ≥ 10 -7clonal types, where the non-rare presence refers to clonal types present in ≥5% of the population.
[0272] Example 1 Detection of Minimal Residual Disease in Human Leukemia and Lymphoma
[0273] A total of 13 post-treatment samples with known pathological information were obtained from Beijing Genepioneer Biotechnologies Inc. and partner hospitals, and 1 cell line was purchased through commercial channels. These included 4 cases of acute lymphoblastic leukemia (ALL) (Samples 1 and 2 were bone marrow gDNA; Samples 3 and 4 were peripheral blood gDNA), 4 cases of chronic lymphocytic leukemia (CLL) (Samples 5 and 6 were bone marrow gDNA; Samples 7 and 8 were peripheral blood gDNA), 3 cases of follicular lymphoma (FL) (Samples 9, 10, and 11 were bone marrow gDNA), 2 cases of multiple myeloma (MM) (Samples 12 and 13 were bone marrow gDNA), and 1 case of mantle cell lymphoma (MCL) (Sample 14 was cell line gDNA).
[0274] Bone marrow DNA, peripheral blood DNA, or cell line DNA extracted from 13 patients before and after treatment were used as starting samples respectively, and according to the technical solution proposed by the present invention, the MRD detection results were obtained.
[0275] Table 7 Sample Detection Results
[0276]
[0277] Within the scope of this detection, the detection of different clonal types ultimately showed positive MRD in all cases. As described above, when using the detection method provided by the present invention to detect positive cases of human leukemia and lymphoma, tracking MRD is effective, and it performs well in terms of the richness of detected clonal types, indicating that this detection method has practical applicability. Among them, the positive clones detected by IGH covered 10 positive samples of 4 pathological types including ALL, CLL, MM, and MCL; the positive clones detected by IGK covered 11 positive samples of 5 pathological types including ALL, CLL, FL, MM, and MCL; the positive clones detected by IGL covered 6 positive samples of 4 pathological types including CLL, FL, MM, and MCL; the positive clones detected by HDJ covered 4 positive samples of 3 pathological types including ALL, CLL, and MM; the positive clones detected by KDE covered 6 positive samples of 4 pathological types including ALL, CLL, FL, and MM; the positive clones detected by BCL-IGHJ rearrangement covered 6 positive samples of 4 pathological types including ALL, CLL, FL, and MCL. Thus, it can be seen that the present invention can effectively detect MRD in human leukemia, lymphoma, and myeloma.
[0278] Example 2 Sample Consistency Detection
[0279] Sixteen patients with known pathological information were selected. Among them, 6 patients selected pre-treatment samples, 5 patients selected post-treatment samples, and 5 patients selected both pre-treatment and post-treatment samples, for a total of 21 samples to be detected. Samples 16 and 17, 18 and 19, 22 and 23, 25 and 26, 28 and 29 were from the same patient. The clonotype detection was performed using the method of the present invention and MFC. Among the detection results of all 21 tests, there were 14 positive results and 7 negative results; 11 pre-treatment samples and 10 post-treatment samples were covered; among them, there were 16 bone marrow samples and 5 peripheral blood samples; the covered pathological types included acute lymphoblastic leukemia (ALL), chronic lymphocytic leukemia (CLL), and multiple myeloma (MM). The purpose was to evaluate the performance of the present invention in detecting MRD from multiple dimensions.
[0280] The specific results are shown in Table 8, where the "NGS result" is the detection result using the method of the present invention, and the "flow cytometry result" is the detection result using multi-color flow cytometry. The results show that the two methods are consistent in the two sample types of bone marrow and peripheral blood detected, and in three different pathological types, whether for pre-treatment or post-treatment samples.
[0281] Table 8 Summary of Detection Results of 21 Clinical Samples
[0282]
[0283] *a: The MRD detection results of both methods are positive. b: The MRD detection results of both methods are negative. c: The MRD detection results of both methods are negative. Among them, the method of the present invention detected a clone, but since it exists in the clones of healthy people or non-homologous patients, it does not belong to disease-specific clones and is not tracked.
[0284] Example 3 Sensitivity Study
[0285] 1. Linear range with a loading amount of 20 μg:
[0286] MFC was used to trace the malignant cell content of ALL cell line (RS4;11), CLL cell line (MEC-1), and MM cell line (IM-9). According to the traced malignant cell frequency, the gDNA of the 3 cell lines was diluted with the gDNA of healthy human whole blood leukocytes. According to a nucleic acid input amount of 20 μg, the number of cells initially input for detection was calculated, and linear reference products with concentrations of 1E-06, 1E-05, 1E-04, 1E-03, and 1E-02 malignant cell contents were configured. The number of cells = the initial nucleic acid input amount (pg) / 6.6 pg / cell; the prepared gradient-diluted samples were detected 3 times in duplicate according to the method of the present invention to investigate the linear range detectable by the method of the present invention.
[0287] When the sample loading amount is 20 μg, the detected disease types are: ALL, CLL, MM, and the detection range is the above gradient reference products. The theoretical frequencies and MRD values are respectively taken -lg for statistics. The data are shown in Table 9 and Figures 5 - 7 as shown, the dilution linearity R of the 3 disease cell lines 2 ≥0.95, which proves that when the sample loading amount is 20 μg, the detection of the present invention has good linearity.
[0288] Table 9 Results of linear relationship at 20 μg sample loading amount
[0289]
[0290] 2. Establishment of LOD at 20 μg sample loading amount:
[0291] Use MFC to trace the malignant cell content of ALL cell line (RS4;11), CLL cell line (MEC-1), and MM cell line (IM-9). Dilute the gDNA of the 3 cell lines and the gDNA of healthy human whole blood leukocytes according to the traced malignant cell frequencies, and dilute to the malignant cell frequencies that need to be confirmed as LOD according to Table 10.
[0292] Table 10 Corresponding relationship between malignant cell frequency and malignant cell number at 20 μg DNA input amount
[0293]
[0294] Note: The DNA quality here is calculated according to the number of bases of the human genome and the diploid characteristics, and 1 cell is calculated according to 6.5 pg.
[0295] Under the condition of a sample loading amount of 20 μg, the method of the present invention is used for 3 repeated detections at 3 MRD frequencies respectively. The LOD is initially established with a 100% detection frequency. It can be seen from the results in Table 11 that the lowest frequency at which positive results are detected three times for the 3 disease types is 6.50E-07.
[0296] Table 11 Detection results for establishing 20 μg-LOD
[0297]
[0298] 3. Confirmation of LOD at 20 μg sample loading amount:
[0299] The lowest detection limit standards ALL, CLL, and MM cell lines established by selecting LOD were respectively tested 13 times under the conditions of 2 malignant cell contents, 3 malignant cell contents, and an input amount of 20 μg DNA using the method of the present invention. The lowest detection limit was determined with a 100% positive detection rate. From the results in Table 12, it was determined that when the sample loading amount was 20 μg in the detection of the present invention, the lowest detection limits LOD of ALL, CLL, and MM were 6.50E-07.
[0300] Table 12 20 μg - LOD confirmation results
[0301]
[0302] Example 6 ctDNA detection
[0303] On the basis of the previous detected sample types, the present invention expanded the sample type of ctDNA. Plasma samples of 8 patients with known pathological information were selected, including 3 cases of diffuse large B-cell lymphoma (DLBCL), 1 case of mantle cell lymphoma (MCL), 1 case of primary central nervous system lymphoma (PCNSL), and 3 cases of multiple myeloma (MM). The ctDNA extracted from the samples of each patient before and after treatment was subjected to MRD detection using the method in Example 1. The results are shown in Table 13.
[0304] Table 13 ctDNA sample detection results
[0305]
[0306] Example 7 Detection of fusion
[0307] Fusion genes are a type of gene abnormality form that was first identified in hematological malignancies and has clear clinical diagnostic and treatment significance. In the latest World Health Organization (WHO) 2016 edition of the classification criteria for hematopoietic and lymphoid tissue tumors, 112 FGs have been clearly listed, and most of them can be used as the basis for diagnosis and classification. The fusion genes included in this tracking were BCL1-IGH and BCL2-IGH.
[0308] Sanger was used to trace the frequency of the fusion gene BCL1-IGH in the JVM-2 cell line. According to the malignant cell frequencies traced from high to low, the gDNA of the JVM-2 cell line and the gDNA of healthy human whole blood leukocytes were diluted to the frequencies in Table 14 below. With an input amount of 20 μg DNA, each frequency standard was detected 3 times to preliminarily establish the LOD.
[0309] Table 14 Corresponding relationship between malignant cell frequency and number of malignant cells at different DNA input amounts
[0310]
[0311] Three different input amounts (disease type MCL), with each frequency standard detected three times, and the LOD was initially established at a 100% detection frequency.
[0312] Table 15 Detection results for the establishment of 20μg - LOD
[0313]
[0314] Example 8 Traceable positive clones, that is, cases of clone types that exist both before and after diagnosis and treatment
[0315] Four patients with known pathological information were selected, and samples were collected at different treatment time points. The MRD was detected using the method of the present invention. The detection results are shown in Table 16. Two significant clone types were detected in the pre - treatment sample of Patient 1. When tracking the sample at 3 months after treatment, the two significant clone types were not detected. When detecting the sample at 6 months after treatment, the two significant clone types before treatment were not traced as residues, but a new clone type was detected, and the clone cell content was 5.43E - 03. Three significant clone types were detected in Patient 2 before treatment. At 3 months after treatment, 10 -6 Frequency residues rebounded at 6 months after treatment, and a high concentration of clone residues was detected. Three significant clone types were detected in Patient 3 before treatment, and these three significant clone types were also traced at 3 months after treatment, but the clone cell content decreased. Two significant clone types were detected in Patient 4 before treatment. At 3 months after treatment, the clone cell content of one of the clone types decreased to 10 -6 , and there was no residue of the IGH - A clone type. There is no data for Patients 3 and 4 at 6 months after treatment.
[0316] Table 16 MRD monitoring results
[0317]
[0318] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A minimal residual monitoring model for human blood tumors, characterized in that, Comprising: A data acquisition module for receiving DNA sequencing data before and after treatment of the same detection object; A nucleated cell count acquisition module for aligning the sequencing data of the reference gene to the reference genome to obtain the nucleated cell count; A clone map and sequence acquisition module for aligning the DNA sequencing data to an immune database including IGH, IGK, and IGL genes through an alignment algorithm, and then obtaining a clone map and corresponding sequences through a clustering algorithm; A fusion genotype and sequence acquisition module for aligning the DNA sequencing data to the reference gene sequence to obtain a fusion genotype and corresponding sequences; A calculation module for calculating clone frequency, clone cell content, fusion frequency, and fusion cell content based on the nucleated cell count, clone map and corresponding sequences, fusion genotype and corresponding sequences, and determining significant clones and significant fusions based on the clone frequency, clone cell content, fusion frequency, and fusion cell content; An MRD calculation and determination module for screening effective clones and fusions based on the clones and fusions obtained before and after treatment of the same object, calculating the MRD value based on the clone cell content and fusion cell content corresponding to the effective clones and fusions, and determining the residual state of human blood tumors based on the MRD value; The clustering algorithm includes: Calculating the edit distance between sequences based on the DNA sequencing data; Performing clustering analysis on the sequenced sequences based on the edit distance and its relationship with the V, D, J regions of the IGH gene, the V, J regions of the IGK gene, and the V, J regions of the IGL gene to obtain a clone map; The calculation method of the edit distance is: Where seq1 and seq2 represent two sequenced sequences, i represents the i-th base in seq1, j represents the j-th base in seq2, both i and j start from 1 and go to the last position, and the initial values are both 0; ms represents the match score, when the i-th base of seq1 and the j-th base of seq2 are the same, the value of m(i,j) is positive; ss represents the mismatch score, when the i-th base of seq1 and the j-th base of seq2 are different, the value of m(i,j) is negative; os represents the score for the start of a gap, gapopening represents the start situation of a gap in sequence matching; es represents the score for the continuation of a gap; gap extending represents the continuation situation of a gap in sequence matching; S represents the clip score, when the bases of seq1 and seq2 are highly mismatched and the negative value is too high, truncation will be selected at this time; The values of ms, ss, os, es, and S satisfy the following formula: Where n is an integer ≥0 and less than the sequencing read length.
2. The minimal residual disease monitoring model for human hematological malignancies according to claim 1, wherein The clustering analysis includes: Based on the comparison result and the edit distance, the best matching method is given, and the sequencing sequences with the same V and J regions of IGH, IGK, and IGL genes are clustered. When IGH-DJ rearrangement occurs, the sequencing sequences with the same D and J regions of the IGH gene are clustered. When IGK-KDE rearrangement occurs, the sequencing sequences with the same V regions of the KDE gene and the IGK gene are clustered; In each VJ class, it is divided according to the CDR3 region consistency; in each DJ class, it is divided according to the 50bp-length region consistency centered on the junction position of the D and J regions of the IGH gene; in each KDE-V class, it is divided according to the 50bp-length region consistency centered on the junction position of the V region of the IGK gene and the KDE gene; the number of sequencing sequences corresponding to each subclass is obtained; The hierarchical clustering method is used to cluster each subclass. In each cluster, the child node and the parent node sequences are highly similar, and the parent node has a significantly higher number of sequencing sequences than the child node; The number of sequencing sequences of all nodes in each cluster is summed, and the root node sequence is reported as a single clonotype; The individual clonotypes reported are combined to form a clonal map.
3. The minimal residual disease monitoring model for human hematological malignancies according to claim 1, wherein The method for the fusion genotype and sequence acquisition module to obtain the fusion genotype and the corresponding sequence includes: The sequence numbers of the IGH, BCL1, and BCL2 genes are obtained respectively through the bwa alignment software; According to the reference sequences of the IGH, BCL1, and BCL2 genes, the FMDIndex data structure is constructed; The FMDIndex is used to find the set of maximum exact matches between each sequencing sequence and the reference sequence; The dynamic programming algorithm is used to find a set of optimal maximum exact match subsets between the reference sequence and the sequencing sequence; The edit distance algorithm is used to align the sequencing sequence to the reference gene sequence corresponding to the optimal exact match subset to obtain the fusion genotype and the corresponding sequence number; The partial-order alignment algorithm is used to obtain the fusion genotype sequence based on the sequencing sequence corresponding to the fusion genotype.
4. A minimal residual monitoring model for human hematological malignancies according to claim 1, wherein, The calculation methods of the clonal frequency and the clonal cell content are as follows: Among them, cfl represents the clonal frequency of a single IGL clonotype, count1 is the number of sequences of a single IGL clonotype, count IGL is the number of sequences of all IGL clonotypes, cfh represents the clonal frequency of a single IGH clonotype, count2 is the number of sequences of a single IGH clonotype, count IGH is the number of sequences of all IGH clonotypes except IGH-DJ clonotypes, count HDJ is the number of sequences of all IGH-DJ clonotypes; cfk represents the clonal frequency of a single IGK clonotype, count3 is the number of sequences of a single IGK clonotype, count IGK is the number of sequences of all IGK clonotypes except IGK-KDE clonotypes, count KDE is the number of sequences of all IGK-KDE clonotypes; the number of clonal cells is the number of sequences of a single clonotype.
5. A minimal residual monitoring model for human hematological malignancies according to claim 1, wherein, The calculation methods of the clonal frequency and the clonal cell content are as follows: clonetype represents the clonotype, and wf(clonetype) is the correction function; λ1 takes the value of 1.03, lengh(plasmid) is the length of the synthetic clonotype plasmid sequence, and length(clonetype) is the length of the clonotype sequence in the test sample; First, a series of plasmids with different combinations of the V regions and J regions of IGH, IGK, and IGL genes were artificially synthesized. When IGH-DJ rearrangement occurred, plasmids with different combinations of the D region and J region of the IGH gene were synthesized. When IGK-KDE rearrangement occurred, plasmids with different combinations of the KDE gene and the J region of the IGK gene were synthesized. After dissolving these plasmids, they were quantified and then randomly mixed in equal amounts to form 10 different plasmid mixture libraries. The 10 plasmid mixture libraries were respectively constructed into libraries and sequenced on the machine, and the obtained data were analyzed to obtain a clone map. For each plasmid, the number of reads was divided by the average number of reads of the plasmids in the plasmid library, and the result was denoted as value a.
6. A monitoring model for minimal residual disease of human hematological malignancies according to claim 1, characterized in that The calculation methods for the fusion frequency and the fusion cell content are as follows: Where fc is the number of sequences of a single fusion type, hc is the total number of IGH sequences, bc is the total number of BCL1 or BCL2 sequences, and the number of fusion cells is the number of sequences of a single fusion type.
7. A monitoring model for minimal residual disease of human hematological malignancies according to claim 1, characterized in that The calculation methods for the fusion frequency and the fusion cell content are as follows: fusion represents the fusion type, and wf(fusion) is the correction function; λ2 takes a value of 1.08, lengh(plasmid) is the length of the synthesized fusion plasmid sequence, and length(fusion) is the length of the fusion type sequence in the test sample; First, a series of plasmids with different combinations of the IGH gene and the BCL1 and BCL2 genes were artificially synthesized. After dissolving these plasmids, they were quantified and then randomly mixed in equal amounts to form 10 different plasmid mixture libraries. The 10 plasmid mixture libraries were respectively constructed into libraries and sequenced on the machine, and the obtained data were analyzed to obtain the fusion genotype. For each plasmid, the number of reads was divided by the average number of reads of the plasmids in the plasmid library, and the result was denoted as value a.
8. A monitoring model for minimal residual disease of human hematological malignancies according to claim 1, characterized in that The screening criteria for significant clones are as follows: (1) The clone frequency ≥ 0.03; (2) The clone cell content ≥ 0.002; (3) The distribution of clone frequencies is discontinuous, and the judgment criteria are as follows: (a) Take the clones with a frequency less than 0.03 and arrange them in descending order; (b) Take 10 times the clone frequency of the 5th clone in the ordered arrangement as the threshold; (c) Compare the clone frequencies of all clones with the threshold. Clones with a frequency greater than or equal to the threshold are considered to meet the requirement of discontinuous clone frequency distribution.
9. A minimal residual disease monitoring model for human hematological malignancies according to claim 1, characterized in that, The screening criteria for significant fusion types are as follows: (1) The fusion frequency ≥ 0.1%; (2) The fusion cell content ≥ 0.02%.
10. A monitoring model for minimal residual disease of human hematological malignancies according to claim 1, characterized in that The calculation method for MRD is: MRD 克隆 = max(sum(IGH clonal cell content), sum(IGK clonal cell content) + sum(IGL clonal cell content)); MRD 融合 = max(sum of BCL1-IGH fusion cell content, sum of BCL2-IGH fusion cell content); MRD = max(MRD 融合 , MRD 克隆 ); or MRD = MRD 融合 + MRD 克隆 。 11. A minimal residual disease monitoring model for human hematological malignancies according to claim 1, characterized in that, The model further includes: Obtaining the clones of healthy people and / or patients with non-homogeneous diseases and calculating the clone cell content of the clones. Based on the clonotypes of the healthy individuals and / or patients with non-homogeneous diseases, the significant clonotypes obtained from the pre-treatment DNA sequencing data are screened for those that are not rarely present and have a clonal cell content ≥ 10 in the healthy individuals and / or patients with non-homogeneous diseases -7 of the clonotypes, and the non-rare presence means clonotypes present in ≥ 5% of the population.
Citation Information
Patent Citations
Primer composition and kit for monitoring minimal residual disease of human leukemia and application of primer composition and kit
CN115927631A