Model for monitoring tiny residues of human hematologic tumors
Through a micro-residue monitoring model for human hematologic tumors, a one-step library building and a single pool amplification strategy were used to calculate cloning and fusion frequency, solving the challenges of existing MRD detection methods in terms of sensitivity, sample demand and cost, and achieving efficient and accurate MRD detection.
Patent Information
- Application Number
- CN202510412727.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
Existing MRD detection methods have challenges in terms of sensitivity, sample demand and cost, and it is difficult to effectively monitor tiny residual lesions in hematologic tumors.
A model for monitoring micro-residues of human blood tumors is provided, using one-step library construction and a single pool amplification strategy, and the cloning frequency and fusion frequency are calculated through clustering algorithms and calculation modules, significant cloning and fusion types are screened, and MRD values are calculated.
It significantly shortens the library construction time, reduces the sample demand, improves detection sensitivity, and can provide accurate MRD detection results in low abundance ctDNA.
Smart Images

Figure CN119943149A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of biotechnology, and in particular to a model for monitoring micro-residues of human blood tumors. Background Art
[0002] Hematological tumors are common malignant tumors with high morbidity and mortality rates, which seriously threaten human health. As hematological tumors enter the era of personalized precision treatment, MRD detection of hematological tumors has played an increasingly important role in the precise monitoring of tumor cells or clones, early intervention, and immune reconstruction evaluation.
[0003] MRD (minimal residual disease) refers to the small amount of cancer cells that remain in the body after cancer treatment, which may lead to disease recurrence. As a routine detection method for acute lymphoblastic leukemia (ALL), MRD is currently the most important prognostic indicator and can also be used to monitor the effectiveness of chemotherapy and biomarkers of drug resistance. Existing MRD detection methods include flow cytometry (MFC), PCR detection, and next-generation sequencing (NGS). MFC detection uses multicolor flow technology to track abnormal immune phenotypes, with a short turnaround time and a wide range of applications. However, its sensitivity is unstable, relies on professional interpretation, and the antigenic phenotype may change after chemotherapy, limiting its early prediction ability. NGS testing is an emerging high-sensitivity method that uses consensus primers to amplify rearranged fragments of immunoglobulin and T cell receptor genes. It does not require patient-specific customization and can detect clonal evolution and track MRD after treatment. Through efficient bioinformatics analysis, NGS can accurately identify specific clonal sequences with a sensitivity of up to 10 -5 ~10 -6 . However, NGS technology also has some challenges and limitations, including the complexity of the library construction process, the limitation of ctDNA sensitivity, the temporal and spatial heterogeneity of tumors, the limited independent prognostic value of IGK and IGL-MRD, and the high cost. Traditional NGS technology for detecting MRD usually adopts multi-step amplification and pooling library construction. These methods also have some defects in practical applications, such as cumbersome library construction process, two rounds of PCR amplification, large sample demand, easy contamination between samples, and high library construction cost. Summary of the invention
[0004] In order to solve the above problems, the present invention provides a model for monitoring micro-residues of human blood tumors.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows: The present invention provides a monitoring model for micro-residue of human blood tumors, comprising: a data acquisition module, used for receiving DNA sequencing data of the same detection object before and after treatment; a nucleated cell number acquisition module, used for comparing the sequencing data of the internal reference gene to the reference genome to obtain the nucleated cell number; a clone map and sequence acquisition module, used for comparing the DNA sequencing data to an immune database including IGH, IGK and IGL genes through an alignment algorithm, and then obtaining the clone map and the corresponding sequence through a clustering algorithm; a fusion genotype and sequence acquisition module, used for comparing the DNA sequencing data to the reference gene sequence, obtaining the fusion genotype and the corresponding sequence a calculation module for calculating the clone frequency, clone cell content, fusion frequency and fusion cell content based on the nucleated cell number, clone map and corresponding sequence, fusion genotype and corresponding sequence, and determining significant clone type and significant fusion type based on the clone frequency, clone cell content, fusion frequency and fusion cell content; an MRD calculation and determination module for screening effective clone type and fusion type according to the clone type and fusion type obtained before and after treatment of the same subject, and calculating the MRD value based on the clone cell content and fusion cell content corresponding to the effective clone type and fusion type, and determining the residual state of the human blood tumor based on the MRD value.
[0006] Furthermore, the clustering algorithm includes: calculating the edit distance between sequences based on the DNA sequencing data; clustering the sequenced sequences based on the edit distance and its relationship with the IGH gene V, D, J regions, IGK gene V, J regions and IGL gene V, J regions to obtain a cloning map.
[0007] Furthermore, the edit distance is calculated as follows: ; ; ; Among them, seq1 and seq2 represent the sequences obtained by two sequencing, i represents the ith base in seq1, j represents the jth base in seq2, i and j both start from 1 and go to the last digit, and the initial value is 0; ms represents the match score, when the ith base of seq1 and the jth base of seq2 are consistent, the m(i, j) value is a positive value; ss represents the mismatch score, when the ith base of seq1 and the jth base of seq2 are inconsistent, the m(i, j) value is a negative value; os represents the gap start score, gap opening represents the start of a gap in sequence matching; es represents the gap extension score, gap extending represents the extension of a gap in sequence matching; S represents the clip score, when the bases of seq1 and seq2 are highly mismatched, the negative value is too high, and truncation will be chosen at this time.
[0008] Furthermore, the cluster analysis includes: giving the best match according to the alignment result and the edit distance, clustering the sequencing sequences with the same V region and J region of IGH, IGK and IGL genes; when IGH-DJ rearrangement occurs, clustering the sequencing sequences with the same D region and J region of IGH genes; when IGK-KDE rearrangement occurs, clustering the sequencing sequences with the same V region of KDE genes and IGK genes; in each VJ class, dividing according to the consistency of CDR3 region; in each DJ class, clustering the centered on the junction position of D region and J region of IGH gene; 50bp length region consistency division; In each KDE-V class, the 50bp length region consistency division centered on the junction position of the V region of the IGK gene and the KDE gene; The number of sequencing sequences corresponding to each subclass is obtained; Each subclass is clustered using a hierarchical clustering method. In each cluster, the child node and parent node sequences are highly similar, and the parent node has a significantly higher number of sequencing sequences than the child node; The number of sequencing sequences of all nodes in each cluster is summed, and the root node sequence is reported as a separate clone type; Each reported separate clone type is combined to form a clone map.
[0009] Furthermore, the calculation method of the clone frequency and clone cell content is: ; ; ; ; Among them, cfl represents the clonal frequency of a single IGL clone, count1 is the sequence number of a single IGL clone, and count IGLis the sequence number of all IGL clones, cfh represents the clonal frequency of a single IGH clone, count2 is the sequence number of a single IGH clone, and count IGH is the number of sequences of all IGH clonotypes except IGH-DJ clonotype, count HDJ is the number of sequences of all IGH-DJ clones; cfk represents the clonal frequency of a single IGK clone, count3 is the number of sequences of a single IGK clone, and count IGK is the number of sequences of all IGK clonotypes except IGK-KDE clonotypes, count KDE is the sequence number of all IGK-KDE clones; the number of cloned cells is the sequence number of a single clone.
[0010] Furthermore, the calculation method of the fusion frequency and the fusion cell content is: ; ; Among them, fc is the number of sequences of a single fusion type, hc is the total number of IGH sequences, bc is the total number of BCL1 or BCL2 sequences, and the number of fusion cells is the number of sequences of a single fusion type.
[0011] Furthermore, the screening criteria for significant clonotypes are: (1) clonotype frequency ≥ 0.03; (2) clonotype cell content ≥ 0.002; (3) clonotype frequency distribution is discontinuous, and the judgment criteria are as follows: (a) clonotypes with frequencies less than 0.03 are selected and arranged in descending order; (b) 10 times the clonotype frequency of the fifth clonotype in the order is used as the threshold; (c) the clonotype frequencies of all clonotypes are compared with the threshold, and the clonotypes with frequencies greater than or equal to the threshold are considered to meet the clonotype frequency distribution discontinuity.
[0012] Furthermore, the screening criteria for significant fusion types are: (1) fusion frequency ≥ 0.1%; (2) fusion cell content ≥ 0.02%.
[0013] Furthermore, the MRD is calculated as follows: MRD 克隆 = max(sum(IGH clone cell content), sum(IGK clone cell content) + sum(IGL clone cell content)); MRD 融合 = max(sum BCL1-IGH fusion cell content), sum(BCL2-IGH fusion cell content)); ; or .
[0014] Furthermore, the model further comprises: obtaining clonotypes of healthy persons and / or patients with non-similar diseases, and calculating the content of clonotypes; based on the clonotypes of the healthy persons and / or patients with non-similar diseases, screening out the clonotypes of the healthy persons and / or patients with non-similar diseases that are not rare and have a clonotype content of ≥10 from the significant clonotypes obtained based on the DNA sequencing data before treatment; -7 The non-rare clonotypes are clonotypes present in ≥5% of the population and form the basis for subsequent MRD calculations.
[0015] The beneficial effects brought by the technical solution provided by the embodiment of the present invention include: The present invention proposes a simple library construction method, which adopts a one-step library construction method, and only one round of PCR amplification is required to obtain the library, which significantly shortens the library construction time. At the same time, through the single pool amplification strategy, the primers of genes such as IGH, IGK, IGL, KDE and IGH-BCL1 and IGH-BCL2 fusion sites are mixed in one pool, so as to achieve the simultaneous detection of immune receptor rearrangement and key fusion gene events in a single PCR reaction, which greatly reduces the sample demand and is particularly conducive to tracking the type of sample rearrangement after treatment. By introducing IGH-BCL fusion detection, the capture ability of specific molecular markers of blood tumors such as lymphoma is further enhanced. In terms of detection system construction, this scheme uses hundreds of highly specific primers to comprehensively cover all immune cell receptor chains (including IGH, IGL, IGK, IGH-DJ, IGK-KDE, etc.) and IGH-BCL fusion hotspots, breaking through the detection blind spots of traditional customized panels. By combining the optimized significant clone screening algorithm with the uniqueness index evaluation model, it is possible to accurately distinguish between malignant clones and physiological rearrangements, effectively solving the clinical pain point of limited prognostic value of IGK / IGL detection. Experimental verification shows that the minimum detection sensitivity of this technical solution reaches 10 -7 The sensitivity of ctDNA is a key challenge in MRD detection. The proportion of ctDNA to cfDNA is only 0.01%, and it is affected by many factors such as tumor type and load. The technical solution proposed in the present invention has excellent detection performance, can effectively solve the problem of low abundance of ctDNA, and provide accurate MRD detection. This method not only improves the detection accuracy, but also provides more comprehensive and in-depth tumor biology information for clinicians. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 A graph showing the relationship between the number of samples and the internal reference ratio provided in an embodiment of the present invention, wherein the horizontal axis is the number of samples and the vertical axis is the ratio of the corrected sequence numbers of internal reference 1 and internal reference 2; Figure 2 A schematic diagram of the quantitative results of purified DNA provided in an embodiment of the present invention; Figure 3 A schematic diagram of the best matching method of seq1 and seq2 provided in an embodiment of the present invention; Figure 4 A schematic diagram of layered clustering provided by an embodiment of the present invention; Figure 5 This is a schematic diagram of the DNA quantification results after purification of the ALL cell line with a loading amount of 20 μg provided in Example 3 of the present invention within the linear range; Figure 6 This is a schematic diagram of the DNA quantification results after purification of the CLL cell line with a loading amount of 20 μg provided in Example 3 of the present invention within the linear range; Figure 7 This is a schematic diagram of the DNA quantification results after linear range purification of the MM cell line with a sample load of 20 μg provided in Example 3 of the present invention. DETAILED DESCRIPTION
[0018] The present invention is further described in detail below by specific embodiments. However, it will be appreciated by those skilled in the art that the following examples are only used to illustrate the present invention and should not be considered as limiting the scope of the present invention. If specific techniques or conditions are not specified in the examples, they are carried out according to the techniques or conditions described in the literature in this area or according to the product instructions. If the manufacturer of the reagents or instruments used is not specified, they are all conventional products that can be obtained commercially.
[0019] As used herein, the words "comprises," "including," "having," or any other variations thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that includes listed elements is not necessarily limited to those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. The singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0020] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meanings as those generally understood by those skilled in the art. In addition to the specific methods, equipment, and materials used in the embodiments, any methods, equipment, and materials of the prior art that are similar or equivalent to the methods, equipment, and materials in the embodiments of the present invention may be used to implement the present invention, based on the prior art mastery of those skilled in the art and the description of the present invention.
[0021] Unless otherwise stated, the experimental methods, detection methods, and preparation methods disclosed in the present invention all adopt conventional techniques in the field of molecular biology, immunology, laboratory science, gene sequencing technology, bioinformatics technology, and related fields.
[0022] The embodiment of the present invention provides a model for monitoring micro-residues of human blood tumors, comprising: (1) A data acquisition module, used to receive DNA sequencing data of the same test subject before and after treatment; (2) A nucleated cell number acquisition module, which is used to align the sequencing data of the internal reference gene with the reference genome to obtain the nucleated cell number; (3) Clone map and sequence acquisition module, which is used to align DNA sequencing data to the immune database including IGH, IGK and IGL genes through an alignment algorithm, and then obtain the clone map and corresponding sequence through a clustering algorithm; (4) Fusion genotype and sequence acquisition module, used to align DNA sequencing data to the reference gene sequence to obtain the fusion genotype and the corresponding sequence; (5) a calculation module, for calculating the clonal frequency, clonal cell content, fusion frequency and fusion cell content based on the number of nucleated cells, the clonal map and the corresponding sequence, the fusion genotype and the corresponding sequence; and determining significant clonal types and significant fusion types based on the clonal frequency, clonal cell content, fusion frequency and fusion cell content; (6) An MRD calculation and determination module, which is used to screen effective clonal types and fusion types based on the clonal types and fusion types obtained before and after treatment of the same subject, calculate the MRD value based on the clonal cell content and fusion cell content corresponding to the effective clonal type and fusion type, and determine the residual status of human blood tumors based on the MRD value.
[0023] Specific, effective clonal and fusion screening includes: (1) Significant clonotypes and fusion types that exist in the same subject after treatment and are the same as those before treatment.
[0024] (2) The clonotype that exists in subsequent tests of the same subject and is the same as the newly discovered significant clonotype after treatment.
[0025] The newly discovered significant clonotype after treatment is a significant clonotype that exists after treatment and is different from the significant clonotype before treatment. Subsequent testing is performed to track the newly discovered significant clonotype after treatment.
[0026] In MRD detection, the sensitivity of ctDNA is a key challenge. The proportion of ctDNA to cfDNA is only 0.01%, and it is affected by multiple factors such as tumor type and load. Especially in early-stage cancer patients, low-abundance ctDNA may lead to MRD-negative false appearance. Based on the BCR database, through significant clone screening and uniqueness index evaluation, it is possible to determine whether it is a cancer cell clone, solving the problem of limited prognostic value of IGK and IGL-MRD. This method not only improves the detection accuracy, but also provides more comprehensive and in-depth tumor biology information for clinicians. It has excellent detection performance and the minimum detection limit can reach 10 -7 .
[0027] Specifically, the data acquisition module realizes DNA sequencing data of several identical test subjects before and after treatment by cooperating with the detection kit and the monitoring device. The same clone type and fusion type refer to the same gene sequence.
[0028] Specifically, the detection kit comprises: an upstream outer primer F1, an upstream inner primer F2 and a downstream outer primer R; the upstream outer primer F1 is composed of a sequencing adapter sequence 1, a barcode sequence and an upstream universal sequence 1, the adapter sequence 1 is shown in SEQ ID No.5, the barcode sequence is shown in SEQ ID No.6-15, and the upstream universal sequence 1 is shown in SEQ ID No.16; The upstream inner primer F2 is composed of an upstream universal sequence 2 and an upstream specific primer sequence, wherein the upstream universal sequence 2 is shown in SEQ ID No. 17, and the upstream specific primer sequence is shown in SEQ ID No. 18-163; The downstream outer primer R consists of a downstream specific primer sequence, a downstream adapter sequence and a downstream sequencing adapter sequence, and the primer is 5' phosphorylated. The downstream specific primer sequence is shown in SEQ ID NO.164-183, the downstream adapter sequence is shown in SEQ ID NO.184, and the downstream sequencing adapter sequence is shown in SEQ ID NO.185.
[0029] Optionally, the kit includes two internal reference genes, the upstream specific primer sequence of the first internal reference gene is shown in SEQ ID NO.1, and the downstream specific primer sequence of the first internal reference gene is shown in SEQ ID NO.2; the upstream specific primer sequence of the second internal reference gene is shown in SEQ ID NO.3, and the downstream specific primer sequence of the second internal reference gene is shown in SEQ ID NO.4.
[0030] 1. Selection of internal reference genes Specifically, considering the diversity of sample types, two amplicons were selected as internal references to ensure the stability of the total number of internal reference sequences. In the selection of internal references, regions with good sequence specificity were given priority for design. In addition, since the interval between pre-treatment samples and post-treatment samples was relatively long, in order to ensure that the same sample could be traced and compared before and after, three SNP sites were specially designed on the internal reference amplicon. Each site had three genotypes. Samples collected from the same individual at different time periods should have the same genotype at the three SNP sites, which significantly reduced the risk of sample confusion caused by sampling or experimental operations.
[0031] Two genes were selected as internal references, and hg19 was used as a reference. The specific information is as follows: (1) Internal reference 1 Position information chr20:15124894-15125033; upstream specific primer sequence GGCCTACAATTCAAATTAATGTAAAAACTGC (SEQ ID NO:1), upstream primer position is chr20:15124894-15124924; downstream primer position is chr20:15125006-15125033, downstream specific primer sequence TCTGATTCAAACTTTTCTTTTGTGGCTG (SEQ ID NO:2).
[0032] (2) Internal reference 2 Position information chr5:17374840-17374979; upstream specific primer sequence GCACTTTCAATCAACTGTGTTAGATTGA (SEQ ID NO:3), upstream primer position is chr5:17374840-17374867; downstream primer position is chr5:17374951-17374979, downstream specific primer sequence TATGATCCACATTGTATGGTTTTTAGGCA (SEQ ID NO:4).
[0033] The sequence specificity of the primer sequence of the internal reference gene is good on the genome, and there is no amplification of pseudogene sequences. After the whole genome alignment, the internal reference 1 and the internal reference 2 each have a 140bp fragment that can be aligned 100%, and the position is in line with expectations. The other two fragments that can be aligned are only partial regions of the amplicon, and the size does not exceed 30bp. It is believed that the probability of not being able to amplify is high, as shown in Table 1.
[0034] Table 1 Internal reference comparison results
[0035] The internal reference gene contains 3 SNP sites, as shown in Table 2, which are stably present in normal and tumor cells and are independent of the cell cycle and whether the cells are activated. They are not affected by any endogenous or exogenous factors (such as any experimental treatment measures). At the same time, they can be used as pairing information for samples before and after treatment to prevent sample contamination.
[0036] Table 2 Internal reference SNP site information
[0037] 136 clinical samples of hematological tumors were selected for library construction, purification, quantification, mixed sample dilution, sequencing and data analysis according to the following steps, and the ratio of the sequence number of internal reference 1 to the sequence number of internal reference 2 was calculated, such as Figure 1 As shown, the results show that the ratio of sequence numbers after correction with the two internal references is relatively stable.
[0038] 2. Library Construction Different samples use different specific adapters; gDNA from the same sample is mixed into a pool and uses the same specific adapter (upstream outer primer); samples are divided into pre-treatment samples and post-treatment samples, and the starting amounts of the two are different.
[0039] Take an equal amount of gDNA from each clinical sample to build a library in a single tube, take a 0.2 mL PCR tube with the number of clinical samples + 1 (control), and prepare the reaction system according to Table 3. The upstream outer primer: upstream inner primer mixture: downstream primer mixture = 5: 3: 20.
[0040] Table 3 Amplification system
[0041] Primer mix preparation: The primer sequences consist of an upstream outer primer F1, an upstream inner primer F2 and a downstream outer primer R, as shown in Tables 4 and 5.
[0042] The upstream external primer F1 consists of sequencing adapter sequence 1 (CAAGCAGAAGACGGCATACGAGAT SEQ ID NO: 5), Barcode sequence (SEQ ID NO: 6-15), and upstream universal sequence 1 (GGCACCCGAGAATTCCA SEQ ID NO: 16).
[0043] The upstream internal primer F2 (i.e., the internal reference upstream primer, IGH-V, IGH-D upstream primer, IGK-V, IGKED upstream primer and IGL-V upstream primer, fusion upstream primer) consists of the upstream universal sequence 2 (CGAACCGCTTTCGGAGG SEQ ID NO: 17) and the upstream specific primer sequence (SEQ ID NO: 18-163).
[0044] The downstream outer primer R (i.e., the internal reference downstream primer, IGH-J downstream primer, IGK-J, IGKDE downstream primer, and IGL-J downstream primer) consists of a downstream specific primer sequence (SEQ ID NO: 164-183), a downstream adapter sequence (TCAGAGTTCTACAGTCCGACGATC SEQ ID NO: 184), and a downstream sequencing adapter sequence (AATGATACGGCGACCACCGAGATCTACACGT SEQ ID NO: 185), and the primer is 5' phosphorylated.
[0045] Dilute the upstream outer primer F1 one by one to 100 μM; The upstream inner primers F2 were diluted one by one to 100 μM, and the upstream inner primers MIX1 (i.e., the mixture of the internal reference and IG, fusion upstream primers) were obtained by mixing them in equal amounts; The downstream outer primers R were diluted one by one to 100 μM, and mixed according to equal amounts of substances to obtain the downstream primer MIX2 (i.e., the mixture of the internal reference and IG downstream primers).
[0046] The PCR amplification reaction conditions are shown in Table 6.
[0047] Table 4 Barcode sequence information
[0048] Table 5 Specific primer sequence information
[0049] Table 6 PCR amplification reaction conditions
[0050] 3. Library Purification The amplified product obtained in step 2 was purified using the MPure XP-PCR product purification kit (Beckman Coulter, A63882) to obtain a purified DNA library. The purified DNA library solution can be stored at 4°C for no more than 16 hours. If it needs to be stored for a longer period of time, it should be stored below -18°C.
[0051] IV. Library Quantification The concentration and band size of the purified DNA library were quantified using TapeStation2200 / 4200. The results of the TapeStation2200 quantification are as follows: Figure 2 .
[0052] Agilent2200 / 4200 quality control: After magnetic bead purification, libraries with a target fragment between 220-400bp and a dimer band (less than 200bp) of no more than 20% are qualified and can be directly loaded onto the machine. Libraries with a dimer band of more than 20% after purification are unqualified and need to be rebuilt.
[0053] 5. Library mixing and dilution (1) All DNA libraries screened in step 4 are mixed together according to the mass ratio and marked as the machine library mixture.
[0054] (2) The total concentration of the target fragments in the 220-400 bp region in the library mixture is the library concentration. The recommended data volume for each sample is 5G.
[0055] Note: The mixed library needs to be vortexed for 1 minute and centrifuged for 2 seconds. Repeat this step at least 3 times until the mixed library is mixed. Make sure that the library is mixed evenly before performing quantitative detection.
[0056] 6. Sequencing After the prepared library mixture is quantified by Qubit, it is sequenced on the GENETRONS2000 platform according to the established data volume requirements. The data off the machine is the fastq file of PE150, which is the raw sequencing data.
[0057] The data acquisition module provided by the present invention for use in the human leukemia minimal residual monitoring model is used to acquire raw sequencing data, that is, to receive DNA sequencing data of the same test subject before and after treatment.
[0058] 7. Data Analysis It should be pointed out that the nucleated cell number acquisition module is used to align the sequencing data (sequencing sequence) of the internal reference gene (chr5:17374840-17374979, chr20:15124894-15125033, human hg19 genome position) to the reference genome (hg19, human hg38) to obtain the nucleated cell number, which is an existing technology and will not be described in detail here.
[0059] The clone map and sequence acquisition module is used to align the DNA sequencing data to the immune database including IGH, IGK and IGL genes through an alignment algorithm, and then obtain the clone map and the corresponding sequence through a clustering algorithm.
[0060] Specifically, the sequencing data were aligned to the immune database (source: IMGT http: / / www.imgt.org) including IGH, IGK and IGL genes through an alignment algorithm (reference https: / / academic.oup.com / nar / article / 41 / 10 / e108 / 1075719), and then a clonal map was obtained through a clustering algorithm. The clonal map includes but is not limited to IGH, IGK, IGL, IGH-DJ, IGK-KDE type clones or subsets of these types.
[0061] The clustering algorithm first calculates the edit distance between sequences, as follows: ; Among them, editdist represents the edit distance between sequences, which is a function related to the sequence. seq1 and seq2 represent two sequences obtained by sequencing.
[0062] The edit distance can be calculated by the following formula: ; Among them, i represents the i-th base in seq1, j represents the j-th base in seq2, i and j both start from 1 and go to the last digit, and their initial values are both 0.
[0063] ; Where ms represents the match score (positive number). When the ith base of seq1 and the jth base of seq2 are consistent, the m(i, j) value is positive. ss represents the mismatch score (negative number). When the ith base of seq1 and the jth base of seq2 are inconsistent, the m(i, j) value is negative.
[0064] ; Among them, os represents the score of the beginning of the gap (a negative number), gap opening represents the beginning of the gap in sequence matching, es represents the score of the gap extension (a negative number), gap extending represents the extension of the gap in sequence matching. Generally, os is greater than es.
[0065] S represents the clip score (negative number). When the base heights of seq1 and seq2 do not match, the negative value is too high, and truncation will be chosen.
[0066] Get the best match, see Figure 3 .
[0067] In order to better cluster the subsequent data, the values of ms, ss, os, es and S satisfy the following formula: ; Here, n is an integer ≥ 0 and is less than the sequencing read length. The sixth inequality in the above formula indicates that there must be a limit n such that the absolute value of the soft truncation (S) must be less than the gap opening (os) plus n gap extending (es) plus 2.
[0068] Preferably, ms can take the value of 1, and other values can be integer or negative multiples of ms.
[0069] Through the above restrictions, weights can be assigned to various situations, which is conducive to the clustering of subsequent data and avoids unclear boundaries and confusion of parent nodes in the clustering process. For example, ms takes the value of 1, ss takes the value of -4, os takes the value of -5, es takes the value of -1, and S takes the value of -6.
[0070] Then, the sequenced sequences were clustered according to the edit distance and the relationship between them and the V, D, J regions of the IGH gene, the V, J regions of the IGK gene, and the V, J regions of the IGL gene to obtain a clonal map. The specific process is as follows: (1) The best matching method is given according to the alignment results and edit distance of the alignment algorithm, that is, the edit distance score is maximized, and the sequencing sequences with the same V and J regions of the IGH, IGK and IGL genes are clustered. When IGH-DJ rearrangement occurs, the sequencing sequences with the same D and J regions of the IGH genes are clustered. When IGK-KDE rearrangement occurs, the sequencing sequences with the same V regions of the KDE and IGK genes are clustered.
[0071] (2) In each VJ class, the division was based on the consistency of the CDR3 region; in each DJ class, the division was based on the consistency of the 50 bp region centered on the junction of the D region and the J region of the IGH gene; in each KDE-V class, the division was based on the consistency of the 50 bp region centered on the junction of the V region of the IGK gene and the KDE gene; and the number of sequencing sequences corresponding to each subclass was obtained.
[0072] (3) A hierarchical clustering method is used to cluster each sub-class. In each cluster, the sequences of the child nodes and parent nodes are highly similar, and the parent node has a significantly higher number of sequenced sequences than the child node.
[0073] Specifically, it includes: obtaining the average error rate of sequencing bases, which is generally an empirical constant related to the sequencing platform, etc., and evaluating the subclass coefficient according to the error rate (hereinafter referred to as factor, which is greater than 1). In each category (VJ, DJ and KDE-V), the sequencing numbers are arranged from high to low. According to the sorting results and the edit distance algorithm, the high sequence number category and other subclasses are compared in turn, and the hierarchical clustering method is used for clustering. In each cluster, the child node and parent node sequences are highly similar, and the parent node has a significantly higher number of sequencing sequences than the child node, that is, the number of parent node sequences divided by the number of child node sequences is greater than factor.
[0074] (4) Sum the number of sequenced sequences of all nodes in each cluster, and report the sequence of the root node (top parent node) as a separate clone type.
[0075] (5) Combine each reported individual clonal type to form a clonal map.
[0076] Figure 4 This is a schematic diagram of the hierarchical clustering tree. Clonetype1 is the parent node of clonetype2 and clonetype3, clonetype4 and clonetype5 are the child nodes of clonetype2, and clonetype6 is the child node of clonetype3.
[0077] The fusion genotype and sequence acquisition module is used to align the DNA sequencing data to the reference gene sequence to obtain the fusion genotype and the corresponding sequence.
[0078] Specifically, they include: (1) The sequence numbers of IGH, BCL1 and BCL2 genes were obtained using the bwa alignment software.
[0079] (2) Based on the reference sequences of IGH, BCL1, and BCL2 genes, the FMDIndex data structure was constructed (https: / / pmc.ncbi.nlm.nih.gov / articles / PMC3389770 / ); Use FMD Index to find the maximum set of exact matches between each sequence and the reference sequence; A dynamic programming algorithm is used to find a set of optimal maximum exact matching subsets between the reference sequence and the sequencing sequence; The edit distance algorithm was used to align the sequencing sequence (seq1) to the reference gene sequence (seq2) corresponding to the best exact match subset to obtain the fusion type and the corresponding sequence number.
[0080] The partial-order alignment algorithm is used to obtain the fusion sequence based on the sequencing sequence corresponding to the fusion type.
[0081] A calculation module is used to calculate the clonal frequency, clonal cell content, fusion frequency and fusion cell content based on the nucleated cell number, clonal map and corresponding sequence, fusion genotype and corresponding sequence; and determine the significant clonal type and significant fusion type based on the clonal frequency, clonal cell content, fusion frequency and fusion cell content.
[0082] Specifically, the calculation method of the clone frequency and clone cell content is: ; ; ; ; Among them, cfl represents the clonal frequency of a single IGL clone, count1 is the sequence number of a single IGL clone, and count IGL is the number of sequences of all IGL clones, cfh represents the clonal frequency of a single IGH clone, count2 is the number of sequences of a single IGH clone, and count IGH is the number of sequences of all IGH clonotypes except IGH-DJ clonotype, count HDJ is the number of sequences of all IGH-DJ clones; cfk represents the clonal frequency of a single IGK clone, count3 is the number of sequences of a single IGK clone, and count IGK is the number of sequences of all IGK clonotypes except IGK-KDE clonotypes, count KDE is the sequence number of all IGK-KDE clones; the number of cloned cells is the sequence number of a single clone.
[0083] The calculation method of the fusion frequency and fusion cell content is: ; ; Among them, fc is the number of sequences of a single fusion type, hc is the total number of IGH sequences, bc is the total number of BCL1 or BCL2 sequences, and the number of fusion cells is the number of sequences of a single fusion type.
[0084] Preferably, the clonal frequency and clonal cell content are calculated as follows: ; ; ; ; clonetype represents the clonal type, wf (clonetype) is the correction function.
[0085] The calculation method of the fusion frequency and fusion cell content is: ; ; fusion Indicates fusion type, wf(fusion) is the correction function.
[0086] It should be pointed out that wf (clonetype) is a function related to the clonal amplification efficiency and the clonal amplification length; wf(fusion) It is a function related to the fusion amplification efficiency and the fusion amplification length.
[0087] The calculation method of clonal expansion efficiency is: First, a series of plasmids (i.e., clonal types) with permutations of the V and J regions of the IGH, IGK, and IGL genes are artificially synthesized. When IGH-DJ rearrangement occurs, a plasmid with permutations of the D and J regions of the IGH gene is synthesized. When IGK-KDE rearrangement occurs, a plasmid with permutations of the KDE gene and the J region of the IGK gene is synthesized. These plasmids are dissolved and quantified, and then randomly mixed into 10 different plasmid mixed libraries with equal mass. The 10 plasmid mixed libraries are constructed and sequenced separately, and the obtained data are analyzed to obtain a clonal map. For each plasmid, the number of reads is divided by the average number of reads of the plasmid in the plasmid library, recorded as the value a, and the amplification coefficient is calculated based on a. b(clonetype) .
[0088] ; The calculation method of fusion amplification efficiency is: First, a series of plasmids containing IGH genes and BCL1 and BCL2 genes were artificially synthesized (i.e., fusion types). These plasmids were dissolved and quantified, and then randomly mixed into 10 different plasmid mixed libraries of equal mass. The 10 plasmid mixed libraries were constructed and sequenced separately, and the obtained data were analyzed to obtain the fusion genotype. For each plasmid, the number of reads was divided by the average number of reads of the plasmid in the plasmid library, recorded as value a, and the amplification coefficient was calculated based on a. b (fusion).
[0089] ; After obtaining the amplification coefficient, we further calculate the correction function according to the amplification length, as follows: ; Among them, λ1 and λ2 take fixed values, preferably, λ1 takes a value of 1.03, and λ2 takes a value of 1.08. lengh(plasmid) is the length of the synthetic clone / fusion (plasmid) sequence, length(clonetype) To detect the length of the clonal sequence in the sample, length(fusion) To detect the length of the fusion sequence in the sample.
[0090] After obtaining the above data, significant clonotypes and significant fusion types were screened. The screening criteria for significant clonotypes were: (1) Clonal frequency ≥ 0.03; (2) Clonal cell content ≥ 0.002; (3) The clone frequency distribution is discontinuous. The judgment criteria are as follows: (a) Take the clonal types with a frequency less than 0.03 and arrange them in descending order; (b) The threshold is 10 times the clonal frequency of the fifth clonotype in the sequence; (c) The clonal frequencies of all clonal types are compared with the threshold, and the clonal types that are greater than or equal to the threshold are considered to satisfy the discontinuity of clonal frequency distribution.
[0091] The significant clonotype can be any one or more combinations and subsets of IGH, IGK, IGL, IGH-DJ, IGK-KDE.
[0092] The screening criteria for significant fusion type are: (1) Fusion frequency ≥ 0.1%; (2) Fusion cell content ≥ 0.02%.
[0093] Sample DNA is extracted from the same test subject before and after treatment to obtain significant clonal types and significant fusion types before treatment and clonal types and fusion types after treatment. The treatment includes but is not limited to surgical treatment, immunotherapy, targeted therapy, chemotherapy or radiotherapy.
[0094] The effective clonal type is screened out according to the clonal type obtained after treatment of the same test subject, and the clonal cell content of the effective clonal type is recorded, thereby calculating the MRD 克隆 .
[0095] Specific and effective clonal type screening methods include: (1) The clonotype that exists after treatment in the same subject and is the same as the significant clonotype before treatment.
[0096] (2) The clonotype that exists in subsequent tests of the same subject and is the same as the newly discovered significant clonotype after treatment.
[0097] The newly discovered significant clonotype after treatment is a significant clonotype that exists after treatment and is different from the significant clonotype before treatment. Subsequent testing is performed to track the newly discovered significant clonotype after treatment.
[0098] MRD 克隆 The values are calculated by tracking heavy chains (IGH) and light chains (IGK, IGL).
[0099] a. When only heavy chain (IGH) clones are traced; MRD 克隆 = sum(IGH clone cell content); b. When only light chain (IGK, IGL) clones are traced; MRD 克隆 = sum(IGK clone cell content) + sum(IGL clone cell content); c. When both heavy chain and light chain clones are tracked; MRD 克隆 = max(sum (IGH clone cell content), sum(IGK clone cell content) + sum(IGL clone cell content)) That is, the maximum value between the heavy chain and the light chain represents MRD 克隆 .
[0100] The effective fusion type is screened out according to the fusion type obtained after treatment of the same test subject, and the fusion cell content of the effective fusion type is recorded, thereby calculating the MRD 融合 The effective fusion type is the fusion type that exists in the same test subject after treatment and is the same as the significant fusion type before treatment.
[0101] a. Only when BCL1-IGH fusion is tracked; MRD 融合 = sum(BCL1-IGH fusion cell content); b. Only when BCL2-IGH fusion is tracked; MRD 融合 = sum(BCL2-IGH fusion cell content); c. When BCL1-IGH and BCL2-IGH fusions are tracked simultaneously; MRD 融合 = max(sum (BCL1-IGH fusion cell content), sum (BCL2-IGH fusion cell content)); MRD is calculated as: MRD 克隆 = max(sum(IGH clone cell content), sum(IGK clone cell content) + sum(IGL clone cell content)); MRD 融合 = max(sum (BCL1-IGH fusion cell content), sum (BCL2-IGH fusion cell content)); ; or .
[0102] When MRD>0, it is judged that there is residual human blood tumor, and when MRD =0, it is judged that there is no residual human blood tumor.
[0103] For example, before treatment, significant clonal types IGH-A (clonal cell content 3.64E-03) and IGH-B (clonal cell content 8.25E-03) were detected in hematological malignancy patient 1, and clonal type IGH-A (clonal cell content 2.58E-06) was detected 3 months after treatment. In this case, clonal type IGH-A is the effective clonal type screened by this test, and MRD = MRD 克隆 =2.58E-06>0, then the result of the test 3 months after treatment is that there is residual disease.
[0104] For example, in hematological malignancy patient 2, significant fusion BCL1-IGHJ (fusion cell content 5.79E-04) was detected before treatment, and fusion BCL1-IGHJ (fusion cell content 6.38E-06) was detected 3 months after treatment. Therefore, fusion BCL1-IGHJ is the effective fusion type screened by this test, and MRD = MRD 融合 =6.38E-06>0, then the result of the test 3 months after treatment is that there is residual disease.
[0105] For example, before treatment, significant clonal types IGH-A (clonal cell content 6.42E-02) and IGH-B (clonal cell content 4.12E-03) were detected in hematological malignancy patient 3. Three months after treatment, no clonal types identical to significant clonal types IGH-A and IGH-B were detected, but a newly discovered significant clonal type IGH-C (clonal cell content 5.35E-03) was detected. MRD=0, so the result of the test three months after treatment was that there was no residual disease. The patient took a new treatment plan, and continued to be tested three months after the new treatment plan. Only clonal type IGH-C (clonal cell content 4.38E-04) was detected. Therefore, clonal type IGH-C was the effective clonal type screened by this test, and MRD= MRD 克隆 =4.38E-04>0, then the result of the test 3 months after the new treatment is that there is residual disease.
[0106] The model further comprises: obtaining clonotypes of healthy persons and / or patients with non-similar diseases, and calculating the clonotype content; based on the clonotypes of healthy persons and / or patients with non-similar diseases, screening out the clonotypes of healthy persons and / or patients with non-similar diseases that are not rare and have a clonotype content of ≥10 from the significant clonotypes obtained based on DNA sequencing data before treatment; -7 The clonotypes formed the basis for subsequent MRD calculations.
[0107] The above method can be run in the calculation module or in the MRD calculation and determination module. By obtaining the clonal type of healthy people and / or patients with non-similar diseases, non-rare clones with a clone cell content of ≥10 -7 In addition, for the same test subject, a new significant clonotype may be generated after treatment, and the significant clonotype is different from the significant clonotype before treatment, but the new significant clonotype needs to be included in the subsequent MRD calculation.
[0108] An embodiment of the present invention further provides a computer storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the following method is implemented, wherein the method includes: Receive DNA sequencing data of the same test subject before and after treatment; The sequencing data of the internal reference gene was aligned to the reference genome to obtain the number of nucleated cells; The DNA sequencing data were aligned to the immune database including IGH, IGK and IGL genes through an alignment algorithm, and then the clone map and corresponding sequence were obtained through a clustering algorithm; Align the DNA sequencing data to the reference gene sequence to obtain the fusion genotype and the corresponding sequence; Calculate the clone frequency, clone cell content, fusion frequency and fusion cell content based on the nucleated cell number, clone map and corresponding sequence, fusion genotype and corresponding sequence; and determine significant clone type and significant fusion type based on the clone frequency, clone cell content, fusion frequency and fusion cell content; Based on the clonal types and fusion types obtained before and after treatment of the same subject, effective clonal types and fusion types are screened, and the MRD value is calculated based on the clonal cell content and fusion cell content corresponding to the effective clonal types and fusion types.
[0109] Furthermore, the clustering algorithm includes: calculating the edit distance between sequences based on the DNA sequencing data; clustering the sequenced sequences based on the edit distance and its relationship with the IGH gene V, D, J regions, IGK gene V, J regions and IGL gene V, J regions to obtain a cloning map.
[0110] Furthermore, the edit distance is calculated as follows: ; ; ; Among them, seq1 and seq2 represent the sequences obtained by two sequencing, i represents the ith base in seq1, j represents the jth base in seq2, i and j both start from 1 and go to the last digit, and the initial value is 0; ms represents the match score, when the ith base of seq1 and the jth base of seq2 are consistent, the m(i, j) value is a positive value; ss represents the mismatch score, when the ith base of seq1 and the jth base of seq2 are inconsistent, the m(i, j) value is a negative value; os represents the gap start score, gap opening represents the start of a gap in sequence matching; es represents the gap extension score, gap extending represents the extension of a gap in sequence matching; S represents the clip score, when the bases of seq1 and seq2 are highly mismatched, the negative value is too high, and truncation will be chosen at this time.
[0111] Furthermore, the cluster analysis includes: giving the best match according to the comparison result and the edit distance, clustering the same sequencing sequences of the V region and J region of the IGH, IGK and IGL genes; when IGH-DJ rearrangement occurs, clustering the same sequencing sequences of the D region and J region of the IGH gene; when IGK-KDE rearrangement occurs, clustering the same sequencing sequences of the J region of the KDE gene and the IGK gene; in each VJ class, dividing according to the consistency of the CDR3 region; in each DJ class, clustering the same sequencing sequences of the D region and J region of the IGH gene as the center; 50bp length region consistency division; In each KDE-V class, the 50bp length region consistency division centered on the junction position of the V region of the IGK gene and the KDE gene; The number of sequencing sequences corresponding to each subclass is obtained; Each subclass is clustered using a hierarchical clustering method. In each cluster, the child node and parent node sequences are highly similar, and the parent node has a significantly higher number of sequencing sequences than the child node; The number of sequencing sequences of all nodes in each cluster is summed, and the root node sequence is reported as a separate clone type; Each reported separate clone type is combined to form a clone map.
[0112] Further, specifically, the calculation method of the cloning frequency and the cloning cell content is: ; ; ; ; Among them, cfl represents the clonal frequency of a single IGL clone, count1 is the sequence number of a single IGL clone, and count IGL is the number of sequences of all IGL clones, cfh represents the clonal frequency of a single IGH clone, count2 is the number of sequences of a single IGH clone, and count IGH is the number of sequences of all IGH clonotypes except IGH-DJ clonotype, count HDJ is the number of sequences of all IGH-DJ clones; cfk represents the clonal frequency of a single IGK clone, count3 is the number of sequences of a single IGK clone, and count IGK is the number of sequences of all IGK clonotypes except IGK-KDE clonotypes, count KDE is the sequence number of all IGH-KDE clonotypes; the number of clonal cells is the sequence number of a single clonotype.
[0113] The calculation method of the fusion frequency and fusion cell content is: ; ; Among them, fc is the number of sequences of a single fusion type, hc is the total number of IGH sequences, bc is the total number of BCL1 or BCL2 sequences, and the number of fusion cells is the number of sequences of a single fusion type.
[0114] Preferably, the clonal frequency and clonal cell content are calculated as follows: ; ; ; ; clonetype represents the clonal type, wf (clonetype) is the correction function.
[0115] The calculation method of the fusion frequency and fusion cell content is: ; ; fusion Indicates fusion type, wf (fusion) is the correction function.
[0116] Furthermore, the screening criteria for significant clonotypes are: (1) Clonal frequency ≥ 0.03; (2) Clonal cell content ≥ 0.002; (3) The clone frequency distribution is discontinuous. The judgment criteria are as follows: (a) Take the clonal types with a frequency less than 0.03 and arrange them in descending order; (b) The threshold is 10 times the clonal frequency of the fifth clonotype in the sequence; (c) The clonal frequencies of all clonal types are compared with the threshold, and the clonal types that are greater than or equal to the threshold are considered to satisfy the discontinuity of clonal frequency distribution.
[0117] Furthermore, the screening criteria for significant fusion are: (1) Fusion frequency ≥ 0.1%; (2) Fusion cell content ≥ 0.02%.
[0118] Furthermore, the MRD is calculated as follows: MRD 克隆 = max(sum(IGH clone cell content), sum(IGK clone cell content) + sum(IGL clone cell content)); MRD 融合= max(sum (BCL1-IGH fusion cell content), sum (BCL2-IGH fusion cell content)); ; or .
[0119] Furthermore, the model further comprises: obtaining clonotypes of healthy persons and / or patients with non-similar diseases, and calculating the clonotype content thereof; based on the clonotypes of the healthy persons and / or patients with non-similar diseases, screening out the clonotypes of the healthy persons and / or patients with non-similar diseases that are not rare and have a clonotype content of ≥10 from the significant clonotypes obtained based on the DNA sequencing data before treatment; -7 The non-rare existence is a clonal type that exists in ≥5% of the population.
[0120] Example 1 Detection of minimal residual disease in human leukemia and lymphoma The 13 post-treatment samples with known pathological information were all from Beijing PanGenomics Gene Technology Co., Ltd. and its cooperative hospitals, and 1 cell line was purchased commercially, including 4 cases of acute lymphoblastic leukemia (ALL) (sample 1 and sample 2 were bone marrow gDNA; sample 3 and sample 4 were peripheral blood gDNA), 4 cases of chronic lymphocytic leukemia (CLL) (sample 5 and sample 6 were bone marrow gDNA; sample 7 and sample 8 were peripheral blood gDNA), 3 cases of follicular lymphoma (FL) (sample 9, sample 10, sample 11 were bone marrow gDNA), 2 cases of multiple myeloma (MM) (sample 12 and sample 13 were bone marrow gDNA), and 1 case of mantle cell lymphoma (MCL) (sample 14 was cell line gDNA).
[0121] Bone marrow DNA, peripheral blood DNA or cell line DNA extracted from 13 patients before and after treatment were used as starting samples, and MRD detection results were obtained according to the technical solution proposed in the present invention.
[0122] Table 7 Sample test results
[0123] Within the coverage of this test, the detection of different clone types finally showed MRD positive. Above, the detection method provided by the present invention is used to detect positive cases of human leukemia and lymphoma, and tracking MRD is effective, and the detection of clone types is well performed, indicating that the detection method has practical applicability. Among them, IGH detected positive clones covering 10 positive samples of four pathological types of ALL, CLL, MM, and MCL, IGK detected positive clones covering 11 positive samples of five pathological types of ALL, CLL, FL, MM, and MCL, IGL detected positive clones covering 6 positive samples of four pathological types of CLL, FL, MM, and MCL, HDJ detected positive clones covering 4 positive samples of three pathological types of ALL, CLL, and MM, KDE detected positive clones covering 6 positive samples of four pathological types of ALL, CLL, FL, and MM, and BCL-IGHJ rearrangement detected positive clones covering 6 positive samples of four pathological types of ALL, CLL, FL, and MCL. It can be seen that the present invention can effectively detect MRD for human leukemia, lymphoma, and myeloma.
[0124] Example 2 Sample consistency detection 16 patients with known pathological information were selected, of which 6 selected pre-treatment samples, 5 selected post-treatment samples, and 5 selected both pre-treatment and post-treatment samples, for a total of 21 samples tested. Samples 16 and 17, 18 and 19, 22 and 23, 25 and 26, 28 and 29 were from the same patient, and the method of the present invention and MFC were used for clonal type detection. The test results of all 21 tests included 14 positive results and 7 negative results; covering 11 pre-treatment samples and 10 post-treatment samples; including 16 bone marrow samples and 5 peripheral blood samples; the pathological types covered included acute lymphocytic leukemia (ALL), chronic lymphocytic leukemia (CLL), and multiple myeloma (MM). It is intended to evaluate the performance of the present invention in detecting MRD from multiple dimensions.
[0125] The specific results are shown in Table 8, where "NGS results" are the test results using the method of the present invention, and "flow results" are the test results using multicolor flow cytometry. The results show that the two are consistent in the two sample types of bone marrow and peripheral blood tested, three different pathological types, regardless of whether the samples are before or after treatment.
[0126] Table 8 Summary of test results of 21 clinical samples
[0127] *a: Both MRD detection results of the two methods are positive, b: Both MRD detection results of the two methods are negative, c: Both MRD detection results of the two methods are negative, wherein the method of the present invention detects a clone, but because it exists in the clonal type of healthy people or non-similar patients, it is not a disease-specific clone and is not tracked.
[0128] Example 3 Sensitivity study 1. Linear range of 20μg loading: MFC was used to trace the malignant cell content of ALL cell lines (RS4;11), CLL cell lines (MEC-1), and MM cell lines (IM-9). The gDNA of the three cell lines and the gDNA of healthy human whole blood leukocytes were diluted according to the traced malignant cell frequency. According to the input amount of 20μg nucleic acid, the number of cells for initial input detection was calculated, and the concentrations of malignant cells were set to 1E-06, 1E-05, 1E-04, 1E-03, and 1E-02, respectively. Linear reference materials were configured. Cell number = initial nucleic acid input (pg) / 6.6pg / cell; the configured gradient dilution samples were tested three times according to the method of the present invention to investigate the linear range of the method of the present invention.
[0129] When the sample volume was 20 μg, the disease types detected were: ALL, CLL, MM, and the detection range was the above-mentioned gradient reference. The theoretical frequency and MRD value were taken as -lg for statistics. The data are shown in Table 9 and Figure 5-7 As shown, the dilution linear R of the three disease cell lines 2 ≥0.95, indicating that the detection of the present invention has good linearity when the sample loading amount is 20 μg.
[0130] Table 9 Linear relationship results of 20 μg sample loading
[0131] 2. Establishment of LOD for 20μg loading: MFC was used to trace the malignant cell content of ALL cell lines (RS4;11), CLL cell lines (MEC-1), and MM cell lines (IM-9). The gDNA of the three cell lines and the gDNA of healthy human whole blood leukocytes were diluted according to the traced malignant cell frequencies and diluted to the malignant cell frequencies that needed to be confirmed as LOD according to Table 10.
[0132] Table 10 Corresponding relationship between the frequency and number of malignant cells at 20 μg DNA input
[0133] Note: The DNA quality here is calculated based on the base number and diploid characteristics of the human genome, and 1 cell is calculated as 6.5pg.
[0134] Under the condition of a sample load of 20 μg, the method of the present invention was used to perform repeated detection three times at three MRD frequencies, and the LOD was preliminarily established with a detection frequency of 100%. It can be seen from the results in Table 11 that the minimum frequency of positive detection in three repeated tests for three disease types was 6.50E-07.
[0135] Table 11 20μg-LOD establishment test results
[0136] 3. LOD confirmation of 20μg loading: The minimum detection limit standard ALL, CLL and MM cell lines established by LOD were selected at 2 malignant cell contents and 3 malignant cell contents, respectively, under the condition of 20 μg DNA input, and the method of the present invention was used for 13 repeated tests, and the minimum detection limit was determined with a 100% positive detection rate. The results in Table 12 determined that the minimum detection limit LOD of ALL, CLL and MM was 6.50E-07 when the sample load was 20 μg.
[0137] Table 12 20μg-LOD confirmation results
[0138] Example 6 ctDNA Detection The present invention expands the sample types of ctDNA on the basis of the previous detection sample types, and selects 8 patient plasma samples with known pathological information, including 3 diffuse large B-cell lymphoma (DLBCL), 1 mantle cell lymphoma (MCL), 1 primary central nervous system lymphoma (PCNSL), and 3 multiple myeloma (MM). The ctDNA extracted from the samples before and after treatment of each patient was subjected to MRD detection using the method in Example 1. The results are shown in Table 13.
[0139] Table 13 ctDNA sample detection results
[0140] Example 7 Detection of Fusion Fusion genes are the earliest type of genetic abnormalities identified in hematological tumors and have clear clinical diagnostic and therapeutic significance. In the latest World Health Organization (WHO) 2016 classification criteria for hematopoietic and lymphoid tumors, there are 112 FGs clearly listed, most of which can be used as a basis for diagnosis and classification. The fusion genes included in this tracking are BCL1-IGH and BCL2-IGH.
[0141] Sanger was used to trace the frequency of the fusion gene BCL1-IGH in the JVM-2 cell line. According to the traced malignant cell frequencies from high to low, the gDNA of the JVM-2 cell line and the gDNA of healthy human whole blood leukocytes were diluted to the frequencies in Table 14 below. The input amount of DNA was 20 μg, and each frequency standard was tested repeatedly 3 times to preliminarily establish the LOD.
[0142] Table 14 Corresponding relationship between malignant cell frequency and number of malignant cells under different DNA input amounts
[0143] Each frequency standard was tested three times at three different input levels (disease type MCL), and the LOD was initially established with a 100% detection frequency.
[0144] Table 15 20μg-LOD establishment test results
[0145] Example 8 Traceable positive clones, i.e., clones that exist both before and after treatment Four patients with known pathological information were selected, and samples were collected at different treatment time points, and MRD detection was performed using the method of the present invention. The test results are shown in Table 16. Two significant clonotypes were detected in the sample of patient 1 before treatment. The sample was tracked 3 months after treatment, and no two significant clonotypes were detected. The sample was tested 6 months after treatment. The 2 significant clonotypes before treatment were not traced to residues, but new clonotypes were detected, and the clonal cell content was 5.43E-03. Patient 2 had 3 significant clonotypes detected before treatment, and 10 were tracked 3 months after treatment. -6 The frequency was residual, and rebounded 6 months after treatment, and a high concentration of clonal residual was detected. Patient 3 had 3 significant clonal types detected before treatment, and these 3 significant clonal types were also tracked 3 months after treatment, but the clonal cell content decreased. Patient 4 had 2 significant clonal types detected before treatment, and 3 months after treatment, the clonal cell content of one of the clonal types decreased to 10 -6 , no residual IGH-A clonotype. No data were available for patients 3 and 4 6 months after treatment.
[0146] Table 16 MRD monitoring results
[0147] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A model for monitoring microresidues of human blood tumors, characterized in that: include: A data acquisition module, used to receive DNA sequencing data of the same test subject before and after treatment; A nucleated cell number acquisition module is used to align the sequencing data of the internal reference gene to the reference genome to obtain the nucleated cell number; The clone map and sequence acquisition module is used to align the DNA sequencing data to the immune database including IGH, IGK and IGL genes through an alignment algorithm, and then obtain the clone map and the corresponding sequence through a clustering algorithm; The fusion genotype and sequence acquisition module is used to align the DNA sequencing data to the reference gene sequence to obtain the fusion genotype and the corresponding sequence; a calculation module, for calculating the clonal frequency, clonal cell content, fusion frequency and fusion cell content based on the nucleated cell number, clonal map and corresponding sequence, fusion genotype and corresponding sequence, and determining significant clonal type and significant fusion type based on the clonal frequency, clonal cell content, fusion frequency and fusion cell content; The MRD calculation and determination module is used to screen effective clonal types and fusion types based on the clonal types and fusion types obtained before and after treatment of the same subject, calculate the MRD value based on the clonal cell content and fusion cell content corresponding to the effective clonal type and fusion type, and determine the residual status of human blood tumors based on the MRD value.
2. A model for monitoring microresidues of human blood tumors according to claim 1, characterized in that: The clustering algorithm includes: Calculate the edit distance between sequences based on the DNA sequencing data; Based on the edit distance and its relationship with the IGH gene V, D, J region, IGK gene V, J region and IGL gene V, J region, the sequenced sequences were clustered and analyzed to obtain a clonal map.
3. A model for monitoring microresidues of human blood tumors according to claim 2, characterized in that: The edit distance is calculated as follows: ; ; ; Among them, seq1 and seq2 represent the sequences obtained by two sequencing, i represents the ith base in seq1, j represents the jth base in seq2, i and j both start from 1 and go to the last digit, and the initial value is 0; ms represents the match score, when the ith base of seq1 and the jth base of seq2 are consistent, the m(i, j) value is a positive value; ss represents the mismatch score, when the ith base of seq1 and the jth base of seq2 are inconsistent, the m(i, j) value is a negative value; os represents the gap starting score, gapoping represents the start of a gap in sequence matching; es represents the gap extension score; gap extending represents the extension of a gap in sequence matching; S represents the clip score, when the bases of seq1 and seq2 are highly mismatched, the negative value is too high, and truncation will be chosen at this time.
4. A model for monitoring microresidues of human blood tumors according to claim 2, characterized in that: The cluster analysis includes: The best matching method is given according to the alignment results and the edit distance. The sequencing sequences with the same V and J regions of the IGH, IGK and IGL genes are clustered. When IGH-DJ rearrangement occurs, the sequencing sequences with the same D and J regions of the IGH genes are clustered. When IGK-KDE rearrangement occurs, the sequencing sequences with the same V regions of the KDE and IGK genes are clustered. In each VJ class, the division was based on the consistency of the CDR3 region; in each DJ class, the division was based on the consistency of the 50bp length region centered on the junction of the D region and the J region of the IGH gene; in each KDE-V class, the division was based on the consistency of the 50bp length region centered on the junction of the V region of the IGK gene and the KDE gene; the number of sequencing sequences corresponding to each subclass was obtained; A hierarchical clustering method is used to cluster each sub-class. In each cluster, the sequences of the child nodes and parent nodes are highly similar, and the parent node has a significantly higher number of sequenced sequences than the child node. Sum the number of sequenced sequences of all nodes in each cluster, and report the root node sequence as a separate clonal type; Each individual clonotype reported is combined to form a clonal map.
5. The human blood tumor microresidue monitoring model according to claim 1, characterized in that: The calculation method of the clone frequency and clone cell content is: ; ; ; ; Among them, cfl represents the clonal frequency of a single IGL clone, count1 is the sequence number of a single IGL clone, and count IGL is the number of sequences of all IGL clones, cfh represents the clonal frequency of a single IGH clone, count2 is the number of sequences of a single IGH clone, and count IGH is the number of sequences of all IGH clonotypes except IGH-DJ clonotype, count HDJ is the number of sequences of all IGH-DJ clones; cfk represents the clonal frequency of a single IGK clone, count3 is the number of sequences of a single IGK clone, and count IGK is the number of sequences of all IGK clonotypes except IGK-KDE clonotypes, count KDE is the sequence number of all IGK-KDE clones; the number of cloned cells is the sequence number of a single clone.
6. The human blood tumor microresidue monitoring model according to claim 1, characterized in that: The calculation method of the fusion frequency and fusion cell content is: ; ; Among them, fc is the number of sequences of a single fusion type, hc is the total number of IGH sequences, bc is the total number of BCL1 or BCL2 sequences, and the number of fusion cells is the number of sequences of a single fusion type.
7. The human blood tumor microresidue monitoring model according to claim 1, characterized in that: The screening criteria for significant clonotypes are: (1) Clonal frequency ≥ 0.03; (2) Clonal cell content ≥ 0.002; (3) The clone frequency distribution is discontinuous. The judgment criteria are as follows: (a) Take the clonal types with a frequency less than 0.03 and arrange them in descending order; (b) The threshold is 10 times the clonal frequency of the fifth clonotype in the sequence; (c) The clonal frequencies of all clonal types are compared with the threshold, and the clonal types that are greater than or equal to the threshold are considered to satisfy the discontinuity of clonal frequency distribution.
8. The human blood tumor microresidue monitoring model according to claim 1, characterized in that: The screening criteria for significant fusion type are: (1) Fusion frequency ≥ 0.1%; (2) Fusion cell content ≥ 0.02%.
9. The human blood tumor microresidue monitoring model according to claim 1, characterized in that: MRD is calculated as: MRD 克隆 = max(sum(IGH clone cell content), sum(IGK clone cell content) + sum(IGL clone cell content)); MRD 融合 = max(sum BCL1-IGH fusion cell content), sum(BCL2-IGH fusion cell content)); ; or 。 10. The human blood tumor microresidue monitoring model according to claim 1, characterized in that: The model also includes: Obtain clonotypes of healthy individuals and / or patients with non-similar diseases, and calculate the clonal cell content of the clonotypes; Based on the clonotypes of the healthy persons and / or patients with non-similar diseases, the significant clonotypes obtained based on the DNA sequencing data before treatment are screened out, and the clonotypes of the healthy persons and / or patients with non-similar diseases are not rare and the clonotype cell content is ≥10 -7 The non-rare existence is a clonal type that exists in ≥5% of the population.
Citation Information
Patent Citations
Detection method, device and equipment for tiny residual focus and storage medium
CN115679000A
Primer composition and kit for monitoring minimal residual disease of human leukemia and application of primer composition and kit
CN115927631A
Composition for clinical diagnosis and treatment of hematological malignant tumors and application thereof
CN116121383A
Tumor minimal residual focus detection method based on customization strategy
CN116580768A
Ultra-high depth sequencing-based tiny residual focus detection method and system
CN119252324A