Blood tumor diagnosis and treatment primer group, kit and application

By designing multiplex PCR primer sets and kits for hematological malignancies, we have achieved efficient amplification and analysis of the CDR3 sequence, which solves the problems of insufficient sensitivity and precision of existing MRD detection technologies and provides accurate immune status monitoring and personalized treatment support.

CN120330335BActive Publication Date: 2026-07-21WUHAN XINO MEDICAL LABORATORY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN XINO MEDICAL LABORATORY CO LTD
Filing Date
2025-04-28
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing MRD detection technologies lack sufficient sensitivity and precision in hematologic malignancies, failing to comprehensively monitor changes in tumor cells, especially when clonal tumor cells exhibit antigenic changes, leading to false negative results. Furthermore, there is a lack of effective technologies that reflect the overall state of a patient's immune system.

Method used

A primer set suitable for multiplex PCR of human immune repertoire was designed, including a first and second primer set, for amplifying genes such as IGH, IGDH, IGK, IGL, TRB VJ, TRB DJ, TRG, and TRD. Combined with an artificial internal reference sequence that mimics the natural CDR3 and a housekeeping gene, the CDR3 sequence can be accurately detected and analyzed through multiplex PCR amplification and high-throughput sequencing.

Benefits of technology

It significantly improves the positive detection rate of clinical samples, accurately reflects the individual's immune status, dynamically monitors changes in patients' tumor cells, provides technical support for precision medicine and personalized treatment, and improves the sensitivity and precision of MRD detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120330335B_ABST
    Figure CN120330335B_ABST
Patent Text Reader

Abstract

The application discloses a blood tumor diagnosis and treatment primer group, a kit and application, the primer group comprises a first primer group and / or a second primer group; the first primer group is used for amplifying IGH, IGDH, IGK and IGL; the second primer group is used for amplifying TRB V-J, TRB D-J, TRG and TRD, after the primer group amplification reordering sequence, the positive detection of a clinical sample can be significantly improved, the CDR3 sequence obtained can be used for clinical monitoring and can independently evaluate the immune group of each patient, accurately reflects the immune state of the individual, dynamically monitors the change of the tumor cells of the patient, and has good reliability and repeatability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of immune repertoire sequencing technology, specifically to primer sets, reagent kits, and applications for the diagnosis and treatment of hematological malignancies. Background Technology

[0002] Lymphocytes can differentiate into B cells, T cells, and natural killer cells. The normal function of B cells and T cells is crucial for protecting the body from invading viruses or bacteria. This proper function depends on the structure of B cells and T cells. The antigen receptor on the surface of T cells is called the T cell receptor (TCR). Approximately 95% of TCRs are composed of α and β chains linked by disulfide bonds, while the remaining 5% are composed of γ and δ chains. The antigen receptor on the surface of B cells is called the B cell receptor (BCR), which consists of two identical heavy chains and two identical light chains. The diversity of TCRs and BCRs determines the diversity of T / B cell antigen recognition.

[0003] The amino acid composition near the N-terminus of the TCR and BCR varies greatly, forming the variable region (V region). The C-terminus amino acid composition is relatively constant, called the constant region (C region). Within the complementarity determining region (CDR) of the variable region, CDR3 exhibits the highest diversity, with the diversity increasing closer to the antigen binding point during antigen recognition. The gene clusters encoding the TCR and BCR exist as clusters of segmented gene fragments on chromosomes during the germline stage before T / B cell development. These gene clusters contain V and J gene fragments at the 5' and 3' ends, respectively, with a D gene fragment in the middle. Each T / B cell has a unique TCR / BCR encoding gene; this diversity results from rearrangements of the V, D, and J genes during T / B cell development. During rearrangement, when different gene segments join, the homologous recombination repair system randomly deletes some nucleotides from the germline gene's terminal portion, while terminal nucleotide transferases randomly add non-template nucleotides to the joining site. These randomly added nucleotides create new sequences not present in the original sequence. In this V(D)J rearrangement, the randomness of the selection of V, D, and J genes and the imprecision of the rearrangement connection result in each T cell or B cell having a unique TCR or BCR encoding gene, forming a large immune repertoire.

[0004] Minimal residual disease (MRD) is defined as residual tumor cells that cannot be detected by conventional methodologies. This result is an important prognostic factor, and treatment strategies adjusted based on MRD results can significantly improve the prognosis of patients with hematologic malignancies. Among the conventional methods for monitoring MRD, multiparameter flow cytometry (MFC) and PCR are the most important methodologies. MFC measurement of MRD depends on the immunophenotype of tumor cell surface, but its sensitivity is usually limited, only about 10%. -3 -10 -4 However, the results are highly dependent on the labeled antibody and may not comprehensively cover all possible case subtypes. PCR methods detect specific fusion genes or related mutated genes in tumor cells, heavily relying on known rearrangements and mutations, and cannot provide comprehensive monitoring. Therefore, some clonal tumor cells may exhibit antigenic changes after treatment, and tumor heterogeneity can lead to false negatives, thus limiting the application of traditional methodologies. Recent studies have shown that in hematologic malignancies, NGS methods offer higher sensitivity and precision than MFC or PCR methods. For example, in leukemia patients, using sensitive and robust NGS methods for MRD monitoring can enhance the identification of relapse risk after hematopoietic stem cell transplantation or chimeric antigen receptor-modified T-cell therapy, and can identify newly emerging clones at each monitoring time point. Furthermore, for MRD quantification results, NGS-based immune repertoire detection technology can achieve results within 10... 6 Even with larger cell numbers, quantification can be performed at the molecular level to accurately assess the burden of MRD. Furthermore, the CDR3 sequence obtained from NGS sequencing will serve as a unique identifier to distinguish different subpopulations within the same type of clonal population. This more detailed differentiation can provide a more refined and accurate assessment for subsequent MRD monitoring of treatment efficacy.

[0005] Currently, domestic products related to immune repertoire testing for MRD (Mean Discharge) are still in the early stages of research and development. Although some progress has been made, existing products still cannot fully meet the needs of clinical patients and researchers, both in terms of market demand and technological maturity. Furthermore, there is a lack of effective technologies in clinical testing that can comprehensively reflect the overall state of a patient's immune system. With the combination of immune repertoire and NGS (Next Generation Sequencing) technology, this innovative approach brings unprecedented opportunities to clinical treatment and related research fields, providing strong technical support for future precision medicine and personalized treatment. Summary of the Invention

[0006] This invention provides a primer set suitable for multiplex PCR of human immune repertoires, used for multiplex amplification of fully rearranged and incompletely rearranged CDR3 sequences in human immune cells. The rearranged sequences include IGH, IGDH, IGK, IGL, TRB VJ, TRB DJ, TRG, and TRD. Clinical data shows that amplifying the rearranged sequences using this multiplex PCR primer set significantly improves the positive detection rate of clinical samples. The obtained CDR3 sequences can be used for clinical monitoring and can independently assess the immune repertoire of each patient, accurately reflecting the individual's immune status.

[0007] In view of this, the solution of the present invention is as follows:

[0008] The first aspect of the present invention is to provide a primer set for the diagnosis and treatment of hematological malignancies, comprising a first primer set and / or a second primer set; the first primer set is used to amplify IGH, IGDH, IGK and IGL; the second primer set is used to amplify TRB VJ, TRB DJ, TRG and TRD;

[0009] The nucleotide sequences of the upstream and downstream primers used to amplify IGH are shown in SEQ ID NO: 1-25 and 26-32, respectively;

[0010] The nucleotide sequences of the upstream and downstream primers used to amplify IGDH are shown in SEQ ID NO: 33-39 and 40-44, respectively.

[0011] The nucleotide sequences of the upstream and downstream primers used to amplify IGK are shown in SEQ ID NO: 45-54 and 55-63, respectively.

[0012] The nucleotide sequences of the upstream and downstream primers used to amplify IGL are shown in SEQ ID NO: 64-73 and 74-80, respectively.

[0013] The nucleotide primer sequences for the upstream and downstream primers used to amplify TRB VJ are shown in SEQ ID NO: 81-119 and 120-129, respectively.

[0014] The nucleotide sequences of the upstream and downstream primers used to amplify TRB DJ are shown in SEQ ID NO: 130-145 and 146-156, respectively.

[0015] The nucleotide sequences of the upstream and downstream primers used to amplify TRD are shown in SEQ ID NO: 157-169 and 170-177, respectively.

[0016] The nucleotide sequences of the upstream and downstream primers used to amplify TRG are shown in SEQ ID NO: 178-196 and 197-205, respectively.

[0017] Furthermore, the primer set also includes at least one of the following: an artificial internal reference sequence set simulating natural CDR3, a housekeeping gene sequence, and sequencing primers.

[0018] Furthermore, the first primer set includes a set of artificial internal reference sequences as shown in SEQ ID NO: 206-281, and the second primer set includes a set of artificial internal reference sequences as shown in SEQ ID NO: 282-352;

[0019] The housekeeping gene is GAPDH, and its sequence is shown in SEQ ID NO: 353-354;

[0020] The sequencing primer sequences are shown in SEQ ID NO: 355-356.

[0021] A second aspect of the present invention is to provide a diagnostic and therapeutic kit for hematologic malignancies, comprising the primer set described in the first aspect.

[0022] Furthermore, the kit consists of two tubes, each containing a first primer set and a second primer set.

[0023] Furthermore, the final primer concentrations for amplifying IGH, IGDH, IGK, and IGL were 1 μM, 0.8 μM, 0.5 μM, and 0.35 μM, respectively; and / or the final primer concentrations for amplifying TRB VJ, TRB DJ, TRG, and TRD were 1.2 μM, 1.0 μM, 0.85 μM, and 0.5 μM, respectively.

[0024] Furthermore, the kit also includes at least one of the following: a PCR amplification reaction system, a nucleic acid extraction reagent, an RNA reverse transcription reagent, and a PCR product sequencing reagent. The nucleic acid extraction reagent is used to extract sample DNA and / or RNA.

[0025] A third aspect of the present invention is to provide the use of the primer set described in the first aspect, or the kit described in the second aspect, in the preparation of hematologic malignancy diagnostic and therapeutic products.

[0026] Furthermore, the hematologic oncology diagnostic and treatment product is used for at least one of the following purposes:

[0027] 1) Immune repertoire analysis;

[0028] 2) Detection and tracking of minimal residual lesions;

[0029] 3) Dynamic monitoring of immune status.

[0030] Furthermore, the application includes the following steps:

[0031] Obtain nucleated cell DNA, cDNA, or cfDNA samples, add artificial internal reference sets, perform multiplex PCR amplification using the amplification primer set described in the first aspect, perform secondary amplification on the PCR products to add sequencing adapters, construct a library, and perform sequencing.

[0032] The immune repertoire files obtained from sequencing need to undergo quality control, preprocessing, and assembly. They are then compared with sequences in the IMGT database, and clones are clustered based on CDR3 sequences or incomplete rearrangements, generating files sorted by clonal abundance. The MRD calculation formula is as follows:

[0033] MRD=[(Corrected_Clones*C_int*n_GC*n_total) / (Corrected_R_int*N_cells)]*F_batch*Q, where:

[0034] Corrected_Clones is the number of tumor-related reads obtained after corrected sequencing; C_int is the copy number of each internal reference set added; η_GC: GC% content difference correction coefficient; η_total is the DNA extraction, amplification, ligation, and sequencing correction coefficient; Corrected_R_int is the number of internal reference reads obtained from sequencing; F_batch is the technical bias correction coefficient; Q is the fragment ratio and genomic DNA ratio weighting coefficient; N_cells is the total number of nucleated cells in the sample;

[0035] The total number of nucleated cells is calculated using the following formula: N_cells=(M_DNA*η_DV / A_DNA)*[(1+0.1*(Batch_effect_score)], where:

[0036] M_DNA represents the total amount of DNA extracted from the experimental sample (ng); η_DV represents the fragment proportion weighting coefficient; A_DNA represents the average amount of DNA per nucleated cell; and Batch_effect_score represents the bias coefficient.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] The primer design process provided by this invention clusters the CDR3 core region to amplify the most complete and comprehensive subtype with the fewest primer combinations, reducing the complexity of primer combinations and avoiding adverse factors such as poor homogenization and multiple dimers caused by too many primer combinations. After amplifying and rearranging the sequence using this multiplex PCR primer set, the positive detection rate of clinical samples can be significantly improved. The obtained CDR3 sequence can be used for clinical monitoring and can independently evaluate the immune repertoire of each patient, accurately reflecting the individual's immune status and dynamically monitoring the changes in the patient's tumor cells, with good reliability and reproducibility.

[0039] In the blood tumor diagnosis and treatment application described in this invention, by reasonably introducing a correction coefficient, the actual residual level of tumor cells can be accurately restored, providing strong technical support for future precision medicine and personalized treatment. Attached Figure Description

[0040] Figure 1 A schematic diagram of the blood tumor detection and analysis process described in this embodiment of the invention.

[0041] Figure 2 This is the result of the pre-treatment bone marrow sample library fragment analysis in an embodiment of the present invention.

[0042] Figure 3 This is a schematic diagram of the bioinformatics analysis process described in an embodiment of the present invention. Detailed Implementation

[0043] The technical solution of the present invention will now be clearly and completely described in conjunction with preferred embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0044] Example 1: Primer set design for blood tumor detection

[0045] This embodiment provides a complete set of primer combinations for the detection of hematological malignancies through extensive testing and analysis. The primers cluster the CDR3 core region, amplifying the most complete and comprehensive subtypes with the fewest primer combinations. This reduces the complexity of primer combinations and avoids the disadvantages of excessive primer combinations, such as poor homogenization and numerous dimers. It improves the accuracy of CDR3 region sequence variation analysis, dynamically monitors changes in patient tumor cells, and indicates disease progression. The primer combination includes at least one tube amplifying IGH, IGDH, IGK (including KDE), and IGL, and one tube amplifying TRB VJ, TRB DJ, TRG, and TRD.

[0046] Specifically, the upstream primer sequence for amplifying IGH is shown in SEQ ID NO:1-25, and the downstream primer sequence is shown in SEQ ID NO:26-32; the upstream primer sequence for amplifying IGDH is shown in SEQ ID NO:33-39, and the downstream primer sequence is shown in SEQ ID NO:40-44; the upstream primer sequence for amplifying IGK is shown in SEQ ID NO:45-54, and the downstream primer sequence is shown in SEQ ID NO:55-63; the upstream primer sequence for amplifying IGL is shown in SEQ ID NO:64-73, and the downstream primer sequence is shown in SEQ ID NO:74-80; the upstream primer sequence for amplifying TRB VJ is shown in SEQ ID NO:81-119, and the downstream primer sequence is shown in SEQ ID NO:120-129; the upstream primer sequence for amplifying TRB DJ is shown in SEQ ID NO:130-145, and the downstream primer sequence is shown in SEQ ID NO:146-156; the upstream primer sequence for amplifying TRD ... The downstream primer sequence is shown in SEQ ID NO:157-169, and the upstream primer sequence is shown in SEQ ID NO:170-177; the downstream primer sequence is shown in SEQ ID NO:178-196, and the downstream primer sequence is shown in SEQ ID NO:197-205.

[0047] In the above embodiments, the primer set further includes a synthetic artificial internal reference set that mimics the natural CDR3 sequence, preferably an artificial internal reference set sequence as shown in SEQ ID NO:206-352.

[0048] In the above embodiments, the primer set further includes the housekeeping gene GAPDH sequence; the preferred housekeeping gene GAPDH sequence is shown in SEQ ID NO:353-354.

[0049] In the above embodiments, the primer set further includes sequencing primers, wherein the upstream sequencing primer sequence is shown in SEQ ID NO:355 and the downstream sequencing primer sequence is shown in SEQ ID NO:356.

[0050] The primer sets mentioned above are shown in Table 1 (where F represents the upstream primer and R represents the downstream primer).

[0051] Table 1:

[0052]

[0053]

[0054]

[0055]

[0056]

[0057]

[0058]

[0059]

[0060]

[0061]

[0062]

[0063]

[0064]

[0065] In the primer set mentioned above, artificial internal controls for the corresponding genes are added when amplifying IGH, IGDH, IGK, IGL, TRB VJ, TRB DJ, TRG, and TRD. For example, when amplifying IGH, IGH-Internal control is added to facilitate accurate copy number calculation.

[0066] For each gene, upstream and downstream primers are randomly combined for amplification to meet the needs of different individual samples and samples at different stages of development.

[0067] For the sequencing primers shown in SEQ ID NO:355-356, the underlined part is the unique number corresponding to each library, N is a degenerate base, which can be A, T, G or C; the other parts are the same sequence.

[0068] Example 2: Diagnostic and therapeutic reagent kits and sample detection and analysis

[0069] This embodiment provides a sample detection and analysis method. It involves extracting DNA and RNA from the sample, performing multiplex PCR amplification using the primer set provided in Example 1, and then adding sequencing adapters for a second round of PCR amplification. High-throughput sequencing using BCR and TCR is then performed, followed by immune repertoire information analysis and master clone sequence labeling. The workflow is as follows: Figure 1 As shown. Specifically:

[0070] 1. Sample Acquisition:

[0071] Bone marrow samples were packaged in EDTA anticoagulant tubes;

[0072] Peripheral blood samples were collected in EDTA-anticoagulated vials.

[0073] 2. Multiplex PCR primer mixing:

[0074] The primers used in the first step of the PCR reaction were mixed in the following proportions to obtain the primer pool. The final concentration of IGH primer was 1 μM, the final concentration of IGDH primer was 0.8 μM, the final concentration of IGK primer was 0.5 μM, the final concentration of IGL primer was 0.35 μM; the final concentration of TRB VJ primer was 1.2 μM, the final concentration of TRB DJ primer was 1.0 μM, the final concentration of TRG primer was 0.85 μM, and the final concentration of TRD primer was 0.5 μM.

[0075] 3. Construction of a DNA-template library:

[0076] use Spin gDNA Extraction Kit extracts genomic DNA from nucleated cells.

[0077] Depending on the concentration of different primer pools, the internal reference set or the internal reference set combined with the housekeeping gene GAPDH primer pool was used for specific amplification of the sample. Different calculation methods were used depending on the amplification strategy. The number of internal reference sets used were: IGH 54, IGDH 20, IGK 18, IGL 23, TRB VJ 62, TRB DJ 12, TRG 16, and TRD 24.

[0078] Using the Multiplex PCR Assay Kit Ver.2, prepare the first step of the multiplex PCR amplification reagent according to the reaction system in the table below. Mix the prepared reagent by inverting, centrifuge briefly, and place it in the PCR instrument. Perform the PCR amplification reaction under the following conditions:

[0079] Reaction system:

[0080] reagents Usage <![CDATA[2×Multiplex PCR Buffer(Mg 2+ ,dNTP plus)]]> 25μL Primer Mix*1 Appropriate amount Multiplex PCR Enzyme Mix 0.25μL DNA template 2μL Internal control 5μL <![CDATA[dH2O (Sterile water)]]> up to 50μL

[0081] Step 1 PCR amplification reaction conditions:

[0082]

[0083] 4. Library amplification using RNA reverse transcription products as templates

[0084] Highly efficient and rapid RNA reverse transcription was performed using the RevertSAid RT reverse transcription kit. The DNA obtained from the reverse transcription was then used as a template for library construction as described in step 3. The reverse transcription system is as follows:

[0085] reagents Usage total RNA 1μg oligo(dT)primer (0.5 μg / μL) 1μL DEPC-treated water to 12μL

[0086] Mix gently and centrifuge for 3-5 seconds.

[0087] Incubate the mixture at 70°C for 5 minutes, cool it on ice, and centrifuge briefly.

[0088] Place the centrifuge tubes on ice and add the following components:

[0089] reagents Usage 5x reaction buffer 4μL RiboLock Ribonuclease Inhibitor(20u / μL) 1μL 10mM dNTP mix 2μL

[0090] Incubate at 37°C for 5 minutes;

[0091] Add 1 μL of RevertAid TMM-MuLV Reverse Transcriptase (200 u / μL), and then add DEPC water to a final volume of 20 μL.

[0092] Incubate the mixture at 42°C for 60 minutes;

[0093] The reaction was terminated by incubating at 70°C for 10 minutes, and the container was then quickly placed on ice.

[0094] 5. Add sequencing primers to the PCR amplification target product.

[0095] Components Volume (μL) 2×KAPA HiFi premixed liquid 12.5 10μM forward primer 0.75 10μM reverse primer 0.75 Template DNA 1 PCR-grade water up to 25

[0096] PCR running conditions:

[0097]

[0098]

[0099] Transfer the PCR reaction mixture to a 1.5 mL centrifuge tube and purify the amplified product 1.0-fold using the AMPure XP DNA Purification Kit (Thermo). The specific method is as follows:

[0100] a) Remove Ampure XP Beads at 4℃ and allow them to stand at room temperature for 30 minutes to equilibrate;

[0101] b) Shake well before use, add magnetic beads at a 1:1 ratio with the sample volume and mix well, then let stand for 5 minutes.

[0102] c) Transfer the 1.5 mL centrifuge tube to a magnetic rack and let it stand for 3-5 minutes until clear;

[0103] d) Keep the centrifuge tubes on the magnetic rack and carefully remove the supernatant, being careful not to touch the magnetic beads;

[0104] e) Add 500 μL of 75% ethanol, gently blow the magnetic beads 2-3 times, wait 30 seconds, and discard the supernatant (when adding ethanol, it should be added slowly, and try not to add the liquid towards the magnetic beads, otherwise the magnetic beads will detach from the tube and be damaged).

[0105] f) Repeat step e to remove as much supernatant as possible (no need to blow the magnetic beads);

[0106] g) Place the magnetic beads in a constant temperature mixer at 37°C for 3-5 minutes to dry until there is no moisture on the surface of the magnetic beads.

[0107] h) Add 31 μL of nuclease-free water to a 1.5 mL centrifuge tube, mix thoroughly, and let stand for 5 minutes.

[0108] min, then place on a magnetic rack for about 5 min until clear;

[0109] i) Transfer 30 μL of liquid to a new 1.5 mL centrifuge tube that has been prepared beforehand.

[0110] The concentration of purified DNA was determined using a Qbit BR analyzer, and the library concentration must be greater than 5 μg / μL.

[0111] Using a pre-treatment bone marrow sample, a library fragment was obtained using the above method. The library fragment length was analyzed, and the results are as follows: Figure 2 As shown, this indicates that the length of the constructed library fragments meets expectations.

[0112] 6. Sequencing and bioinformatics analysis

[0113] DNA libraries were sequenced using an Illumina Novaseq™ 6000 platform with PE150. The library denaturation concentration was 2 nM, and the sequencing concentration was 20 nM. After quality control and assembly, the sequencing data were compared with the IMGT (http: / / www.imgt.org / ) database. Clones were clustered based on CDR3 sequence, and result files were generated according to clonal abundance. The workflow is as follows: Figure 3 As shown, the specific steps are as follows:

[0114] 1) Construction of reference sequence

[0115] The complete sequences of the V, D, and J genes were downloaded from the International Immunogenetic Information System (IMGT). Specific primers for each strand were compared with the V, D, and J sequences, retaining those with primer-template mismatches of 2 or fewer bases. Based on the start and stop positions of the primers, the portions capable of generating PCR products were selected as reference sequences. For primer-template mismatches, two reference sequences were recorded, the reference sequence libraries were merged, duplicate reference sequences were removed, and only unique sequences were retained. To improve sequence alignment accuracy, the start and stop positions of the CDR3 region were further clarified during reference sequence construction. This step facilitates subsequent immune repertoire analysis, especially in-depth analysis of antibody-receptor recombination processes on B and T cells, enabling more precise resolution of key sequences related to antibody production and immune responses. Furthermore, by combining algorithm improvements and machine learning techniques, the generation and screening process of reference sequences is automatically optimized, significantly improving efficiency and accuracy in large-scale data processing.

[0116] 2) Sequencing data processing

[0117] Sequencing data needs to be filtered to remove adapters, low-quality sequences, and sequences with low alignment rates;

[0118] 3) Merging data from both ends

[0119] The sequences obtained from sequencing are spliced ​​and merged, meaning there must be at least 5 bp of overlapping region.

[0120] 4) Alignment of merged data with reference sequences

[0121] When aligning the assembled sequence with the constructed reference sequence, the overlapping regions are first compared with the V, D, and J sequences separately. To improve alignment accuracy, a distributed alignment strategy is adopted: the V region extends towards the 3' end, the D region extends bidirectionally, and the J region extends towards the 5' end. After the initial alignment, the alignment position is optimized using BLAST results. Through re-alignment, score calculation, and identity verification, the V, D, and J regions with the highest scores are selected as the best results. Specific mismatch tolerance ranges are set for the V, D, and J regions, and the number of mismatches allowed for each chain varies depending on the sequence characteristics. Combining machine learning algorithms, the mismatch tolerance is automatically adjusted based on historical alignment results and sequence characteristics. Especially when dealing with the highly variable CDR3 region, this dynamic adjustment helps improve alignment accuracy. Utilizing efficient parallel computing and alignment algorithms accelerates large-scale data processing, further ensuring the accuracy and reliability of the analysis results.

[0122] 5) Filtering the comparison results

[0123] Sequences with low frequency (<3%) will be removed. Sequences will also be removed if the alignment results of the V gene and J gene do not match the positive and negative strands of the genome. Pseudogene alignment results will be removed. If a CDR3 sequence is identified and a terminator is present during the translation of amino acids, and the complete sequence cannot be translated normally, it will be removed.

[0124] 6) Data Statistics

[0125] Based on the CDR3 region location, DNA sequences are translated into amino acid sequences, and peptide frequencies are statistically analyzed. After filtering out sequences that do not meet the standard through the above steps, the usage frequency and variation (insertion, deletion, mutation, etc.) of the V, D, and J genes are statistically analyzed to gain a deeper understanding of the relationship between gene mutation patterns and immune responses, as well as sequence length and base composition. Furthermore, the diversity of the immune repertoire is assessed by visualizing V and J gene expression patterns through bar charts and heatmaps, comprehensively evaluating the diversity characteristics of the immune repertoire.

[0126] Clonal clustering is performed based on CDR3 sequences or incomplete rearranged sequences, and files are generated sorted by clonal abundance; the MRD formula is as follows:

[0127] MRD=[(Corrected_Clones*C_int*n_GC*n_total) / (Corrected_R_int*N_cells)]*F_batch*Q, where:

[0128] Corrected_Clones is the number of tumor-related reads obtained after corrected sequencing; C_int is the copy number of each internal reference set added; η_GC: GC% content difference correction coefficient; η_total is the DNA extraction, amplification, ligation, and sequencing correction coefficient; Corrected_R_int is the number of internal reference reads obtained from sequencing; F_batch is the technical bias correction coefficient; Q is the fragment ratio and genomic DNA ratio weighting coefficient; N_cells is the total number of nucleated cells in the sample;

[0129] The total number of nucleated cells is calculated using the following formula: N_cells=(M_DNA*η_DV / A_DNA)*[(1+0.1*(Batch_effect_score)], where:

[0130] M_DNA represents the total amount of DNA extracted from the experimental sample (ng); η_DV represents the fragment proportion weighting coefficient; A_DNA represents the average amount of DNA per nucleated cell; and Batch_effect_score represents the bias coefficient.

[0131] Validated with 50 samples, the coefficient of variation (CV) between the percentage of tumor cells monitored by flow cytometry and the calculated value was very low (CV < 20%) after introducing the correction coefficient, while the CV reached 30%-115% without correction. The results indicate that the correction coefficient can accurately reflect the actual residual level of tumor cells.

[0132] Example 3 Primer Pair Working Results Analysis

[0133] Peripheral blood samples were collected from 100 healthy individuals and processed according to the steps in Example 2. Bioinformatics analysis was performed on two tubes (one for IGH, IGDH, IGK, and IGL; the other for TRB VJ, TRB DJ, TRD, and TRG) to assess whether primer pairs within the same gene functioned correctly and whether cross-amplification occurred between primers from different genes. Results showed that each primer in each tube amplified only its own gene fragment, and different upstream and downstream primers were used depending on template differences in the samples. NGS sequencing results showed that the upstream and downstream primers for each gene randomly combined to amplify the target sequence in the sample, and each primer was detected to participate in the amplification of different target sequences in the same sample. The primer amplification results for IGH and TRB VJ are shown in Table 2 (in Table 2, primer names F_IGH 1 to 25 are the upstream primers numbered 1 to 25 in Table 1, and R_IGH 1 to 7 are the downstream primers numbered 26 to 32 in Table 1; the primer names for TRB VJ are similar). The primer amplification results for other genes (IGDH, IGK, IGL, TRB DJ, TRD, TRG) are similar to those for these two genes.

[0134] Table 2:

[0135]

[0136]

[0137]

[0138] Example 4: Clinical Sample Analysis

[0139] Samples from 110 patients clinically diagnosed with ALL and testing positive by flow cytometry were analyzed according to the method provided in Example 2. The results are shown in the table below.

[0140]

[0141]

[0142] The results are as follows:

[0143] Disease type Initial diagnosis consistency rate MRD Consistency Rate B-ALL 95.4% 92.3% T-ALL 96.3% 95%

[0144] The results showed that the composition used in the clinical diagnosis and treatment of hematological malignancies had a positive rate of over 95% for both B-ALL and T-ALL patients, and the monitoring positive rate was over 90%, demonstrating good reliability and reproducibility.

[0145] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. The application of the reagent kit in the preparation of hematologic oncology diagnostic and therapeutic products, characterized in that, The kit includes two primer sets: the first primer set is used to amplify IGH, IGDH, IGK, and IGL; the second primer set is used to amplify TRB VJ, TRB DJ, TRG, and TRD. The nucleotide sequences of the upstream and downstream primers used to amplify IGH are shown in SEQ ID NO: 1-25 and 26-32, respectively; The nucleotide sequences of the upstream and downstream primers used to amplify IGDH are shown in SEQ ID NO: 33-39 and 40-44, respectively. The nucleotide sequences of the upstream and downstream primers used to amplify IGK are shown in SEQ ID NO: 45-54 and 55-63, respectively. The nucleotide sequences of the upstream and downstream primers used to amplify IGL are shown in SEQ ID NO: 64-73 and 74-80, respectively; The nucleotide primer sequences for the upstream and downstream primers used to amplify TRB VJ are shown in SEQ ID NO: 81-119 and 120-129, respectively. The nucleotide sequences of the upstream and downstream primers used to amplify TRB DJ are shown in SEQ ID NO: 130-145 and 146-156, respectively. The nucleotide sequences of the upstream and downstream primers used to amplify TRD are shown in SEQ ID NO:157-169 and 170-177, respectively. The nucleotide sequences of the upstream and downstream primers used to amplify TRG are shown in SEQ ID NO: 178-196 and 197-205, respectively. The first primer set further includes a set of artificial internal reference sequences as shown in SEQ ID NO: 206-281, and the second primer set further includes a set of artificial internal reference sequences as shown in SEQ ID NO: 282-352; The application includes the following steps: Obtain nucleated cell DNA, cDNA, or cfDNA samples and add artificial internal reference sets. Perform multiplex PCR amplification using the first and second primer sets, respectively. Perform secondary amplification on the PCR products to add sequencing adapters, construct libraries, and perform sequencing. The immunoreperfusion library files obtained from sequencing underwent quality control, preprocessing, and assembly. They were then aligned with sequences in the IMGT database. Clones were clustered based on CDR3 sequences or incomplete rearrangements, and files sorted by clonal abundance were generated. The MRD calculation formula is as follows: MRD= [(Corrected_Clones * C_int * η_GC * η_total) / (Corrected_R_int * N_cells)] * F_batch * Q, where: Corrected_Clones is the number of tumor-related reads obtained after corrected sequencing; C_int is the copy number of each internal reference set added; η_GC: GC% content difference correction coefficient; η_total is the DNA extraction, amplification, ligation, and sequencing correction coefficient; Corrected_R_int is the number of internal reference reads obtained from sequencing; F_batch is the technical bias correction coefficient; Q is the fragment ratio and genomic DNA ratio weighting coefficient; N_cells is the total number of nucleated cells in the sample; The total number of nucleated cells is calculated using the following formula: N_cells = (M_DNA * η_DV / A_DNA) * [(1+0.1*(Batch_effect_score)], where: M_DNA represents the total amount of DNA extracted from the experimental sample (ng); η_DV represents the fragment proportion weighting coefficient; A_DNA represents the average amount of DNA in a single nucleated cell; and Batch_effect_score represents the bias coefficient.

2. The application according to claim 1, characterized in that, It also includes at least one of the housekeeper gene sequence and sequencing primers.

3. The application according to claim 2, characterized in that, The housekeeping gene is GAPDH, and its sequence is shown in SEQ ID NO: 353-354; And / or, the sequencing primer sequences are as shown in SEQ ID NO: 355-356.

4. The application according to claim 1, characterized in that, The final primer concentrations for amplifying IGH, IGDH, IGK, and IGL were 1 μM, 0.8 μM, 0.5 μM, and 0.35 μM, respectively; and / or the final primer concentrations for amplifying TRB VJ, TRB DJ, TRG, and TRD were 1.2 μM, 1.0 μM, 0.85 μM, and 0.5 μM, respectively.

5. The application according to claim 1, characterized in that, The kit also includes at least one of the following: a PCR amplification reaction system, a nucleic acid extraction reagent, an RNA reverse transcription reagent, and a PCR product sequencing reagent.

6. The application according to claim 1, characterized in that, The hematologic oncology diagnostic and treatment product is used for at least one of the following purposes: 1) Immune repertoire analysis; 2) Detection and tracking of minimal residual lesions; 3) Dynamic monitoring of immune status.