Multiplex PCR primer sets, design methods, and applications for IGH gene clonotype detection
By designing a multiplex PCR primer set for IGH gene clonotype detection, screening primers based on conserved regions and a weighted scoring mechanism, constructing a candidate database and adjusting concentrations, we solved the problems of incomplete coverage, slow updating, poor adaptability and low sensitivity in IGH gene clonotype detection, and achieved efficient, accurate and quantitative clonotype monitoring.
Patent Information
- Application Number
- CN202411282310.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-09-12
AI Technical Summary
Existing technologies for IGH gene clonotype detection suffer from problems such as incomplete primer pool coverage, difficulty in rapid updating, poor adaptability of amplification reagents, inconsistent amplification efficiency, low detection sensitivity, and low sequencing data utilization, making it difficult to accurately monitor minimal residual disease in patients with hematological malignancies.
A multiplex PCR primer set for IGH gene clonotype detection was designed. Primers were designed based on the conserved regions of the IGH gene, a weighted scoring mechanism was set up to screen primers, and a candidate primer database was constructed. The amplification efficiency and sensitivity were ensured by adjusting the primer concentration, and a multiplex PCR amplification system was used for precise quantification.
It achieves efficient coverage and accurate quantification of IGH gene clonotypes, can quickly adapt to new clonotypes and changes in amplification reagents, improves the sensitivity and accuracy of detection, and ensures the integrity and detection rate of the CDR3 region.
Smart Images

Figure CN118932067B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of gene detection technology, and specifically relates to a multiplex PCR primer set for IGH gene clonotype detection, a design method, and an application thereof. Background Art
[0002] For patients with hematologic malignancies, although remission rates are currently very high—reaching 95% for children with acute lymphoblastic leukemia and 70% for adults—tumor recurrence remains a major problem for both clinicians and patients. Therefore, earlier and more sensitive detection of residual tumor cells in the body after remission, known as minimal residual disease (MRD), is crucial to identifying signs of tumor recurrence earlier.
[0003] The immune repertoire can be used to monitor patients for signs of relapse. The immune repertoire refers to the sum of all functionally diverse T / B lymphocytes of an individual at a specific point in time. T / B cells use cell surface receptors (TCR / BCR) to recognize and bind to antigens, mediating the body's production of targeted immune responses. There is a region in TCR / BCR called the complementary determining region (CDR). Different CDR sequences can recognize different antigens, and it plays a key role in antigen recognition. CDR consists of CDR1, CDR2, and CDR3, of which CDR3 has the highest degree of diversity. CDR1 and CDR2 are encoded by V genes, and CDR3 is encoded by three genes in the variable region V, D, and J. The immune repertoire characteristics of patients with hematological tumors at the time of relapse are consistent with those at the time of initial diagnosis, that is, the content of a certain clonal type sequence of the immune repertoire gene is significantly increased. Based on this feature, the immune repertoire characteristics can be used to monitor whether patients have signs of relapse after remission, intervene in treatment earlier, improve the patient's quality of life, and increase survival rate.
[0004] There are also products on the market that use DNA capture and high-depth sequencing to monitor minimal residual disease. However, if the patient's mutation site does not fall within the design range of the DNA capture probe, the capture probe needs to be redesigned. Compared with immune repertoire technology, this is more expensive, time-consuming, and has a higher probability of false negatives.
[0005] There are two approaches to monitoring minimal residual disease in the immune repertoire. One involves capillary electrophoresis fragment analysis on a traditional first-generation sequencing platform; the other involves multiplex PCR amplification of immune repertoire sequences followed by high-throughput sequencing analysis. In comparison, first-generation sequencing methods have low sensitivity and are unable to distinguish between immune repertoire clonal types of similar length. Multiplex PCR and second-generation sequencing methods offer both sensitivity and a wide range of immune repertoire detection.
[0006] The immune repertoire contains genes such as IGH, IGK, IGL, TRB, TRD, and TRG. Among them, the IGH gene has the most clonal types, and its clonal characteristics are often used clinically to monitor tumor recurrence. However, the following problems exist when designing multiplex PCR primers for the IGH gene in the immune repertoire:
[0007] (1) There are many IGH clonal types, making it difficult to obtain a primer pool with comprehensive coverage and high working efficiency;
[0008] (2) The IMGT database is frequently updated, making it difficult to quickly update the primer pool adapted to new clonal types while maintaining compatibility with the original primer pool;
[0009] (3) The primer pool is difficult to adapt to different multiplex PCR amplification reagent systems. The amplification primers have different affinities for different multiplex PCR amplification reagents, making it difficult to quickly adjust the primer pool to adapt to various multiplex PCR amplification reagents;
[0010] (4) It is difficult to obtain the true content of clonotypes. Monitoring requires obtaining the true content ratio of different IGH clonotypes, but the amplification efficiency of primers varies. The content of amplified clonotypes is highly correlated with the amplification efficiency of primers, so it is difficult to obtain the true proportion of each IGH clonotype. If UMI tags are introduced, although the problem of quantitative accuracy can be solved, it will lead to a significant decrease in detection sensitivity, that is, a significant reduction in detection performance;
[0011] (5) The utilization rate of sequencing data of amplified products is not high, and some sequencing data cannot detect clonal types. When sequencing the amplicon sequence using a second-generation sequencing platform, the base sequence information of 150bp on both sides of the amplicon can be measured. The key CDR3 region is located in the middle of the amplicon. If the amplicon is too long, the CDR3 region is easily missed, resulting in the inability to detect the corresponding clonal type. If the amplicon is too short, the CDR3 region may be incomplete, resulting in the inability to confirm which clonal type it is. Summary of the Invention
[0012] The purpose of the present invention is to propose that it is necessary to provide a multiplex PCR primer set and design method for IGH gene clonotype detection to at least solve one of the above-mentioned deficiencies of the prior art.
[0013] In view of this, the scheme of the present invention is as follows:
[0014] In a first aspect of the present invention, a PCR primer set for detecting IGH gene clonotypes is provided, comprising primers having nucleotide sequences as shown in SEQ ID NOs: 1-57.
[0015] The second aspect of the present invention provides an IGH gene clonotype detection kit comprising the PCR primer set described in the first aspect.
[0016] The third aspect of the present invention provides the use of the PCR primer set described in the first aspect in preparing an IGH gene clonotype detection product.
[0017] A fourth aspect of the present invention provides a method for designing a multiplex PCR primer set for IGH gene clonotype detection, comprising:
[0018] A primer pool was designed based on the conserved region of the IGH gene;
[0019] Each primer is weighted and scored based on the scoring criteria, which include primer quality parameters, highly conserved base content, and the number of covered clones.
[0020] At least one primer with a score higher than the threshold is selected as the candidate primer for each clonotype;
[0021] Penalty screening is performed on each clonotype primer. Primers with a penalty score greater than 0 are replaced with primers with the next highest score and the screening process is repeated until the primer penalty score for each clonotype is 0, thereby obtaining a multiplex PCR primer set. Penalty criteria include primer dimer formation, excessive sequence length, and incomplete CDR3.
[0022] Furthermore, the primer quality parameters include primer length, primer GC content, primer terminal base type, primer terminal base type, primer base complexity, possibility of forming a hairpin structure, and whether there are five consecutive GC bases.
[0023] Preferably, the weighted scoring formula is: s = a + b + c + d + e - f + g + 5h + 2i; wherein:
[0024] a is the primer length score,
[0025] b value is the GC base content ratio of the primer sequence;
[0026] c is the score of the base at the end of the primer. If the base is A, then c = 0, otherwise c = 1;
[0027] d is the type score of the three bases at the end of the primer. If the three bases are abbreviated as SSA, the value of d is 1.25. If the three bases are abbreviated as SWC or SWG, the value of d is 1.5. In other cases, the value of d is 1.
[0028] e is the primer base complexity score, which is calculated based on kmer, with 0≤e≤1;
[0029] F is the probability score of forming a hairpin structure;
[0030] g is the score for whether there are 5 consecutive GC bases. If so, the value is 0, otherwise it is 1;
[0031] h is the highly conserved base content score, which is the ratio of the number of bases in the conserved region to the primer length;
[0032] i is the score for the number of covered clonotypes, which is the ratio of the number of covered clonotypes to the total number of clonotypes.
[0033] Furthermore, the method for judging whether the sequence length is too long is that the sum of the length of the amplified sequence corresponding to the V / D region primer and the length of the amplified sequence corresponding to the J region primer is greater than a length threshold.
[0034] Furthermore, the multiplex PCR primer set for detecting IGH gene clonotypes includes primers having nucleotide sequences as shown in SEQ ID NOs: 1-57.
[0035] A fifth aspect of the present invention provides a method for constructing a multiplex PCR amplification system for detecting IGH gene clonotypes, comprising the steps of primer design and concentration adjustment; the primers are designed using the primer set described in the first aspect, or according to the design method described in the fourth aspect.
[0036] Furthermore, the concentration adjustment process includes constructing plasmid sequences of different primers, adjusting the primer concentration by analyzing the actual content of the amplicon and the content of the amplified product corresponding to the different primers, thereby achieving accurate quantification while also ensuring the sensitivity of detection.
[0037] Compared with the prior art, the beneficial effects of the present invention include but are not limited to:
[0038] The multiplex PCR primer set described in the present invention can specifically amplify each clonotype sequence of the IGH gene, effectively detecting nearly all clonotypes, with high efficiency and coverage of IGH clonotypes, meeting the clinical need for monitoring tumor recurrence based on clonotype characteristics.
[0039] The primer set design method described in the present invention establishes an IGH candidate primer database by setting up a weighted scoring mechanism and a penalty mechanism for screening primers. When a new clonal type emerges or the amplification reagent changes, the primers can be quickly replaced. The length range of the amplification product corresponding to the F-terminal primer and the length range of the amplification product corresponding to the R-terminal primer are strictly controlled to ensure that various clonal type sequences are well amplified, sequenced, and detected. By constructing plasmid sequences with different primers and analyzing the actual content of the amplicon and the content of the amplification product corresponding to different primers, the primer concentration is adjusted to achieve accurate quantification while also ensuring detection sensitivity. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0041] Figure 1 Flowchart for designing the PCR primer pool for IGH gene clonotype detection according to the present invention.
[0042] Figure 2 Schematic diagram of the multiple sequence alignment results of the IGH gene according to the present invention.
[0043] Figure 3 Schematic diagram of primer amplification deviation before and after adjustment of the multiplex PCR amplification system of the present invention. DETAILED DESCRIPTION
[0044] The following provides definitions of some terms used in this specification. Unless otherwise specified, all terms used herein have the meanings commonly understood by those skilled in the art to which this solution belongs.
[0045] Explanation of terms:
[0046] Primer dimer: refers to the double-stranded DNA structure formed by two primers pairing with each other in the polymerase chain reaction (PCR), rather than the structure formed by the primers pairing with the target DNA template.
[0047] In one embodiment, the inventors discovered that, based on multiple scoring metrics, a high-quality primer pool for each IGH clonotype can be obtained, ensuring that the IGH gene primer pool can better cover all clonotypes and has good amplification efficiency; an IGH candidate primer database is established, and when new clonotypes appear or amplification reagents change, primers can be quickly replaced; strictly controlling the length range of amplification products corresponding to the F-terminal primers and the length range of amplification products corresponding to the R-terminal primers can ensure that various clonotype sequences are well amplified, sequenced, and detected; constructing plasmid sequences with different primers, and adjusting primer concentrations by analyzing the actual content of amplicons and amplification products corresponding to different primers, achieving accurate quantification while also ensuring detection sensitivity.
[0048] In the above embodiment, the primer design process is as follows Figure 1 As shown, the specific steps are as follows:
[0049] 1. Identification of IGH conserved regions
[0050] Download the reference sequences of each IGH clonotype from the IMGT database, perform multiple sequence alignment on these sequences, and obtain the conservation score of each base. Prioritize the conservative region as the target, use as few primers as possible to cover all clonotypes, so as to minimize the impact of primer amplification efficiency on clonotype quantification; prioritize the conservative region for primer design to avoid mutations in the primer region, which may cause the clonotype to be unable to be amplified. The multiple sequence alignment result is shown in the following example: Figure 2 shown.
[0051] 2. Construction of candidate primer pool
[0052] Primer3 was used to design primer sequences based on the conserved regions to obtain a candidate primer pool.
[0053] 3. Primer database construction
[0054] Each primer is scored based on primer length, primer GC content, primer terminal base type, primer terminal base type, primer base complexity, hairpin structure possibility score, presence of five consecutive GC bases, highly conserved base content, and number of clonal types covered. Primers for each clonal type are ranked from high to low based on their scores to form a candidate database of IGH gene primers for the immune repertoire. The database contains the aforementioned scores, the overall score, and the length of the corresponding amplified sequences.
[0055] The primer length score is recorded as a, The GC content score of the primer is recorded as b, and its value is equal to the GC base content ratio of the primer sequence. The score of the base at the end of the primer is recorded as c. If the base is A, then, and vice versa. The score of the three base types at the end of the primer is recorded as d. If the three bases are abbreviated as "SSA", the value of d is 1.25. If the three bases are abbreviated as "SWC" or "SWG", the value of d is 1.5. Other cases are recorded as 1 (W and S are degenerate bases, "W" represents the base "A" or "T", and "S" represents the base "C" or "G"). The primer base complexity score is recorded as e, and its score is calculated based on kmer, with a maximum of 1 and a minimum of 0. The score for the possibility of forming a hairpin structure is recorded as f, which is calculated by primer3. The higher the value, the more likely it is to form a hairpin structure. Whether there are 5 consecutive GC bases is scored as g. If so, the value is 0, otherwise it is 1. The score for highly conserved base content is recorded as h, The score of the number of covering clonotypes is denoted as i,
[0056] The total primer score s is scored as follows:
[0057] s=a+b+c+d+e-f+g+5h+2i.
[0058] 4. Select the initial IGH primer pool
[0059] The primer with the highest score for each clonotype in the database was taken as the initial primer pool.
[0060] 5. Initial Primer Pool Optimization
[0061] ① Primer dimer assessment. Use the software mfeprimer to assess primer dimers in the initial primer pool. For each primer involved in the formation of a primer dimer, the penalty score is increased by one.
[0062] ② Simulate the length of all clonal types and add the length of the amplified sequence corresponding to the V / D region primers and the length of the amplified sequence corresponding to the J region primers. If the added length is greater than 280 bp, the amplified product is considered to be too long and a penalty of one is added.
[0063] ③Simulate all clonal sequences and use igblast to evaluate them to ensure that the sequences amplified by the primers contain the complete CDR3 region. If the CDR3 is incomplete, the corresponding primer penalty is increased by one.
[0064] For each clonotype, primers with a penalty score greater than 0 were replaced with primers with suboptimal scores for that clonotype in the primer database.
[0065] 6. Iterative optimization of primer pool
[0066] Repeat step 5 of the optimization process until all primer penalties in the primer pool are 0, thus obtaining the final primer pool for the immune repertoire IGH gene. At this point, the primer pool contains no primer dimers, the primer amplification efficiency is good, and the amplification products of all clonal types meet the expected length and contain complete CDR3 sequences.
[0067] 7. Adjustment of Multiplex PCR Amplification System
[0068] Construct fragments amplified by different primers and insert them into the plasmid vector. The plasmid is then transformed into Escherichia coli. Positive clones are confirmed by antibiotic screening and sequencing to ensure that the inserted fragment sequence is correct. Plasmid DNA is extracted from the positive clones, and the concentration of the plasmid DNA is quantified using a spectrophotometer. The plasmid is diluted in a certain ratio to establish a standard curve. The final primer pool of the immune repertoire IGH gene is used to perform PCR amplification on plasmids of different dilutions. The concentration of the PCR product is then analyzed using qPCR. According to the concentration of the PCR product, the working concentration of each primer is adjusted so that the content of the clonal sequence amplification product is consistent with the actual concentration ratio.
[0069] The ratio of the amplified product concentration to the actual concentration is compared, and this is recorded as the deviation. As shown in the figure below, after the primer concentration is adjusted, the deviation between the amplified product and the actual concentration approaches 1, indicating that after the primer concentration is adjusted, the multiplex PCR amplification system can accurately quantify each IGH clonotype.
[0070] At this point, a multiplex PCR primer pool for the IGH gene in the immune repertoire has been generated, which can effectively amplify individual IGH clonotype sequences, ensuring sensitivity and accurate quantification. If new clonotypes emerge, the existing primer database is updated and steps 4-7 are repeated to quickly construct a new primer pool. If amplification reagents need to be replaced, repeat step 7 and adjust the concentrations of each primer to ensure accurate quantification with the new multiplex PCR amplification system.
[0071] In a preferred embodiment, the above-mentioned design method is used to obtain a series of multiplex PCR primer pools for IGH gene clonotype detection, as shown in Table 1.
[0072] Table 1: IGH multiplex PCR primer pool
[0073]
[0074]
[0075] Peripheral blood samples were selected for nucleic acid extraction, multiplex PCR primer amplification, and NGS sequencing to verify the effectiveness of the primer pool. The multiplex PCR primer amplification step used a two-round amplification protocol. For the first round, VAHTS Pathogen DNA & RNA Multiplex PCR Mix was dissolved on ice and vortexed to mix. The first-round PCR amplification system was prepared according to Table 2, and the amplification reaction was performed according to Table 3. The second round of PCR used the Novozymes VAHTS HiFi Amplification Mix, and VAHTS DNA Clean Beads were used for purification of PCR products.
[0076] Table 2: PCR amplification reaction system
[0077] Reagent components Volume (μL) 5×VAHTS RT Multi-PCR Mix 5 Primerpool-1-F / R 4 DNA template 10 <![CDATA[Nuclease-free H2O]]> 6 Total 25
[0078] Table 3: PCR amplification reaction program
[0079]
[0080] PCR product purification steps are as follows: VAHTS Pathogen DNA & RNA Multiplex PCR Mix Amplification System: Make up the PCR mixture to 50 μl with Nuclease-free ddH2O. Vortex VAHTS DNA Clean Beads to mix thoroughly. Add 50 μl of VAHTS DNA Clean Beads (1×) and pipette or vortex to mix thoroughly. Incubate at room temperature for 5 minutes. Centrifuge briefly. Place the PCR tube on a magnetic rack for 3 minutes to allow the solution to clear. Completely remove the supernatant. Remove the PCR tube from the magnetic rack and add 50 μl of Buffer YF to the tube. Pipet and mix thoroughly. Incubate at room temperature for 5 minutes. Keep the PCR tube on the magnetic rack and carefully discard the supernatant. Add 180 μl of 80% ethanol to the PCR tube and let it stand for 30 seconds. Keep the PCR tube on the magnetic rack and discard the supernatant. Add 180 μl of 80% ethanol to the PCR tube again. Let it stand for 30 seconds and discard the supernatant. Cap the tube and centrifuge briefly to remove any remaining ethanol to the bottom of the tube. Place the PCR tube on a magnetic rack and carefully remove any remaining ethanol with a 10 μL pipette, taking care not to absorb the magnetic beads. Add 30 μL of NFW, remove the PCR tube from the magnetic rack, pipette or vortex to mix, and let it stand at room temperature for 2 minutes. Use a pipette to aspirate 28 μL of the supernatant and transfer it to a new PCR tube. The supernatant in the tube is the prepared first-round PCR purification product. Take 1 μL and measure the concentration using a Qubit 4.0 Fluorometer (Qubit dsDNA HS Assay Kit) and record the product concentration.
[0081] The working conditions of each primer are shown in Table 4. It is not difficult to see that all primers work normally.
[0082] Table 4: Primer performance
[0083]
[0084]
[0085] The detection of individual IGH gene clonotypes is shown in Table 5. Over 90% of reads detected the CDR3 region of the IGH gene, enabling clonotype determination. The robust detection of nearly all clonotypes demonstrates the high efficiency of this primer pool and its high coverage of IGH gene clonotypes.
[0086] Table 5: Statistics of the number of reads for IGH gene clonotype detection
[0087]
[0088]
[0089]
[0090] Although the present invention is disclosed as above, the scope of protection disclosed by the present invention is not limited thereto. Those skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will fall within the scope of protection of the present invention.
Claims
1. A PCR primer set for detecting IGH gene clonotypes, characterized in that: Primers having nucleotide sequences as shown in SEQ ID NOs: 1-57 are included.
2. A kit for detecting IGH gene clonotypes comprising the PCR primer set according to claim 1.
3. Use of the PCR primer set according to claim 1 in preparing an IGH gene clonotype detection product.
4. A method for constructing a multiplex PCR amplification system for IGH gene clonotype detection, characterized in that: The method comprises using the PCR primer set according to claim 1 and adjusting the primer concentration.
5. The construction method according to claim 4, characterized in that: The primer concentration adjustment process includes: constructing a plasmid containing different primer amplification fragments, and adjusting the primer concentration by analyzing the actual content of the amplicon corresponding to the different primers and the content of the amplification product.
Citation Information
Patent Citations
Preparation method and application of library, linker and kit
CN117822130A
Nucleic acid amplification primers for PCR-based clonality studies
CN1965089A