CDR3 simulation tag and application thereof in evaluating working efficiency of primer in immune repertoire

By constructing a CDR3-based simulated tag system and combining it with homologous recombination technology to link primers, the problems of primer amplification bias and detection sensitivity in multiplex PCR were solved, enabling high-precision immune repertoire analysis.

CN120924563APending Publication Date: 2025-11-11WUHAN XINO MEDICAL LABORATORY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511086109.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing multiplex PCR methods suffer from uneven amplification and decreased detection sensitivity due to primer amplification bias in immune repertoire detection. Furthermore, existing technologies have issues with synthesizing a large number of gene sequences and insufficient starting amounts when evaluating the amplification efficiency of different primer combinations.

Method used

A CDR3 simulated tag system was constructed. By synthesizing V(D)J artificial templates with known proportions and combining amplification efficiency analysis of different primer combinations, primer preference was quantitatively evaluated and primer concentration ratios were optimized. Homologous recombination technology was used to connect upstream and downstream primers, avoiding enzyme digestion and recovery steps, and enabling flexible evaluation of primer efficiency.

Benefits of technology

It reduces amplification bias, improves detection accuracy and sensitivity, simplifies calculation methods, reduces human error, and is suitable for multiple detection needs of TCR and BCR, providing a methodological basis for high-precision immune repertoire analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120924563A_ABST
    Figure CN120924563A_ABST
Patent Text Reader

Abstract

The invention discloses a CDR3 simulation tag and application thereof in evaluating the working efficiency of primers in an immune repertoire, the CDR3 simulation tag can be randomly connected with an upstream primer and a downstream primer, and the cost is reduced; in the process of evaluating the working efficiency of the primers by the immune repertoire, unnecessary calculation errors caused by excessive calculation can be avoided while the preference problem is solved through a calculation analysis method, so that result inaccuracy caused by normal errors on an experiment and a sequencer is avoided; the method can avoid the problem that the initial quantities of different primer templates are different due to introduction of restriction enzyme, reduces artificial introduction errors, can be compatible with multiple detection requirements of TCR and BCR at the same time, and provides a methodological basis for high-precision immune repertoire analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of immune repertoire detection technology, specifically to the CDR3 mimic tag and its application in evaluating primer efficiency in immune repertoire testing. Background Technology

[0002] An immune repertoire is the diverse collection of all T-cell and B-cell receptors (TCRs and BCRs) within an individual, reflecting the adaptive immune system's potential to recognize antigens. T-cell receptors (TCRs) consist of α and β chains (or γ and δ chains), while B-cell receptors (BCRs) consist of heavy chains (IgH) and light chains (Igκ / Igλ). Both form variable regions through recombination of V (Variable), D (Diversity, present only in the BCR heavy chain and TCR β / δ chains), and J (Joining) gene segments. The complementarity-determining region 3 (CDR3) is a key region for direct antigen contact. For example, the CDR3 of the TCR β chain is encoded by VDJ gene recombination, while the CDR3 of the BCR heavy chain is determined by both VDJ rearrangement and N nucleotide insertion. The genetic composition of immune repertoires is highly diverse; the human genome contains approximately 50 TCR V genes and over 100 BCR V genes, which can generate more than 1000 different types through recombination. 15 A unique receptor.

[0003] The core of immune repertoire diversity stems from the random recombination, linkage flexibility, and somatic high-frequency mutation (SHM) of the V(D)J gene. First, during lymphocyte development, V, D, and J gene fragments randomly combine via recombinases (RAG1 / RAG2) to form variable region sequences. Second, during gene fragment linkage, exonucleases randomly cleave and insert template-independent N nucleotides, further increasing sequence complexity. Furthermore, the BCR undergoes SHM upon antigen stimulation, optimizing antigen affinity through point mutations. For example, the CDR3 diversity of the TCR β chain is mainly contributed by VDJ recombination and N insertion, while the diversity of the BCR heavy chain also involves multiple reading frame selections of the D gene and SHM. These mechanisms work together to enable the immune repertoire to recognize a near-infinite number of antigen types, forming the core foundation of immune defense.

[0004] Currently, the main methods for immune repertoire analysis based on next-generation sequencing (NGS) include multiplex PCR and 5'RACE. Multiplex PCR: This method designs dozens to hundreds of primer pairs to specifically amplify the VJ or VDJ regions of the TCR / BCR. Its advantage lies in its compatibility with both DNA and RNA templates, but it is susceptible to primer amplification bias, leading to insufficient coverage of some V / J genes. 5'RACE: Using mRNA as a template, it uses reverse transcription primers to anchor the constant region (C region) and amplifies the full-length variable region sequence using a single primer. This method avoids primer bias in multiplex PCR and is particularly suitable for full-length capture of BCR / TCR, but it relies on high-quality RNA and cannot analyze DNA templates. Both methods require the integration of NGS technology to analyze clonal composition and diversity characteristics by sequencing the CDR3 region.

[0005] Although multiplex PCR is the mainstream technique for immune repertoire detection, limitations in its primer design (such as differences in primer-template binding efficiency) can lead to uneven amplification, significantly affecting the true distribution of clone frequencies. For example, some V genes may experience low amplification efficiency due to primer mismatches or abnormal GC content, resulting in decreased detection sensitivity.

[0006] Patent document CN111850016A discloses immune repertoire standard materials corresponding to immune repertoire standard material sequences. By artificially simulating the CDR3 sequence and combining it with different combinations of V, J, and C genes, this standard material is used to simulate a T / B cell receptor library, allowing for the optimization of multiplex PCR primer systems to reduce bias. However, this approach has two significant drawbacks: 1. Each independently synthesized gene sequence can only be amplified with a single primer pair, which is not conducive to evaluating multiple primer combinations in a multiplex PCR system. Specifically, existing technologies synthesize all the required V, J, and C genes, along with the intermediate simulated CDR3 sequence, and then construct them into vectors via homologous recombination to form different template sequences. When used as templates for amplification, each vector contains only one sequence of the V, J, and C genes, thus only a single pair of upstream and downstream primers can be matched. The more primers there are, the more combinations of upstream and downstream primers there are, and the more gene sequences need to be synthesized. This significantly limits the evaluation of flexible combinations between different primers. 2. After single-enzyme digestion and recovery of the synthesized gene sequences, the starting amount of each template is greatly reduced, leading to significant uncertainty in subsequent quantification and primer efficiency evaluation. Specifically, when using vectors constructed with existing techniques as templates for amplification, equal amounts of the required template sequences are mixed, fragments are cut using a single enzyme digestion method, and then the fragments are mixed again. This process is inevitably affected by the activity of restriction endonucleases and the efficiency of fragment recovery.

[0007] Therefore, it is necessary to improve existing technologies so that they can be evaluated more flexibly and accurately in multiplex PCR systems after free combination of different primer pairs. Summary of the Invention

[0008] To address the problems of the prior art, this invention constructs an evaluation system that allows for flexible replacement of primer sequences. This system synthesizes V(D)J artificial templates in known proportions and, combined with amplification efficiency analysis of different primer combinations, quantitatively assesses primer bias and optimizes primer concentration ratios. Results show that this method can reduce amplification bias and simultaneously meet the multiple detection requirements of TCR and BCR, providing a methodological basis for high-precision immune repertoire analysis.

[0009] In view of this, the solution of the present invention is as follows: A first aspect of the present invention is to provide a CDR3 mimic tag, the sequence of which comprises, from the 5' end to the 3' end, an upstream primer sequence, a V gene mimic sequence, a CDR3 mimic sequence, a J gene mimic sequence, and a downstream primer sequence; the CDR3 mimic sequence includes a non-human tag sequence and a numbering sequence; the non-human tag sequence is used to distinguish whether the sequence belongs to a vector sequence or a human sequence during amplification, the system identification sequences are B(b) and J(j) sequences, and the numbering sequence is used to distinguish the source of the sequence.

[0010] The CDR3 simulated tag sequence of this invention is designed with reference to the TCR (TRB VJ, TRB DJ, TRG, and TRD) and BCR (IGH, IGDH, IGK, and IGL) gene sequences, combined with the distribution patterns of TCR / BCR in healthy individuals. By artificially simulating the CDR3 sequence and combining it with different combinations of simulated V and J gene sequences, a simulated lymphocyte receptor library was constructed. Simultaneously, the proportions of the multiplex PCR primer system were adjusted and optimized to reduce experimental bias.

[0011] In a preferred embodiment, each system identification sequence is 5-30 bp in length, and the numbering sequence is 6-15 bp in length. The above-mentioned starting amino acid, system identification sequence, numbering sequence, and termination amino acid combination together simulates the real sample CDR3.

[0012] In a preferred embodiment, the simulated sequence of the TCR encoding gene CDR3 is CCINQB(b)J(j)DVW or CCINQB(b)J(j)DVF from the 5' end to the 3' end. And / or, the simulated sequence of the BCR encoding gene CDR3 from the 5' end to the 3' end is CCINQB(b)J(j)DVF or CCINQB(b)J(j)DVWGXG; in: C represents the nucleotide corresponding to the codon of cysteine, the CDR3 initiation conserved amino acid; W indicates the nucleotide corresponding to the codon of the conserved amino acid tryptophan, which is terminated by CDR3. F indicates the nucleotide corresponding to the codon of the conserved amino acid phenylalanine, which is terminated by CDR3. G indicates the nucleotide corresponding to the codon of glycine in the CDR3-terminated conserved amino acid motif; CINQDV represent the nucleotides corresponding to the codons of cysteine, isoleucine, asparagine, glutamine, aspartic acid, and valine, respectively. B and J can be independently selected from adenine, guanine, cytosine, or thymine; X is any amino acid; b and j are selected from natural numbers from 5 to 30.

[0013] In some embodiments, when the CDR3 analog tag sequence is an IGH sequence, the system identification sequence is selected from the nucleotide sequences shown in SEQ ID NO: 1-12; And / or, when the CDR3 simulated tag sequence is a TRB sequence, the system recognition sequence is selected from the nucleotide sequences shown in SEQ ID NO: 13-34; In some embodiments, the probability of the CDR3 simulated label sequence appearing in real samples is ≤0.0001%, which can effectively distinguish it from real samples.

[0014] In some embodiments, the GC content of the immune repertoire simulated sequence corresponding to the CDR3 simulated tag sequence is 0.58±0.07, and the A, T, C, and G bases cannot appear consecutively for more than 3 times.

[0015] In some embodiments, the sequence from the 5' end to the 3' end is as follows: upstream of the 5' end of the upstream primer is sequentially composed of a system backbone homology arm 1 and a restriction endonuclease site A, and downstream of the 3' end of the upstream primer is sequentially composed of a restriction endonuclease B and a system backbone homology arm 2. And / or, the downstream primer contains, upstream of the 5' end, a system backbone homology arm 3 and a restriction endonuclease site C, and downstream of the 3' end, a restriction endonuclease D and a system backbone homology arm 4.

[0016] Preferably, the total length of any homologous arms 1-4 of the system backbone is approximately 30 bp; the restriction endonuclease cleavage site is located approximately 10 bp from the 3' or 5' end of the homologous arm of the adjacent system backbone structure.

[0017] In some embodiments, the simulated V gene sequence and the simulated J gene sequence are obtained as follows: 1) Alignment with reference sequences: Align the known human BCR and TCR gene sequences to select the sequences with the highest similarity among human sequences, while ensuring that the selected sequences are functional genes and open reading frame regions, retaining about 100 bp from the 3' end of the V gene and about 100 bp from the 5' end of the J gene; 2) CDR3 sequence position determination: The 5' start amino acid of the CDR3 sequence of the BCR and TCR coding genes V gene reference sequence is conserved cysteine ​​(C), the 3' end of the CDR3 sequence of the TCR coding gene J gene is conserved tryptophan (W) / phenylalanine (F), and the 3' end of the CDR3 sequence of the BCR coding gene J gene is conserved phenylalanine (F) or tryptophan-glycine-X-glycine (WGXG, where X is any amino acid) motif; The steps of the CDR3 simulated tag design method for evaluating the working efficiency of primers in immune repertoires are as follows: 1) Download the reference sequences of genes V, D, and J from IMGT or a database containing complete BCR / TCR related sequences; 2) Only retain the CDR3 sequence between the start codon and the stop codon; 3) Insert a CDR3 simulated tag sequence between the start codon and the stop codon; 4) The CDR3 mimic tag sequence includes, from the 5' end to the 3' end, the following sequence in order: system backbone homology arm 1, restriction site A, upstream primer sequence, restriction site B, system backbone homology arm 2, V gene mimic sequence, CDR3 mimic sequence, J gene mimic sequence, system backbone homology arm 3, restriction site C, downstream primer sequence, restriction site D, and system backbone homology arm 4.

[0018] This invention also discloses a method for preparing CDR3 simulated tags to evaluate the working efficiency of different primers, comprising the following steps: 1) The upstream portion of the CDR3 mimic tag contains the upstream primer. The 5' end of the upstream primer contains the system backbone homology arm 1 and restriction endonuclease site A in sequence, and the 3' end contains restriction endonuclease B and system backbone homology arm 2 in sequence.

[0019] 2) The downstream portion of the CDR3 mimic tag contains a downstream primer. The 5' end of the complementary strand of the downstream primer contains, in sequence, a system backbone homology arm 3 and a restriction endonuclease site C, and the 3' end contains, in sequence, a restriction endonuclease D and a system backbone homology arm 4.

[0020] 3) According to experimental requirements, the 5' end of the CDR3 simulated sequence also contains different synthesized V gene simulated sequences, and the 3' end contains different synthesized J gene simulated sequences.

[0021] 4) Introduce the 5' end of the upstream primer into the 3' end 10bp sequence of homologous arm 1, and introduce the 3' end of the upstream primer into the 5' end 10bp sequence of homologous arm 2.

[0022] 5) The vector is cut with restriction endonuclease A and restriction endonuclease B to expose a gap, and the upstream primer is recombined into the vector according to the homologous recombination method.

[0023] 6) In step 5), the cyclic product is transformed into DH5α competent Escherichia coli cells, shaken at 37°C for 2 hours, then spread on plates containing the corresponding antibiotics, and incubated in an oven at 37°C upside down overnight. 7) After culturing for 10 hours, select single clones, perform bacterial PCR and Sanger sequencing, and expand the culture of the single clones with correct sequencing to extract plasmid DNA.

[0024] 8) Introduce a 10bp sequence from the 3' end of the 5' end of the complementary strand of the downstream primer into the 3' end of the homologous arm 3, and introduce a 10bp sequence from the 5' end of the 3' end of the complementary strand of the downstream primer into the 5' end of the homologous arm 4.

[0025] 9) Cut the vector in which the upstream primer was successfully recombined in step 7) using restriction endonuclease C and restriction endonuclease D to expose the gap, and then recombined the downstream primer into the vector according to the homologous recombination method.

[0026] 10) Repeat steps 6) and 7), and you will obtain a CDR3 analog tag system with specific upstream and downstream primers.

[0027] The reason for introducing different sequences into the CDR3 simulated tag sequence (e.g., sequences B(b) and J(j)) is to avoid the influence of human factors causing a single intermediate sequence to affect amplification efficiency, and to maximize diversity. Combining different F and R primers allows for a more realistic evaluation of the amplification efficiency of primer combinations while meeting diverse amplification needs. Alternatively, depending on actual requirements, the diversity of sequences B(b) and J(j) can be reduced, using the same CDR3 simulated tag sequence, replacing only the different F and R primers, to evaluate the amplification efficiency between primers.

[0028] Understandably, after fixing one upstream or downstream primer, the downstream or upstream primer can be flexibly replaced through homologous recombination to achieve rapid evaluation of differences in primer amplification efficiency. At the same time, different upstream and downstream primer sets can be flexibly added to the immune repertoire detection system as internal references for amplification correction.

[0029] After constructing the aforementioned artificial CDR3 simulated tag sequences, we used a data-fitted method to evaluate primer efficiency, determining the bias of each primer and providing reasonable adjustments to avoid amplification bias caused by primer bias. In the first primer efficiency test, primers to be used for testing were mixed together at a working solution concentration of 100 nM. Different CDR3 simulated tag sequences were amplified in the first round of PCR, with the same copy number of each CDR3 simulated tag sequence plasmid (same amount of each template). The sequences were then ligated into sequencing adapters for NGS sequencing in the second round of PCR. Based on the number of plasmids used, the theoretical data proportion of each primer, i.e., the theoretical abundance result, could be calculated. In actual sequencing, there are differences in amplification efficiency between different primer pairs; therefore, the actual data proportion of each primer pair, i.e., the relative abundance result, could be calculated. Each upstream primer will yield multiple different theoretical and relative abundance results due to pairing with various downstream primers. The efficiency deviation value of each primer can be calculated using the efficiency factor formula, and the concentration adjustment formula can be used to calculate the required concentration adjustment for each primer. Repeat the above steps with the adjusted concentrations to obtain a new number of reads for each plasmid. Data before and after primer concentration adjustment are homogenized for comparison, allowing evaluation of whether the effect of adjusting primer concentrations meets expectations. This invention significantly surpasses other similar methods in addressing primer preference issues, providing clear experimental, computational, and analytical logic, and avoiding the need to use fragmented fragments as amplification templates. This greatly mitigates the problem of varying starting amounts of different primer templates caused by the introduction of restriction endonucleases. Furthermore, it significantly reduces human-introduced errors in evaluating primer amplification efficiency. Moreover, this system and its computational logic can be parallelly ported to other multiplex PCR projects, solving a widespread problem from specific points to a broader scope.

[0030] The second aspect of this invention is to propose the application of the CDR3 mimicry tag described in the first aspect above in the evaluation of the working efficiency of immune repertoire primers or in the preparation of immune repertoire detection products.

[0031] Furthermore, the CDR3 mimic tag in the immune repertoire detection product exists in the form of circular or linear, single-stranded or double-stranded DNA.

[0032] A third aspect of this invention is to provide a method for evaluating the working efficiency of primers for an immune repertoire, comprising: Multiple CDR3 mimic tag sequence plasmids were constructed, each containing different upstream and downstream primers; the CDR3 mimic tag is the CDR3 mimic tag described in the first aspect; Primer efficiency was evaluated based on amplification data from multiple primers. The evaluation process was as follows: Calculate relative abundance: the ratio of the number of sequencing reads for any primer pair to the total number of reads, where the total number of reads is the sum of the number of sequencing reads for all upstream and downstream primer combinations; Theoretical abundance is calculated as one-tenth of the total number of plasmids used for primer efficiency evaluation. Calculate the efficiency factors of upstream and downstream primers separately: the product of the ratio of the relative abundance of all upstream or downstream primers to the theoretical abundance is raised to the power of 1 / 2 as the corresponding efficiency factor. The ratio of the original primer concentration to the efficiency factor is calculated as the adjusted primer concentration.

[0033] Furthermore, the total plasmid concentration of the CDR3 simulated tag is 10. 2 Up to 10 15 Copy / μL.

[0034] Compared with the prior art, the present invention has the following beneficial effects: The CDR3 mimic tag provided by this invention allows for the arbitrary connection of upstream and downstream primers based on the principle of homologous recombination after synthesizing a CDR3 mimic tag sequence. Both the time cost and the economic cost of gene synthesis are far lower than existing technologies.

[0035] The computational analysis method for evaluating the working efficiency of immune repertoire primers proposed in this invention greatly simplifies the calculation method while ensuring the resolution of primer preference issues, avoiding unnecessary computational errors introduced by over-calculation, and thus avoiding inaccurate results caused by normal errors in experiments and sequencing instruments. It addresses primer preference issues with clear experimental, computational, and analytical logic, and eliminates the need to use fragmented fragments as amplification templates, greatly avoiding the problem of different starting amounts of different primer templates caused by the introduction of restriction endonucleases. Simultaneously, it significantly reduces human-introduced errors in primer amplification efficiency evaluation, and can simultaneously meet the multiple detection requirements of TCR and BCR, providing a methodological foundation for high-precision immune repertoire analysis. Attached Figure Description

[0036] Figure 1 This is a schematic diagram of the immune repertoire CDR3 simulated tag sequence lymphoid clone analysis system described in Example 1.

[0037] Figure 2 This is a schematic diagram showing the deviations of the IGH gene amplification before (A), after (B), and amplification efficiency (C) in the experimental system of Example 4.

[0038] Figure 3 This is a schematic diagram showing the deviations of TRB gene amplification before (A), after (B), and amplification efficiency (C) in the experimental system of Example 5. Detailed Implementation

[0039] The technical solution of the present invention will now be clearly and completely described in conjunction with preferred embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0040] Example 1

[0041] Provide CDR3 analog tag sequences such as Figure 1 As shown, the system can use the same V gene, CDR3 mimic sequence, and J gene sequence, requiring only the upstream and downstream primers to be changed. Furthermore, when it is necessary to simultaneously evaluate the amplification efficiency of a particular upstream primer against multiple downstream primers, only the downstream primer sequences need to be replaced for subsequent operations; there is no need to resynthesize new gene sequences. This significantly reduces time and economic costs. Additionally, the CDR3 mimic sequence can be used to distinguish different primer types. When using the constructed system as a template for amplification, there is no need for enzyme digestion and fragment recovery, greatly improving the accuracy of quantification and primer efficiency evaluation (details can be found in Examples 2 and 3).

[0042] Artificially synthesized CDR3 mimic tag sequences from immune repertoires, such as Figure 1 As shown. The operation is as follows: 1) The CDR3 mimic tag sequence from 5' to 3' includes: system backbone homology arm 1, restriction site A, upstream primer sequence, restriction site B, system backbone homology arm 2, V gene mimic sequence, CDR3 mimic sequence, J gene mimic sequence, system backbone homology arm 3, restriction site C, downstream primer sequence, restriction site D, and system backbone homology arm 4. 2) Download the reference sequences of the V, D, and J genes of the IGH, IGK, IGL, TRB, TRD, and TRG genes from IMGT or a database containing complete BCR / TCR related sequences; 3) Select the human sequence with the highest similarity, while ensuring that the selected sequence is a functional group and an open reading frame region, retaining about 100 bp from the 3' end of the V gene and about 100 bp from the 5' end of the J gene; 4) CDR3 sequence position: The 5' start amino acid of the V gene reference sequence is conserved cysteine ​​(C), the 3' end of the TCR coding gene CDR3 sequence J gene is conserved tryptophan (W) / phenylalanine (F), and the 3' end of the BCR coding gene CDR3 sequence J gene is conserved phenylalanine (F) or tryptophan-glycine-X-glycine (WGXG, where X is any amino acid) motif. 5) The CDR3 mimic sequence is retained. From the start amino acid and the end amino acid sequence, a 100bp V gene sequence is included at the 5' end of the CDR3 mimic tag sequence, and a 100bp J gene sequence is included at the 3' end of the CDR3 mimic tag sequence. A unique mimic CDR3 sequence is added between the start amino acid and the end amino acid. The CDR3 mimic sequence of the TCR encoding gene is CCINQB(b)J(j)DVW or CCINQB(b)J(j)DVF from the 5' end to the 3' end. The CDR3 mimic sequence of the BCR encoding gene is CCINQB(b)J(j)DVF or CCINQB(b)J(j)DVWGXG from the 5' end to the 3' end. CINQDV is an amino acid sequence, which is an abbreviation of CinoDx of Shanghai Xinnobaishi Medical Laboratory Co., Ltd., and can also be replaced by other sequences with a length of 3-15 amino acids or 9-45 nucleotides to serve as a tag.

[0043] Wherein, C represents the nucleotide of the codon corresponding to the CDR3 initiation conserved amino acid cysteine; W indicates the nucleotide corresponding to the codon of the conserved amino acid tryptophan, which is terminated by CDR3. F indicates the nucleotide corresponding to the codon of the conserved amino acid phenylalanine, which is terminated by CDR3. G indicates the nucleotide corresponding to the codon of glycine in the CDR3-terminated conserved amino acid motif; CINQDV represent the nucleotides corresponding to the codons of cysteine, isoleucine, asparagine, glutamine, aspartic acid, and valine, respectively. B and J can be independently selected from: adenine, guanine, cytosine, or thymine; X is any amino acid; b and j are selected from natural numbers between 5 and 30. 6) The GC content of the simulated sequence of the immune repertoire is 0.58±0.07, and the A, T, C and G bases cannot appear consecutively for more than 3 times.

[0044] Example 2

[0045] A method for evaluating the working efficiency of primers for IGH immune repertoires is proposed, as follows: The CDR3 simulated tag sequence of the IGH chain is CCINQB(b)J(j)DVF, and its length is within the range of the actual rearranged sequence. In this example, the restriction endonuclease A used is EcoRI (GAATTC), restriction endonuclease B is KpnI (GGTACC), restriction endonuclease C is SacI (GAGCTC), and restriction endonuclease D is Xhol (CTCGAG). The B(b)J(j) sequence is shown in Table 1, and the examples shown in SEQ ID NO: 1-12 are not limited to the sequences shown. All sequences that conform to the description in Example 1 are within the scope of protection of this sequence.

[0046] Table 1. IGH CDR3 Simulated Tag B(b)J(j) Sequence

[0047] The specific steps are as follows: 1) The upstream portion of the CDR3 mimic tag contains the upstream primer IGHF1. The 5' end of the upstream primer contains, in sequence, the system backbone homology arm 1 and the restriction endonuclease EcoRI cleavage site, and the 3' end contains, in sequence, the restriction endonuclease KpnI cleavage site and the system backbone homology arm 2; 2) The downstream portion of the CDR3 mimic tag contains the downstream primer IGHR1. The 5' end of the complementary strand of the downstream primer contains, in sequence, the system backbone homology arm 3 and the restriction endonuclease SacI cleavage site, and the 3' end contains, in sequence, the restriction endonuclease Xhol cleavage site and the system backbone homology arm 4; 3) The 5' end of the CDR3 simulated sequence also contains synthetically produced different V gene simulated sequences, and the 3' end contains synthetically produced different J gene simulated sequences; 4) The 5' end of the upstream primer IHF1 was introduced into the 3' end 10bp sequence of homologous arm 1, and the 3' end of the upstream primer was introduced into the 5' end 10bp sequence of homologous arm 2. The complementary strands were synthesized in a similar manner and mixed at room temperature for subsequent steps. 5) The vector was cut with restriction endonucleases EcoRI and KpnI to expose a gap, and the upstream primer was recombined into the vector according to the homologous recombination method; 6) In step 5), the cyclic product is transformed into DH5α competent Escherichia coli cells, shaken at 37°C for 2 hours, then spread on plates containing the corresponding antibiotics, and incubated in an oven at 37°C upside down overnight. 7) After culturing for 10 hours, select single clones, perform bacterial PCR and Sanger sequencing, and expand the single clones with correct sequencing before extracting plasmid DNA. 8) The 5' end of the complementary strand of the downstream primer IGHR1 was introduced into the 3' end 10bp sequence of homologous arm 3, and the 3' end of the complementary strand of the downstream primer IGHR1 was introduced into the 5' end 10bp sequence of homologous arm 4. The complementary strands were synthesized in a similar manner and mixed at room temperature for subsequent steps. 9) The vector in which the upstream primer IGH1 was successfully recombined in step 7) was cut with restriction endonuclease SacI and restriction endonuclease Xhol to expose the gap, and the downstream primer IGH1 was recombined into the vector according to the homologous recombination method. 10) Repeat steps 6) and 7), and you will obtain a CDR3 analog tag system with a specific upstream primer IGHF1 and a downstream primer IGHR1. In addition, the CDR3 analog tag sequence of the IGH chain can also be CCINQB(b)J(j)DVWGXG. The design can be completed in the same way as described above, only replacing CCINQB(b)J(j)DVF with this sequence, while keeping other structures unchanged.

[0048] Obtaining a CDR3 mimic tag sequence using the improved technique requires only two steps: 1. Synthesize a CDR3 mimic tag sequence; 2. Based on the principle of homologous recombination, arbitrarily connect the upstream and downstream primers. This step is significantly less time-consuming and less costly than existing techniques for gene synthesis.

[0049] The artificially synthesized CDR3 mimicry tag from the immune repertoire exists in circular or linear DNA form, with a total concentration of 10-1 of the artificially synthesized CDR3 mimicry tag from the immune repertoire. 2 Up to 10 15 Copy / μL; The reason for introducing different sequences into the CDR3 simulated tag sequence (e.g., sequences B(b) and J(j)) is to avoid the influence of human factors causing a single intermediate sequence to affect amplification efficiency, and to maximize diversity. Combining different F and R primers allows for a more realistic evaluation of the amplification efficiency of primer combinations while meeting diverse amplification needs. Alternatively, depending on actual requirements, the diversity of sequences B(b) and J(j) can be reduced, using the same CDR3 simulated tag sequence, replacing only the different F and R primers, to evaluate the amplification efficiency between primers.

[0050] Example 3

[0051] A method for evaluating the working efficiency of TRB immune repertoire primers, comprising the following steps: The CDR3 simulated tag sequence of the TRB chain is CCINQB(b)J(j)DVW, and its length is within the range of the actual rearranged sequence. In this example, the restriction endonuclease A used is NcoI (CCATGG), restriction endonuclease B is NdeI (CATATG), restriction endonuclease C is SphI (GCATGC), and restriction endonuclease D is XbaI (TCTAGA). The B(b)J(j) sequence is shown in Table 2. Examples of SEQ ID No: 13-SEQ ID No: 34 are provided. The sequences are not limited to those shown and are consistent with the description in Example 1 and are all within the scope of protection of this sequence.

[0052] Table 2. TRB CDR3 Simulated Tag B(b)J(j) Sequence

[0053] The specific steps are as follows: 1) The upstream portion of the CDR3 mimic tag contains the upstream primer TRBF1. The 5' end of the upstream primer contains, in sequence, the system backbone homology arm 1 and the restriction endonuclease Ncol cleavage site, and the 3' end contains, in sequence, the restriction endonuclease Ndel cleavage site and the system backbone homology arm 2; 2) The downstream portion of the CDR3 mimic tag contains the downstream primer TRBR1. The 5' end of the complementary strand of the downstream primer contains, in sequence, the system backbone homology arm 3 and the restriction endonuclease site SphI, and the 3' end contains, in sequence, the restriction endonuclease Xbal cleavage site and the system backbone homology arm 4; 3) The 5' end of the CDR3 simulated sequence also contains synthetically produced different V gene simulated sequences, and the 3' end contains synthetically produced different J gene simulated sequences; 4) The 5' end of the upstream primer TRBF1 was introduced into the 3' end 10bp sequence of homologous arm 1, and the 3' end of the upstream primer was introduced into the 5' end 10bp sequence of homologous arm 2. The complementary strands were synthesized in a similar manner and mixed at room temperature for subsequent steps. 5) The vector was cut with restriction endonucleases NcoI and Ndel to expose a gap, and the upstream primer was recombined into the vector according to the homologous recombination method; 6) In step 5), the cyclic product is transformed into DH5α competent Escherichia coli cells, shaken at 37°C for 2 hours, then spread on plates containing the corresponding antibiotics, and incubated in an oven at 37°C upside down overnight. 7) After culturing for 10 hours, select single clones, perform bacterial PCR and Sanger sequencing, and expand the single clones with correct sequencing before extracting plasmid DNA. 8) The 5' end of the complementary strand of the downstream primer TRBR1 was introduced into the 3' end 10bp sequence of homologous arm 3, and the 3' end of the complementary strand of the downstream primer TRBR1 was introduced into the 5' end 10bp sequence of homologous arm 4. The complementary strands were synthesized in a similar manner and mixed at room temperature for subsequent steps. 9) The vector in which the upstream primer TRBF1 was successfully recombined in step 7) was cut with restriction endonuclease SphI and restriction endonuclease Xbal to expose the gap, and the downstream primer TRBR1 was recombined into the vector according to the homologous recombination method. 10) Repeat steps 6) and 7), and you will obtain a CDR3 analog tag system with a specific upstream primer TRBF1 and a downstream primer TRBR1. In addition, the CDR3 analog tag sequence of the TRB chain can also use CCINQB(b)J(j)DVF or, as described above, complete the design by replacing CCINQB(b)J(j)DVW with this sequence, without changing other structures.

[0054] Obtaining a CDR3 mimic tag sequence using the improved technique requires only two steps: 1. Synthesize a CDR3 mimic tag sequence; 2. Based on the principle of homologous recombination, arbitrarily connect the upstream and downstream primers. This step is significantly less time-consuming and less costly than the original technique.

[0055] The artificially synthesized CDR3 mimicry tag from the immune repertoire exists in circular or linear DNA form, with a total concentration of 10-1 of the artificially synthesized CDR3 mimicry tag from the immune repertoire. 2 Up to 10 15 Copy / μL; The reason for introducing different sequences into the CDR3 simulated tag sequence (e.g., sequences B(b) and J(j)) is to avoid the influence of human factors causing a single intermediate sequence to affect amplification efficiency, and to maximize diversity. Combining different F and R primers allows for a more realistic evaluation of the amplification efficiency of primer combinations while meeting diverse amplification needs. Alternatively, depending on actual requirements, the diversity of sequences B(b) and J(j) can be reduced, using the same CDR3 simulated tag sequence, replacing only the different F and R primers, to evaluate the amplification efficiency between primers.

[0056] Example 4

[0057] Based on the cloning analysis system in Example 2, the efficiency and correction of IGH multiplex PCR primers were evaluated.

[0058] 1) Quantitative

[0059] According to the method in Example 2, different primers were constructed into the same or one system to form artificially synthesized CDR3 mimic tag material particles. The universal sequence on the plasmid was used by digital PCR to quantify all system plasmids twice. The CV value of the two quantifications was <10%. The average value was taken and diluted to the same copy number. All IGH artificially synthesized CDR3 mimic tag material particles were then mixed for subsequent steps.

[0060] 2) IGH multiplex PCR amplification

[0061] The reaction system contained the same copy number of artificially synthesized CDR3 mimic tag plasmids. Each IGH primer working solution was mixed at a concentration of 100 nM. Using the Multiplex PCR Assay Kit Ver.2, the reaction system was prepared according to the table below. The prepared reagents were mixed by inversion and centrifugation, then placed in a PCR instrument and subjected to the following PCR amplification conditions: Reaction system:

[0062] Step 1 PCR amplification reaction conditions:

[0063] After the first round of PCR is completed, the second round of PCR to add adapters can be performed.

[0064] PCR running conditions:

[0065] Transfer the PCR reaction mixture to a 1.5 mL centrifuge tube and purify the amplified product 1.0-fold using the AMPure XP DNA Purification Kit (Thermo). The specific method is as follows: a) Remove Ampure XP Beads at 4℃ and allow them to stand at room temperature for 30 min to equilibrate; b) Shake well before use, add magnetic beads at a 1:1 ratio with the sample volume and mix well, then let stand for 5 min; c) Transfer the 1.5 mL centrifuge tube to a magnetic rack and let it stand for 3-5 minutes until clear; d) Keep the centrifuge tubes on the magnetic rack and carefully remove the supernatant, being careful not to touch the magnetic beads; e) Add 500 μL of 75% ethanol, gently blow the magnetic beads 2-3 times, wait 30 s, and discard the supernatant (when adding ethanol, add it slowly and try not to add the liquid towards the magnetic beads, otherwise the magnetic beads will detach from the tube and be damaged). f) Repeat step e to remove as much supernatant as possible (no need to blow the magnetic beads); g) Place the magnetic beads in a constant temperature mixer at 37℃ for 3-5 minutes to dry until there is no moisture on the surface of the magnetic beads. h) Add 31 μL of nuclease-free water to a 1.5 mL centrifuge tube, mix thoroughly, let stand for 5 min, and then place on a magnetic rack for about 5 min until clear; i) Transfer 30 μL of liquid into a new 1.5 mL centrifuge tube that has been prepared beforehand.

[0066] The concentration of purified DNA was determined using a Qbit BR analyzer, and the library concentration needed to be greater than 5 ng / μL.

[0067] 3) Off-machine data analysis

[0068] The data after NGS analysis were used to calculate the relative abundance and theoretical abundance of primers, evaluate the working efficiency of each primer, and then use this data to correct for the working efficiency differences between different primers. The specific method is as follows: Relative abundance formula:

[0069] Wherein, numerator: number of sequencing reads for primer pair (F_i and R_j); denominator: total number of reads for traversing F primers and R primer combinations, and m and n represent the number of F primers and the number of R primers, respectively.

[0070] In this embodiment, m and n are 15 and 6 respectively.

[0071] Theoretical abundance formula:

[0072] Wherein, N(total): the total number of plasmids used for primer efficiency evaluation.

[0073] Efficiency factor: …

[0074] …

[0075] Wherein, numerator: relative abundance of R_1 to R_n; denominator: theoretical abundance of R_1 to R_n.

[0076] Concentration adjustment:

[0077] Among them, Original conc(i) : Original primer concentration.

[0078] The results of Effector(F_i) and Effector(R_i) do not need to be adjusted if they are in the range of 0.8 to 1.2. If they are in the range of 0 to 0.8 or greater than 1.2, the concentrations of F_i and R_i should be adjusted to NewConc(i).

[0079] NGS sequencing can be performed using the Illumina or MGI platforms.

[0080] Table 3. IGH primers and their efficiency

[0081] Based on the analysis of the above formula and the adjustments made according to the proportions in Table 3, the above experimental procedure was repeated. As can be seen from the results in Table 4, even when the same template copy number exists, the bias before primer concentration adjustment deviates greatly from 1. When the bias is calculated again after primer concentration adjustment, the bias between different primers tends to 1. This indicates that the primer concentration adjusted using the experimental and calculation methods in this invention can greatly avoid the amplification differences caused by the amplification efficiency between primer pairs.

[0082] Table 4. Amplification deviation before and after IGH primer adjustment

[0083] exist Figure 2 The results shown in Figure A indicate that before primer concentration adjustment, the absolute number of primer-amplified reads tended towards two extremes (extremely high and extremely low), indicating a clear primer amplification bias; after primer concentration adjustment ( Figure 2 B) It can be found that the absolute difference in the number of amplified reads between different primers is significantly reduced, and in the bias results ( Figure 2As shown in C), before primer concentration adjustment (black column), different primers significantly deviated from "1," while after adjustment (gray column), different primers significantly tended towards "1." These results demonstrate that the computational analysis method for evaluating primer working efficiency in this system can greatly simplify the calculation method while addressing bias issues and avoiding unnecessary computational errors introduced by over-calculation. Furthermore, several experiments can be repeated for comprehensive calculation to avoid inaccuracies caused by normal errors in the experiment and the sequencer. Similar processes have been used in adjusting primer concentrations for IGDH, IGK (including KDE), and IGL, yielding similar results, but not all are presented due to space limitations. This invention is significantly superior to other similar solutions in addressing primer preference issues, providing clear experimental, computational, and analytical logic, and eliminating the need to use fragmented fragments as amplification templates. This greatly avoids the problem of varying starting amounts of different primer templates caused by the introduction of restriction endonucleases. Furthermore, it significantly reduces human-introduced errors in evaluating primer amplification efficiency. Moreover, this system and its computational logic can be parallelly ported to other multiplex PCR projects, solving a common problem from specific to general.

[0084] In this embodiment, the improved scheme directly mixes the quantified system in equal amounts before amplification when using the template. This can more realistically reflect the primer efficiency deviation when amplifying each template using different primer pairs, thus avoiding the problem of amplification efficiency deviation caused by differences in the amount of template used in the original technology.

[0085] This embodiment uses all IGH F and R primers, totaling 90 combinations. However, not all combinations are necessary in practice. It is essential to ensure that all F and R primers used in the primer pool participate in the combination to form different CDR3 simulated tag particles, facilitating subsequent comprehensive evaluation of the efficiency of each primer. The formulas used are described in the calculation process above.

[0086] Example 5

[0087] Based on the cloning analysis system in Example 3, the efficiency and correction of TRB multiplex primer PCR primers were evaluated.

[0088] 1) Quantitative

[0089] According to the method in Example 3, different primers were constructed into the same or one system to form artificially synthesized CDR3 mimic tag material particles. The universal sequence on the plasmid was used by digital PCR to quantify all system plasmids twice. The CV value of the two quantifications was <10%. The average value was taken and diluted to the same copy number. All TRB artificially synthesized CDR3 mimic tag material particles were then mixed for subsequent steps.

[0090] 2) TRB multiplex PCR amplification

[0091] The reaction system contained the same copy number of artificially synthesized CDR3 mimic tag plasmids. Each TRB primer working solution was mixed at a concentration of 100 nM. Using the Multiplex PCR Assay Kit Ver.2, the reaction system was prepared according to the table below. The first step of the multiplex PCR amplification reagents was prepared by inverting and mixing the prepared reagents, centrifuging, and then placing them in a PCR instrument under the following PCR amplification conditions: Reaction system:

[0092] Step 1 PCR amplification reaction conditions:

[0093] After the first round of PCR is completed, the second round of PCR to add adapters can be performed.

[0094] PCR running conditions:

[0095] Transfer the PCR reaction mixture to a 1.5 mL centrifuge tube and purify the amplified product 1.0-fold using the AMPure XP DNA Purification Kit (Thermo). The specific method is as follows: a) Remove Ampure XP Beads at 4℃ and allow them to stand at room temperature for 30 min to equilibrate; b) Shake well before use, add magnetic beads at a 1:1 ratio with the sample volume and mix well, then let stand for 5 min; c) Transfer the 1.5 mL centrifuge tube to a magnetic rack and let it stand for 3-5 minutes until clear; d) Keep the centrifuge tubes on the magnetic rack and carefully remove the supernatant, being careful not to touch the magnetic beads; e) Add 500 μL of 75% ethanol, gently blow the magnetic beads 2-3 times, wait 30 s, and discard the supernatant (when adding ethanol, add it slowly and try not to add the liquid towards the magnetic beads, otherwise the magnetic beads will detach from the tube and be damaged). f) Repeat step e to remove as much supernatant as possible (no need to blow the magnetic beads); g) Place the magnetic beads in a constant temperature mixer at 37℃ for 3-5 minutes to dry until there is no moisture on the surface of the magnetic beads. h) Add 31 μL of nuclease-free water to a 1.5 mL centrifuge tube, mix thoroughly, let stand for 5 min, and then place on a magnetic rack for about 5 min until clear; i) Transfer 30 μL of liquid into a new 1.5 mL centrifuge tube that has been prepared beforehand.

[0096] The concentration of purified DNA was determined using a Qbit BR analyzer, and the library concentration needed to be greater than 5 ng / μL.

[0097] 3) Off-machine data analysis

[0098] The data after NGS analysis were used to calculate the relative abundance and theoretical abundance of primers, evaluate the working efficiency of each primer, and then use this data to correct for the working efficiency differences between different primers. The specific method is as follows: Relative abundance formula:

[0099] Wherein, numerator: number of sequencing reads for primer pair (F_i and R_j); denominator: total number of reads for traversing F primers and R primer combinations, and m and n represent the number of F primers and the number of R primers, respectively.

[0100] In this embodiment, m and n are 35 and 18 respectively.

[0101] Theoretical abundance formula:

[0102] Wherein, N(total): the total number of plasmids used for primer efficiency evaluation.

[0103] Efficiency factor: …

[0104] …

[0105] Wherein, numerator: relative abundance of R_1 to R_n; denominator: theoretical abundance of R_1 to R_n.

[0106] Concentration adjustment:

[0107] Among them, Original conc(i) : Original primer concentration.

[0108] Effector(F_i) and Effector(R_i) results in the range of 0.75 to 1.25 do not require adjustment. When the values ​​are between 0 and 0.75 or greater than 1.25, the concentrations of F_i and R_i should be adjusted to New. Conc(i) .

[0109] NGS sequencing can be performed using the Illumina or MGI platforms.

[0110] Table 5. TRB primers and their efficiency

[0111]

[0112] Based on the analysis of the above formula and the adjustments made according to the proportions in Table 5, the above experimental procedure was repeated. As can be seen from the results in Table 6, even when the same template copy number exists, the bias before primer concentration adjustment deviates greatly from 1. When the bias is calculated again after primer concentration adjustment, the bias between different primers tends to 1, indicating that the primer concentration adjusted using the experimental and calculation methods in this invention can greatly avoid the amplification differences caused by the amplification efficiency between primer pairs.

[0113] Table 6. Amplification deviation before and after TRB primer adjustment

[0114]

[0115] exist Figure 3 The results shown in Figure A indicate that before primer concentration adjustment, the absolute number of primer-amplified reads tended towards two extremes (extremely high and extremely low), indicating a clear primer amplification bias; after primer concentration adjustment ( Figure 3 B) It can be found that the absolute difference in the number of amplified reads between different primers is significantly reduced, and in the bias results ( Figure 3 As shown in C), before primer concentration adjustment (black column), different primers significantly deviated from "1," while after adjustment (gray column), different primers significantly tended towards "1." These results demonstrate that the computational analysis method for primer efficiency evaluation in this system can greatly simplify the calculation method while resolving the preference problem, avoiding unnecessary computational errors introduced by over-calculation. Furthermore, several experiments can be repeated for comprehensive calculation to avoid inaccuracies caused by normal errors in the experiment and the sequencer. A similar process was also used in adjusting the primer concentrations of TRB+, TRD, and TRG, yielding similar results, but not all were shown due to space limitations. This invention is significantly superior to other similar schemes in resolving primer preference problems, with clear experimental, computational, and analytical logic. It eliminates the need to use fragmented fragments as amplification templates, greatly avoiding the problem of different starting amounts of templates caused by the introduction of restriction endonucleases. Simultaneously, it significantly reduces human-introduced errors in primer amplification efficiency evaluation. Moreover, this system and computational logic can be parallelly transferred to other multiplex PCR projects, solving common problems from specific points to a broader scope.

[0116] In this embodiment, the improved technology directly mixes the quantified system in equal amounts before amplification when using the template, which can more realistically reflect the primer efficiency deviation when amplifying each template using different primer pairs, thus avoiding the problem of amplification efficiency deviation caused by the difference in the amount of template used in the original technology.

[0117] This embodiment uses all F and R primers of TRB, totaling 630 combinations. However, not all combinations are necessary in practice. It is essential to ensure that all F and R primers used in the primer pool participate in the combination to form different CDR3 simulated tag particles, facilitating subsequent comprehensive evaluation of the efficiency of each primer. The formulas used are described in the calculation process above.

[0118] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A CDR3 analog tag, characterized in that, Its sequence, from the 5' end to the 3' end, includes: upstream primer sequence, V gene mimic sequence, CDR3 mimic sequence, J gene mimic sequence, and downstream primer sequence; the CDR3 mimic sequence contains a non-human tag sequence and a numbering sequence; the non-human tag sequence contains a system identification sequence; the CDR3 mimic tag sequence includes a TCR-encoding gene CDR3 mimic sequence or a BCR-encoding gene CDR3 mimic sequence.

2. The CDR3 analog tag according to claim 1, characterized in that, The simulated sequence of the TCR encoding gene CDR3 is CCINQB(b)J(j)DVW or CCINQB(b)J(j)DVF from the 5' end to the 3' end. And / or, the simulated sequence of the BCR encoding gene CDR3 from the 5' end to the 3' end is CCINQB(b)J(j)DVF or CCINQB(b)J(j)DVWGXG; in: C represents the nucleotide corresponding to the codon of cysteine, the CDR3 initiation conserved amino acid; W indicates the nucleotide corresponding to the codon of the conserved amino acid tryptophan, which is terminated by CDR3. F indicates the nucleotide corresponding to the codon of the conserved amino acid phenylalanine, which is terminated by CDR3. G indicates the nucleotide corresponding to the codon of glycine in the CDR3-terminated conserved amino acid motif; CINQDV represent the nucleotides corresponding to the codons of cysteine, isoleucine, asparagine, glutamine, aspartic acid, and valine, respectively. B and J can be independently selected from adenine, guanine, cytosine, or thymine; X is any amino acid; b and j are selected from natural numbers from 5 to 30.

3. The CDR3 analog tag according to claim 1, characterized in that, When the CDR3 simulated tag sequence is an IGH sequence, the system recognition sequence is selected from the nucleotide sequences shown in SEQ ID NO: 1-12; And / or, when the CDR3 simulated tag sequence is a TRB sequence, the system identification sequence is selected from the nucleotide sequences shown in SEQ ID NO: 13-34.

4. The CDR3 analog tag according to claim 1, characterized in that, The probability of the CDR3 simulated label sequence appearing in real samples is ≤0.0001%.

5. The CDR3 analog tag according to claim 1, characterized in that, The GC content of the immune repertoire simulated sequence corresponding to the CDR3 simulated tag sequence is 0.58±0.07, and the A, T, C, and G bases cannot appear consecutively for more than 3 times.

6. The CDR3 analog tag according to claim 1, characterized in that, Its sequence from the 5' end to the 3' end: the upstream primer contains, in sequence, a system backbone homology arm 1 and a restriction endonuclease site A upstream of the 5' end, and a restriction endonuclease B and a system backbone homology arm 2 downstream of the 3' end; And / or, the downstream primer contains, upstream of the 5' end, a system backbone homology arm 3 and a restriction endonuclease site C, and downstream of the 3' end, a restriction endonuclease D and a system backbone homology arm 4.

7. The use of the CDR3 mimic tag according to any one of claims 1-6 in evaluating the working efficiency of immune repertoire primers or in preparing immune repertoire detection products.

8. The application according to claim 7, characterized in that, The CDR3 mimic tag in the immune repertoire detection product exists in the form of circular or linear, single-stranded or double-stranded DNA.

9. A method for evaluating the working efficiency of primers for an immune repertoire, characterized in that, include: Construct various CDR3 mimic tag sequence plasmids, each containing different upstream and downstream primers; the CDR3 mimic tag is the CDR3 mimic tag described in any one of claims 1-6; Primer efficiency was evaluated based on amplification data from multiple primers. The evaluation process was as follows: Calculate relative abundance: the ratio of the number of sequencing reads for any primer pair to the total number of reads, where the total number of reads is the sum of the number of sequencing reads for all upstream and downstream primer combinations; Theoretical abundance is calculated as one-tenth of the total number of plasmids used for primer efficiency evaluation. Calculate the efficiency factors of upstream and downstream primers separately: the product of the ratio of the relative abundance of all upstream or downstream primers to the theoretical abundance is raised to the power of 1 / 2 as the corresponding efficiency factor. The ratio of the original primer concentration to the efficiency factor is calculated as the adjusted primer concentration.

10. The evaluation method according to claim 9, characterized in that, The total plasmid concentration of the CDR3 simulated tag was 10. 2 Up to 10 15 Copy / μL.

Citation Information

Patent Citations

  • Immune repertoire standard substance sequence as well as design method and application thereof

    CN111850016A