A method for correcting PCR bias and a method for quantitative analysis of immune repertoires.

By adjusting the concentration of multiplex PCR primer combinations and optimizing the PCR process, combined with next-generation sequencing technology, the PCR bias problem in the quantitative analysis of TCR-β immune repertoire was solved, achieving efficient and accurate quantitative analysis of TCR-β immune repertoire.

CN119673282BActive Publication Date: 2025-12-02SHANGHAI INNOSTAR BIO TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411741719.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-11-29
Publication Date
2025-12-02
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

The existing technology lacks an accurate method for quantitative analysis of TCR-β immune repertoire gDNA. PCR bias leads to inaccurate quantitative results that cannot be effectively corrected. Furthermore, the low binding efficiency of existing molecular tag sequence methods results in rare clonal loss.

Method used

This invention provides a method for PCR bias correction and a combination of artificially synthesized DNA sequences for TCR-β immune repertoires. By adjusting the concentration ratio of multiplex PCR primer combinations and optimizing primer concentrations, combined with next-generation sequencing technology, efficient and quantitative analysis of TCR-β immune repertoires can be achieved.

Benefits of technology

This method enables accurate quantitative analysis of TCR-β immune repertoires at the DNA level, corrects for biases in the PCR amplification process, improves the precision and accuracy of quantitative analysis, and avoids the binding efficiency problems associated with molecular tag sequence methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005162744520000031
    Figure BDA0005162744520000031
  • Figure BDA0005162744520000032
    Figure BDA0005162744520000032
  • Figure BDA0005162744520000033
    Figure BDA0005162744520000033
Patent Text Reader

Abstract

This invention discloses a PCR bias correction method, a synthetic DNA sequence combination for a TCR-β immune repertoire, a primer combination for amplifying the TCR-β immune repertoire, a quantitative analysis method for a TCR-β immune repertoire from gDNA samples, a detection kit, and their applications. The PCR bias correction method can effectively correct biases generated during PCR amplification. The synthetic DNA sequence combination can effectively adjust the primer concentration ratio of multiplex PCR primer combinations. The primer combination can efficiently, quantitatively, and with minimal interference with each other amplify the TCR-β immune repertoire. The method can accurately quantify the immune repertoire at the DNA level.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese patent application CN202311838286.6, filed on December 28, 2023. The entire contents of the aforementioned Chinese patent application are incorporated herein by reference. Technical Field

[0002] This invention belongs to the field of biological detection, specifically relating to a PCR bias correction method, a combination of artificially synthesized DNA sequences for a TCR-β immune library, a set of primer combinations for amplifying the TCR-β immune library, a quantitative analysis method for a TCR-β immune library from gDNA samples, a detection kit, and their applications. Background Technology

[0003] The diversity of T cell receptors (TCRs), B cell receptors (BCRs), and secreted antibodies forms the core of the complex immune system, playing a crucial role in protecting the body from viral, bacterial, and other pathogen invasions. TCRs are proteins on the surface of T cells responsible for specifically recognizing antigenic peptides that bind to the major histocompatibility complex (MHC). When a TCR binds to an antigenic peptide and MHC, T lymphocytes are activated through signal transduction, initiating an immune response. In humans, TCRs consist of a heterodimeric αβ chain (approximately 95%, TRA, TRB) or γδ chain (approximately 5%, TRD, TRG). Structurally, each chain can be divided into a constant region and a variable region. [1,2] CDR1 and CDR2 are relatively conserved and are responsible for recognizing MHC; CDR3 is the TCR region that directly contacts the antigen. CDR3 is encoded by a portion of V, all D and J, and the linker region between VD and DJ, therefore CDR3 has the highest degree of variability. Because the V (65-100 types), D (2 types), and J (13 types) gene fragments themselves are diverse, and furthermore, during rearrangement, the linker regions of VD and DJ often have random insertions or deletions of non-template nucleotides, the diversity of the TCR CDR3 region is further increased, theoretically resulting in 2 × 10-1 TCR regions. 18 TCR-αβ [3] Because lymphocytes are constantly generated, die, and proliferate when stimulated, the TCR lineage is dynamic, reflecting an individual's immune potential and history. Developing precise methods to study the complete individual immune repertoire is of great significance.

[0004] Numerous TCR immune repertoire sequencing and analysis methods have been developed. Based on the research material, these methods can be categorized into DNA-based (gDNA) and RNA-based methods. The former can be further divided into multiPCR using V and J primer sets or probe hybridization capture methods, while the latter can be divided into multiplex PCR or RACE-PCR methods (which can optionally incorporate UMI to limit PCR amplification bias and sequencing errors). Multiplex PCR is compatible with both gDNA and RNA materials. DNA has higher stability, and the TCR DNA template for a single cell is singular, facilitating the quantification of individual TCR clones. However, due to PCR bias, there is currently no mature quantification method. At the RNA level, each cell may have multiple copies of TCR or BCR transcripts. Therefore, using RNA as a template is more sensitive for identifying rare receptor expressions and provides more complete analysis of specific receptor variants. However, quantifying individual TCR clones is extremely challenging. [4] Target region capture methods require fewer PCR steps and have less PCR bias, but the TCR region itself is highly complex and variable, making primer design for hybridization capture challenging. RACE PCR methods typically preserve the complete TCRVDJ region, avoiding amplification bias. However, because it uses RNA as the starting material, it has relatively higher operational requirements compared to other techniques, and the entire process is complex, potentially affecting reproducibility.

[0005] In recent years, cell and gene therapy technologies have been widely applied in fields such as genetic diseases, immune system diseases, tumor treatment, and neurodegenerative diseases. Among them, tumor-infiltrating lymphocytes (TILs) are an important cell therapy approach. TILs are lymphocytes isolated from tumor tissue, possessing the ability to infiltrate tumor tissue and specifically recognize tumor cells. Typically, TIL cell therapy involves extracting lymphocytes from the patient's cancerous tissue, expanding them in large quantities in vitro, and then reinfusing them into the patient. Because this process does not involve genetic modification, the expanded cells are indistinguishable from the patient's own cells in terms of gene sequence and protein expression. After infusion, PCR and flow cytometry methods are almost completely unable to distinguish TILs from native cells, let alone detect the concentration, expansion, and distribution of the infused TILs. Although there is currently no pharmacokinetic (PK) detection method for TILs, changes in the T cell receptor repertoire (TCR repertoire) can serve as a similar PK detection method for TILs, and it is currently the only way to detect TILs. However, current TCR immune repertoire sequencing quantitative analysis methods mainly target RNA for quantitative analysis. Due to differences in gene expression levels, RNA sequencing results cannot be directly correlated with cell number. Each cell has only two TCR genome copies; therefore, for PK analysis of TILs, the gDNA method is a more suitable choice.

[0006] Currently, there are patents in the published literature regarding primer compositions for amplifying TCR-β (patent application publication numbers CN103205420A, CN108602874A, etc.), proposing primer combinations capable of amplifying the TCR-β CDR3 coding sequence. However, with the supplementation of TCR sequences through sequencing, primers need to be designed for novel TCR types, and these patents cannot quantify different types of TCR-β immune repertoires. In addition, there are related products for quantitative analysis of TCR immune repertoires. A patent (patent application publication number CN113122618A) introduces molecular tag sequences containing random bases to eliminate the problem of heterogeneous PCR amplification and achieve quantification of TCR clonal clusters in DNA samples. However, due to the low binding efficiency of molecular tag sequences, they are mainly used for RNA quantification analysis. Similarly, due to the binding efficiency issue, using molecular tag sequence methods to correct PCR bias can lead to the loss of rare clonal types. [5] .

[0007] In summary, quantitative analysis of immune repertoires, including TCR-β, is of significant research importance, and further improvements are needed to better quantify immune repertoires. Summary of the Invention

[0008] To address the technical problems of PCR bias and the lack of accurate quantitative analysis methods for TCR-β immune repertoire gDNA in the quantitative analysis of T cell receptor β chain immune repertoire, this invention provides a PCR bias correction method, a synthetic DNA sequence combination for TCR-β immune repertoire, a primer combination for amplifying TCR-β immune repertoire, a method for quantitative analysis of TCR-β immune repertoire from gDNA samples, a detection kit, and their applications. The PCR bias correction method can effectively correct bias generated during PCR amplification. The synthetic DNA sequence combination can effectively adjust the primer concentration ratio in multiplex PCR primer combinations. The primer combination can efficiently, quantitatively, and with minimal interference with each other amplify the TCR-β immune repertoire. The method can accurately quantify immune repertoires (e.g., TCR-β) at the DNA level.

[0009] To solve the above-mentioned technical problems, the present invention provides a technical solution as follows: a method for correcting PCR bias in an immune repertoire, wherein the PCR bias correction algorithm is used to correct the PCR reaction bias of the number of each unique CDR3 sequence in an unknown sample (unique CDR3 refers to sequences whose CDR3 regions have the same amino acid sequence and whose corresponding V and J genes are also the same) according to their corresponding VJ combinations. The correction method includes the following steps: 1) Analyzing the number of each unique CDR3 sequence (denoted as CDR3i) in an unknown sample containing cellular genomic DNA and denoting it as N. i,k (The number of sequences without bias correction is generally the direct data after sequencing, or the number of sequences after sequencing with only low-quality data removal, such as the number of sequences after removing primer dimers and multimers); 2) Obtain the V and J genes corresponding to each CDR3 i; 3) Based on the baseline bias correction coefficient b for each VJ combination. k The number of sequences for each CDR3i is corrected according to the following formula:

[0010]

[0011] i, k, and c are all integers greater than or equal to 1, where i represents the corresponding CDR3 i, and Correction N. i,k This indicates the number of corrected sequences for CDR3i, k represents the VJ combination corresponding to CDR3i (which can also be denoted as VJ combination k), and b k c represents the baseline bias correction coefficient corresponding to combination k in VJ, and b represents the number of cycles in multiplex PCR. k c-1 This represents the preference correction coefficient corresponding to combination k in VJ during multiplex PCR cycle c.

[0012] The b k The following steps were taken: 1) DNA was extracted from all M training samples and subjected to c... low 1) First multiplex PCR amplification with low cycle number; 2) c low After multiplex PCR amplification, all products were divided into two portions. One portion was designated as the low-cycle sample. The multiplex PCR amplification products from this low-cycle sample were then subjected to next-generation sequencing for library construction. The concentration of the constructed library was denoted as Con. low The other product was further subjected to multiplex PCR amplification and designated as the high-cycle sample. The total number of cycles for the two amplifications of the high-cycle sample was denoted as c. high The multiplex PCR amplification products of the high-cycling samples were subjected to next-generation sequencing for library construction. The concentration after library construction was denoted as Con. high3) Sequencing of high- and low-cycle samples, obtaining at least 6 million sequences from each sample; 4) Sampling of the sequencing results from each high- and low-cycle sample, ensuring that the proportion of sampled data is proportional to the Con... high / Con low Consistent; 5) Perform bioinformatics analysis on the sequencing results and obtain the N of the two amplification products. i,k 6) Obtain the baseline preference correction coefficient b for each VJ combination k according to the following formula. k :

[0013]

[0014] m is an integer greater than 1 and less than M, representing one of the M training samples (training samples containing cellular genomic DNA), b m,k b represents the preference correction coefficient for combination k of VJ in sample m. m,k It is obtained from the following formula:

[0015]

[0016] B m,k For c high -c low The total preference value for each round is obtained using the following formula:

[0017]

[0018] This represents the number of sequences of all CDR3i corresponding to VJ combination k in sample m in the low-cycle samples. B represents the number of sequences in the high-cycle samples corresponding to all CDR3 i of VJ combination k in sample m. m,k By analyzing vectors and B is obtained by performing a linear fit with a zero intercept. m,k Let be the correlation coefficient of the fitting function of VJ combination k in sample m.

[0019] In this invention, except for Correction N i,k Apart from the number of sequences corrected for bias, all other sequence numbers are the number of sequences obtained from next-generation sequencing that have not been corrected for bias (which may be processed by removing low-quality sequences, such as the number of sequences after removing primer dimers and multimers).

[0020] The correction method provided by this invention utilizes the baseline preference correction coefficient b for each VJ combination k obtained from M training samples. kThis invention can correct the number of next-generation sequencing sequences (NGS) of CDR3i in PCR-constructed immune repertoires, i.e., correct for bias in performing the same PCR reaction on unknown samples. The correction method described in this invention has broad applicability, not limited to any specific PCR primer, reaction system, or reaction process. It can correct for biases introduced by the entire PCR reaction and is widely applicable to all databases of sequence numbers obtained from next-generation sequencing.

[0021] In this invention, CDR refers to the complementarity determining regions of an antibody or TCR.

[0022] In this invention, the next-generation sequencing refers to a high-throughput sequencing method that can rapidly sequence base pairs in DNA or RNA samples, preferably using an Illumina sequencing instrument.

[0023] In this invention, the immune repertoire refers to the sum of all functionally diverse B lymphocytes and T lymphocytes in an individual's circulatory system at any specific point in time, that is, the sum of all rearranged TCR and BCR encoding genes (clones) in an individual, including TCR-α, TCR-β, TCR-δ, TCR-γ, IGH, and IGL.

[0024] Optionally, the sampling is performed by downsampling the sequencing sequences of the samples using seqtk(v1,4) to make the sequencing quantities of the samples comparable.

[0025] Preferably, M is a positive integer, such as a positive integer greater than 2, 3, 4, 5, 6, 7, 8, 9, or 10. The larger M is, the more b is obtained. k The more representative the sequence, the better; when M is small, higher sequencing depth can be used to compensate, resulting in more representative b. k value.

[0026] Preferably, the linear fitting is performed using least squares approximation.

[0027] Preferably, the 5' ends of the primers used in the multiplex PCR all contain the same adapter. Preferably, the adapter of the V region primer contains the sequence shown in SEQ ID NO: 1 (TCGTCGGCAGCGTCAGATGTGTATAAGAGAC AG), and the adapter of the J region primer contains the sequence shown in SEQ ID NO: 49 (GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG).

[0028] Preferably, the primer combination used in the multiplex PCR is the primer combination as described below in this invention.

[0029] Preferably, the primer combination concentration ratio in multiplex PCR is optimized using artificially synthesized TCR-β DNA sequences. More preferably, the artificially synthesized TCR-β DNA is a combination of artificially synthesized DNA sequences as described below in this invention.

[0030] Preferably, the immune repertoire is a TCR-β immune repertoire.

[0031] Optionally, the unknown samples and / or the training samples are obtained using conventional methods such as magnetic bead method or column method.

[0032] Preferably, the bioinformatics analysis includes: 1) removal of data adapters after sampling; 2) splicing of paired-end sequencing data read1 and read2; 3) TCR-βV / J gene annotation and CDR3 region identification; 4) removal of primer dimers and multimers; and 5) CDR3 i sequence number counting.

[0033] Optionally, the connector removal is performed using the trim_galore software, which has the advantage of removing some primer dimers or polymers.

[0034] Optionally, the splicing is performed using Flash software, which has the advantage of obtaining a full-length CDR3 region sequence.

[0035] Optionally, the TCR-βV / J gene annotation and CDR3 region identification were performed using igblast software.

[0036] Optionally, the CDR3 i sequence count is obtained through the following steps: 1) Using V and J region genes from the IMGT database as reference sequences, the sequence lengths are compared with the reference genome after splicing, while also taking into account the lengths of primer combinations used in multiplex PCR; 2) Spliced ​​sequences that are too short to match the reference genome are removed.

[0037] To address the bias caused by primer binding efficiency and other issues in multiplex PCR primer combinations, this invention provides a set of artificially synthesized DNA sequences for TCR-β immune repertoires. The artificially synthesized DNA sequence combination contains multiple nucleotide sequences, each of which contains the following components: (1) a partial sequence of a V gene, including the entire V region in the CDR3 region and the entire V gene region amplified by the multiple primer combination; (2) the complete sequence of a J gene, including the entire J region in the CDR3 region and the entire J gene region amplified by the multiple primer combination; (3) the complete sequence of a D gene; and (4) a tag sequence.

[0038] Preferably, the synthetic DNA sequence assembly comprises two parts, one of which is a partial sequence of the same V gene at its 5' end, denoted as V. high frequency The 3' end of one part represents the complete sequence of all J genes; the other part is a synthetic DNA sequence whose 5' end represents partial sequences of all V genes, and whose 3' end represents the complete sequence of the same J gene, denoted as J. high frequency Theoretically, any V gene and any J gene can serve as a V gene. high freq uency and J high frequency Preferably, the V high frequency and J high frequency These are V28 (NCBI Accession number: NC_000007.14:142720874-142721160) and J2-7 (NCBI Accession number: NC_000007.14:142797456-142797502), respectively. Preferably, V28 contains the nucleotide sequence shown in NCBI Accession number: NC_000007.14:142720874-142721160, and / or, J2-7 contains the nucleotide sequence shown in NCBI Accession number: NC_000007.14:142797456-142797502.

[0039] Preferably, the tag sequence is AAACCTGAGAAACCA (SEQ ID NO:125), located between the V gene and the D gene.

[0040] More preferably, the DNA artificially synthesized sequence combination comprises the following nucleotide sequences:

[0041]

[0042]

[0043]

[0044] Taking SEQ ID NO:63 as an example, the sequence “CTCCCTGATTCTGGAGTCCGCCAGCACCAACCAGACATCTATGTACCTCTGTGCCAGCAGTTTATG” (SEQ ID NO:124) is a partial sequence of a V gene, the sequence “AAACCTGAGAAACCAT” (SEQ ID NO:125) is a tag sequence, the sequence “GGGACAGGGGGC” is a complete sequence of a D gene, and the sequence “TGAACACTGAAGCTTTCTTTGGACAAGGCACCAGACTCACAGTTGTAG” (SEQ ID NO:126) is a complete sequence of a J gene.

[0045] To address the aforementioned technical problems, this invention provides a set of primer combinations for amplifying a TCR-β immune repertoire, wherein the primer combinations comprise the following nucleotide sequences:

[0046]

[0047]

[0048]

[0049] Preferably, the amount of each primer in the primer combination is as follows:

[0050] Primers Number of copies Primers Number of copies TRBV2 2 TRBV15 2 TRBV3-1 2 TRBV16 2 TRBV4-1 2 TRBV17 2 TRBV(4-2,4-3) 2 TRBV18 2 TRBV5-1 2 TRBV19 2 TRBV5-3 2 TRBV20-1 2 TRBV(5-4,5-5,5-6,5-7,5-8) 2 TRBV23-1 2 TRBV6-1 2 TRBV24-1 2 TRBV(6-2,6-3) 2 TRBV25-1 2 TRBV6-4 2 TRBV27 2 TRBV6-5 2 TRBV28 2 TRBV6-6 2 TRBV29-1 2 TRBV6-7 2 TRBV30 2 TRBV6-8 2 TRBV3-1*02 2 TRBV6-9 2 TRBV20-OR9-2 2 TRBV7-1 2 TRBV23 / OR9-2*02-2 2 TRBV7-2 2 TRBJ1-1 6.5 TRBV7-3 2 TRBJ1-2 6.5 TRBV7-4 2 TRBJ1-3 6.5 TRBV7-6 2 TRBJ1-4 6.5 TRBV7-7 2 TRBJ1-6 6.5 TRBV7-8 2 TRBJ2-1 6.5 TRBV7-9 2 TRBJ2-2 6.5 TRBV9 2 TRBJ2-3 6.5 TRBV10-1 2 TRBJ2-4 6.5 TRBV10-2 2 TRBJ2-5 8.8 TRBV10-3 2 TRBJ2-6 6.5 TRBV(11-1,11-3) 2 TRBJ2-7 6.5 TRBV11-2 2 TRBJ1-5*01-1 19.5 TRBV(12-3,12-4,12-5) 3 TRBJ2-2P*01-2 6.5 TRBV13 2 TRBV14 2 .

[0051] To address the aforementioned technical problems, this invention provides a quantitative analysis method for an immune repertoire. The quantitative analysis method includes performing next-generation sequencing on an immune repertoire library constructed based on multiplex PCR, and correcting the number of sequencing sequences in the next-generation sequencing using the correction method described in this invention.

[0052] Preferably, the immune repertoire is a T-cell receptor β-chain immune repertoire from a DNA sample.

[0053] Preferably, the primers used in the next-generation sequencing are primer combinations as described in this invention.

[0054] More preferably, the quantitative analysis method further includes using the Jurkat cell line to verify the results of the next-generation sequencing corrected by the correction method as described in this invention.

[0055] Furthermore, and more preferably, in the verification, genomic DNA extracts from the Jurkat cell line are compared with CD3... +Genomic DNA extracts from T cells were mixed in different mass ratios to obtain mixed samples. Each mixed sample was subjected to PCR amplification and next-generation sequencing as defined in the correction method described in this invention, and the Correction N of the Jurkat cell line CDR3i in each mixed sample was calculated according to the correction method. i,k The ratio of this value to the mixed concentration was used for verification.

[0056] To address the aforementioned technical problems, the present invention provides a detection kit comprising a combination of artificially synthesized DNA sequences as described in the present invention and / or a combination of primers as described in the present invention. Preferably, the detection kit further comprises reagents for next-generation sequencing.

[0057] To address the aforementioned technical problems, this invention provides an application of the DNA artificially synthesized sequence combination and / or primer combination as described herein in the quantitative analysis of a T cell receptor β-chain immune repertoire. Preferably, the application is for non-diagnostic or therapeutic purposes.

[0058] In this invention, the T cell receptor β chain and TCR-β can be used in combination.

[0059] Based on common knowledge in the field, the above-mentioned preferred conditions can be combined arbitrarily to obtain various preferred embodiments of the present invention.

[0060] The reagents and raw materials used in this invention are all commercially available.

[0061] The positive and progressive effects of this invention are as follows:

[0062] The PCR bias correction method provided by this invention can effectively correct biases generated during the PCR reaction, making the quantitative analysis of gDNA immune repertoires based on multiplex PCR library construction and next-generation sequencing more accurate. Furthermore, using a set of artificially synthesized DNA sequence combinations for TCR-β immune repertoires designed in this invention, the primer concentration for PCR can be effectively optimized, and the effectiveness of PCR primers can be verified, further correcting the bias caused by primer amplification efficiency in PCR, and helping to achieve more accurate quantitative analysis of TCR-β immune repertoires. The primer combinations for amplifying TCR-β immune repertoires designed in this invention can also improve the accuracy of TCR-β immune repertoire quantitative analysis. Attached Figure Description

[0063] Figure 1 The concentration of the product before primer optimization.

[0064] Figure 2 This represents the concentration of the product after primer optimization.

[0065] Figure 3 This is a distribution diagram of the library product fragments before and after purification in Example 2.

[0066] Figure 4 The result of optimization for removing primer dimers and multimers is shown in the figure. Detailed Implementation

[0067] The present invention is further illustrated below by way of embodiments, but the invention is not limited to the scope of the embodiments described herein. Experimental methods in the following embodiments that do not specify specific conditions were performed according to conventional methods and conditions, or as selected according to the product instructions.

[0068] The multiPCR enzyme reaction system used in this invention ( The Multiplex PCR Kit (#206143) was purchased from Qiager, and the PCR enzyme reaction system for library construction (KAPA HiFi HotStart ReadyMix(2x)#KK2601) was purchased from Roche. Other instruments, reagents and consumables are commercially available unless otherwise specified. Unless otherwise specified, the experimental methods involved are conventional methods.

[0069] Example 1: Design of Primers and Artificially Synthesized DNA Sequences

[0070] 1. Design and synthesis of multiplex PCR primers

[0071] In this embodiment, multiplex PCR primers were designed and synthesized for all functional genes and ORFs in regions V and J of the IMGT database. The forward primers TRBV3-1*02, TRBV20-OR9-2, and TRBV23 / OR9-2*02, and the reverse primers TRBJ1-5*01 and TRBJ2-2P*01 were designed based on the recently discovered V-3-1*02, V20-OR9-2, V23 / OR9-2*02, and J1-5*01 and J2-2P*01. In summary, a total of 48 forward primers for region V and 14 reverse primers for region J were developed, as shown in Table 1.

[0072] Table 1. Detailed information on 48 V-region forward primers and 14 J-region reverse primers.

[0073]

[0074]

[0075]

[0076] 2. Design and synthesis of artificial DNA sequences

[0077] In this embodiment, artificially synthesized VDJ DNA sequences were designed and synthesized for certain V and J combinations to adjust the primer concentration for multiplex PCR: 1) Based on the VJ combination frequency of the healthy cohort in the TCRdb database, the high-frequency V28 and J2-7 sequences were selected to bind all or part of the remaining V and J gene sequences, respectively; 2) A D gene sequence was randomly inserted in between; 3) A tag sequence (AAACCTGAGAAACCAT, SEQ ID NO: 124) was inserted between the V and D sequences to distinguish the artificially synthesized DNA sequences from the TCR-β sequences in the samples in subsequent sequencing results. Sixty-one artificially designed and synthesized VDJ DNA sequences are shown in Table 2, where the V / J sequences correspond to all the primers mentioned above.

[0078] Table 2. Detailed information on the VDJ DNA artificially synthesized sequence.

[0079]

[0080]

[0081]

[0082] Example 2: Optimization of multiplex PCR primer combinations and primer concentration ratios based on artificially synthesized DNA sequences.

[0083] Human peripheral blood was collected and stored in sterile blood collection tubes containing anticoagulant at 2-8°C (shelf life 5 days). Genomic DNA was extracted from the peripheral blood. Specifically, magnetic bead digestion based on proteinase K was used. Genomic DNA was extracted using the Blood Genomic DNA Kit (EC101).

[0084] 1. Optimization of multiplex PCR primer concentrations

[0085] In this embodiment, the artificially synthesized DNA sequence from Example 1 was used to adjust the primer concentration in order to correct for PCR bias caused by differences in primer amplification efficiency.

[0086] Using artificially synthesized DNA sequences as PCR templates, PCR amplification was performed using primers from the aforementioned primer library. Each artificially synthesized DNA sequence reaction system included: a multiplex PCR enzyme reaction system, an artificially synthesized DNA sequence combination, and a combination of amplification primers corresponding to the V and J regions of the artificially synthesized DNA sequence combination.

[0087] The concentration of the amplified product was determined using Qubit, and the primer amplification efficiency was positively correlated with the product concentration. The product concentration was denoted as Con. s The average product concentration was then obtained using the following formula, and denoted as .

[0088] Through experience and experimentation, Con s and By comparing the results, samples with abnormal amplification efficiency were identified, and a relevant formula was established to normalize primer amplification efficiency. If the product concentration differs from the average product concentration by more than 20%, the sample is considered to have abnormal amplification efficiency. Specific screening is performed using the following formula: if f(s) is less than 0, the sample amplification efficiency is too low, and the primer concentration needs to be increased until f(s) is greater than or equal to 0.

[0089]

[0090]

[0091] Prepare 10x primer mix according to the optimized primer concentration ratio, as shown in Table 3:

[0092] Table 3. Ratios of different primer concentrations

[0093]

[0094]

[0095] The comparison before and after primer concentration optimization is as follows: Figure 1 and Figure 2 As shown, the consistency of the product is significantly increased after optimization.

[0096] Example 3: Obtaining baseline preference correction coefficients by analyzing training samples

[0097] 1. High- and low-cycle multiplex PCR reaction and sequencing library construction

[0098] In this embodiment, the aforementioned primer library is used to capture and enrich the target fragment from the extracted DNA sample. The overall process consists of the following steps: 1) First-round multiplex PCR amplification; 2) Second-round multiplex PCR amplification; 3) Multiplex PCR product purification with magnetic beads; 4) Multiplex PCR product library construction; 5) Library purification with magnetic beads; 6) Library quality control. Details are as follows:

[0099] The specific PCR reaction system for multiplex PCR amplification is shown in Table 4. In this embodiment, QIAGEN Multi PCR Master Mix was selected. Multiplex PCR Kit (#206143) is used as a PCR ReactionMix.

[0100] Table 4 PCR reaction system

[0101] Components / tube volume PCR Reaction Mix 25μl 10x Primer Mix 4μl Enzyme-free nucleic acid water 21μl-x Template DNA (1 μg) x Total volume 50μl

[0102] The procedure for the first PCR reaction is shown in Table 5. The number of cycles for the first PCR reaction is denoted as c. low :

[0103] Table 5. Initial PCR reaction procedure

[0104]

[0105] The product from the initial PCR amplification was divided into two portions. One portion was used for multiplex PCR product purification via magnetic beads, skipping step 2. The other portion underwent a second multiplex PCR amplification. The total number of cycles for both amplifications was denoted as c. high The specific reaction procedure is shown in Table 6:

[0106] Table 6 Multiplex PCR reaction procedure

[0107]

[0108] In this embodiment, the purification of multiplex PCR products using magnetic beads is the same as the purification of the library using magnetic beads, both employing conventional methods in the art: DNA sorting magnetic beads are used for one-sided DNA purification to remove small fragments of impurities within 100 bp, such as primers and primer dimers. Each purification process must ensure that both sample volumes are identical.

[0109] After initial purification, all multiplex PCR amplification products were used for next-generation sequencing library construction using sequencing adapter primers complementary to the universal adapter (i.e., Nextera XT (IDT)). In this embodiment, KAPA HiFi HotStart ReadyMix (KAPA HiFi HotStart ReadyMix(2x)#KK2601) was selected as the PCR reaction mix, and Nextera XT (IDT) primers were used as library construction primers. The specific library construction reaction system is shown in Table 7 below:

[0110] Table 7 Library Construction Reaction System

[0111] Components volume DNA 5μl Library Construction Primer 5 5μl Library Construction Priemr 7 5μl PCR Reaction Mix 25μl ddWater 10μl Total volume 50μl

[0112] The library construction reaction procedure is shown in Table 8 below:

[0113] Table 8 Library Construction Reaction Procedure

[0114]

[0115] After library construction, the concentration of the product library after only one round of multiPCR is denoted as Con. lowThe concentration of the product library after two rounds of multiPCR is denoted as Con. high .

[0116] In this embodiment, quality control was performed on the purified product. Specifically, quantification was performed using Qubit and fragment distribution analysis was performed using Qsep100. This invention recommends that the concentration of the library product be no less than 3 ng / μL, the total volume no less than 15 μL, the target fragment peaks of the library product be dispersed, the main peak be concentrated around 250-300 bp, and there be virtually no distribution within 100 bp. The fragment distribution of the library product before and after purification is shown below. Figure 3 As shown (the top image is before purification, and the bottom image is after purification), if there is still a significant distribution within 100bp in the purified library product, a magnetic bead purification step is added to remove fragments within 100bp.

[0117] 2. Sequencing

[0118] The DNA fragments from the sequencing library were loaded onto a sequencing chip, and sequencing was performed simultaneously using a next-generation sequencer. In this embodiment, the Illumina Nova-seq sequencer was selected, and the PE150 paired-end sequencing strategy was chosen. To ensure that at least 6 million sequences were obtained from both high- and low-cycle samples, the data volume per sample was no less than 3GB. Considering the special sample distribution of multiplex PCR, it is recommended that the data volume per sample be no less than 8GB.

[0119] 3. Data Analysis

[0120] The sequencing results of multiplex PCR high and low cycle samples were analyzed according to the following data analysis steps:

[0121] Step A: Sampling is performed on the sequencing results of high and low cycle samples respectively, so that the data volume ratio is maintained with Con. high / Con low Consistent, sampling was performed using the software seqtk(v1,4); where Con high Con low Obtained through the steps described above.

[0122] Step B: Use trim_galore (Version 0.6.4) to assemble the sequencing reads. [6] The software removes low-quality bases and linkers with the following parameters: (-q 25--phred33--length 36--stringency 3--paired).

[0123] Step C, using FLASH (Fast Length Adjustment of Short Reads) [7]The software splices read1 and read2 after removing the connectors to obtain the full-length CDR3 region sequence.

[0124] Step D, using IgBLAST (1.13.0) [8] The software performed V / J gene annotation and CDR3 region identification for TCR-β. Then, using V and J region genes from the IMGT database as reference sequences, and referring to the primer lengths in the aforementioned multiplex PCR primer combinations, the spliced ​​sequences were aligned with the reference genome sequence length. Spliced ​​sequences with excessively short alignment lengths were removed to eliminate primer dimers and multimers. The results of primer dimer and multimer removal optimization are shown below. Figure 4 As shown, Figure 4 The upper bar chart represents the sequence length distribution without removing dimers and multimers. Figure 4 The lower half of the bar chart shows the sequence length distribution after removing dimers and multimers.

[0125] 4. Calculation of baseline preference correction coefficient

[0126] For each CDR3 i (i.e., a unique CDR3, meaning that the amino acid sequences of these sequences in the CDR3 region are identical, and their corresponding V and J genes are also identical), PCR bias correction is performed according to its corresponding VJ combination. The specific steps include: 1) Analyzing the number of each CDR3 i in the training sample and recording it as N. i,k 2) Obtain the V and J genes corresponding to each CDR3 i. 3) Obtain the baseline bias correction coefficient of the VJ combination k based on the number of high and low cycle sequences of each CDR3 i in the training sample, as shown in the following formula.

[0127]

[0128] m is one of the M DNA training samples, k represents the VJ combination corresponding to the CDR3 i, and b m,k b represents the preference correction coefficient for combination k of VJ in sample m. m,k It is obtained from the following formula:

[0129]

[0130] B m,k For c high -c low The overall preference coefficient for the round cycle is obtained using the following formula:

[0131]

[0132] This represents the number of all CDR3i sequences with VJ combination k in the low-cycle sample m. B represents the number of high-cycle sample sequences corresponding to all CDR3 i with VJ combination k in sample m. m,k By analyzing vectors and B is obtained by performing a linear fit with an intercept of 0. m,k Let be the correlation coefficient of the fitting function for the VJ combination k in sample m;

[0133] The fitting function parameters for each VJ combination k in the training samples are shown in Table 9 below. In the table, k represents different VJ combinations, B k Let be the slope of the fitting function for VJ combination k in this sample, r-square be the R-squared value of the linear fitting function, and p-value be the p-value of the correlation coefficient of the fitting curve. The results selected VJ combinations k with a CDR3i greater than 10, r-square greater than 0.9, and p-value less than 0.01.

[0134] Table 9 shows the fitting function parameters for each VJ combination k in the training samples.

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143] Finally, the baseline preference correction coefficients for each VJ combination k are shown in Table 10 below. In the table, NA indicates that no corresponding VJ combination was detected in this embodiment, so the preference coefficient for that VJ combination cannot be provided.

[0144] Table 10 Baseline bias correction coefficients for each VJ combination k

[0145]

[0146]

[0147] Example 4: Correcting PCR bias based on baseline bias correction coefficient and quantitatively analyzing T cell lines

[0148] 1. DNA extraction from Jurkat cell line

[0149] In this embodiment, a magnetic bead method based on proteinase K digestion is used. Genomic DNA was extracted from the Jurkat (Precella, CL-0129) cell line using the Blood Genomic DNA Kit #EC101. Specifically, the Jurkat cell line is a human T-lymphoblastic leukemia cell line (Clone E6-1). The V and J genes of this Jurkat cell line are TRBV12-3 (NCBI Accession number: NC_000007.14:142560642-142560931) and TRBJ1-2 (NCBI Accession number: NC_000007.14:142787017-142787064), respectively.

[0150] 2. Human peripheral blood PBMC CD3 + T-cell sorting and DNA extraction

[0151] In this embodiment, a flow cytometer (BD FACSAria) is used. TM Fusion (a method for separating CD3 from human peripheral blood PBMCs) + T cells, sorted CD3 + T cells, using a proteinase K-based digestion method with magnetic beads ( Genomic DNA was extracted using the BloodGenomic DNA Kit#EC101.

[0152] 3. In CD3 + Different concentrations of Jurkat cell line genomic DNA were added to the DNA of T cells.

[0153] 100ng CD3 + Different mass ratios (10%, 1%, 0.1%, 0.01%) of Jurkat T cell DNA were incorporated into T cell DNA samples to obtain various mixed samples.

[0154] 4. Multiplex PCR reaction and sequencing library construction

[0155] The above primer library was used to capture and enrich the target fragments incorporated into the above mixed samples. The specific experimental steps and reaction procedures were the same as in Example 3.

[0156] In this embodiment, the purification steps for multiplex PCR products using magnetic beads and the purification steps for libraries using magnetic beads are the same as in Example 3.

[0157] After initial purification, all multiplex PCR amplification products were used for next-generation sequencing library construction using sequencing adapter primers complementary to the universal adapter portion. The reagents and specific library construction reaction systems used were consistent with those in Example 3.

[0158] 5. Sequencing

[0159] The DNA fragments from the sequencing library were loaded onto a sequencing chip, and sequencing was performed simultaneously using a next-generation sequencer. In this embodiment, the Illumina Nova-seq next-generation sequencer was selected, and the PE150 paired-end sequencing strategy was chosen. At least 6 million sequences were obtained for each sample, and the data volume of a single sample was no less than 3GB.

[0160] 6. Data Analysis

[0161] The sequencing results of multiplex PCR high and low cycle samples were analyzed according to the following data analysis steps:

[0162] Step A: Use trim_galore (Version 0.6.4)[6] software to remove low-quality bases and adapters from the sequencing reads. The parameters are as follows (-q 25--phred 33--length 36--stringency 3--paired).

[0163] Step B, using FLASH (Fast Length Adjustment of Short Reads) [7] The software splices read1 and read2 after removing the connectors to obtain the full-length CDR3 region sequence.

[0164] Step C, using IgBLAST (1.13.0) [8] The software performs V / J gene annotation and CDR3 region identification for TCR-β. Then, the V and J region genes in the IMGT database are used as reference sequences. At the same time, the primer lengths in the above multiplex PCR primer combinations are referenced, and the sequence lengths of the spliced ​​sequences are compared with the reference genome. Spliced ​​sequences with too short alignment lengths are removed to eliminate primer dimers and multimers.

[0165] 7. Calculate the percentage of Jurkat cell lines in each pooled sample.

[0166] 1) Analyze the number of CDR3i sequences in each pooled sample and denote it as N. i,k(The number of sequences without bias correction is generally the direct data after sequencing, or the number of sequences after sequencing with only low-quality data removal, such as the number of sequences after removing primer dimers and multimers); 2) Obtain the V and J genes corresponding to each CDR3 i; 3) Based on the baseline bias correction coefficient b for each VJ combination. k The number of sequences for each CDR3i is corrected according to the following formula:

[0167]

[0168] i, k, and c are all integers greater than or equal to 1, where i represents the corresponding CDR3 i, and Correction N. i,k This indicates the number of corrected sequences for CDR3i, k represents the VJ combination corresponding to CDR3i (which can also be denoted as VJ combination k), and b k c represents the baseline bias correction coefficient corresponding to combination k in VJ, and b represents the number of cycles in multiplex PCR. k c-1 This represents the preference correction coefficient corresponding to combination k in VJ during multiplex PCR cycle c.

[0169] The proportion of Jurkat cell lines in each sample was calculated based on the corrected CDR3i sequence number. The results are shown in Table 11:

[0170]

[0171]

[0172] In the table, the proportion of reads aligned to Jurkat in the sequencing results of mixed samples with a Jurkat T cell DNA mass ratio of 10% was 7.1%-8.4%; the proportion was 1.3%-1.7%; the proportion was 0.1%-0.2%; and the proportion was 0.02%-0.04%. The experimental data show a good correlation between the detected values ​​and the actual incorporation ratios (CV < 20%, RE < 50%) at different incorporation ratios of 10%, 1%, 0.1%, and 0.01%. Even at the lowest concentration of 0.01%, this method can still detect the target cells.

[0173] The above validations demonstrate that when the CDR3 frequency is greater than 0.1%, the method established in this study can accurately detect the content of Jurkat cells in T cells at different proportions (CV < 20%, RE < 50%), and when the CDR3 frequency is greater than 0.01%, the method can detect CDR3. The experiments showed good parallelism and the results were stable and reliable, indicating that the method has good accuracy and sensitivity and can be used for quantitative analysis of small amounts of specific cells in a sample.

[0174] References

[0175] [1] Lefranc MP, Lefranc G, 2001a The Immunoglobulin Factsbook. Academic Press: London; 448 pages.

[0176] [2] Lefranc MP, Lefranc G, 2001b The T cell receptor Factsbook. Academic Press: London; 384 pages.

[0177] [3] Robins HS, Campregher PV, Srivastava SK, et al. Comprehensiveassessment of T-cell receptor beta-chain diversity in alphabeta T cells. [J]. Blood, 2009, 114(19): 4099-4107. DOI: 10.1182 / blood-2009-04-217604.

[0178] [4] Rosati, E., Dowds, CM, Liaskou, E. et al. Overview of methodologies for T-cell receptor repertoire analysis. BMC Biotechnol 17, 61 (2017). https: / / doi.org / 10.1186 / s12896-017-0379-9

[0179] [5]Barennes,P.,Quiniou,V.,Shugay,M.et al.Benchmarking of T cellreceptor repertoire profiling methods reveals large systematic biases.NatBiotechnol 39,236–245(2021).https: / / doi.org / 10.1038 / s41587-020-0656-3

[0180] [6]https: / / www.bioinformatics.babraham.ac.uk / projects / trim_galore /

[0181] [7]FLASH:Fast length adjustment of short reads to improve genomeassemblies.T.Magoc and S.Salzberg.Bioinformatics 27:21(2011),2957-63.

[0182] [8]Ye J,Ma N,Madden TL,Ostell JM.IgBLAST:an immunoglobulin variabledomain sequence analysis tool.Nucleic Acids Res.2013 Jul;41(Web Serverissue):W34-40.doi:10.1093 / nar / gkt382.Epub 2013 May 13.PMID:23671333;PMCID:PMC3692102。

Claims

1. A method for correcting PCR bias in an immune repertoire, characterized in that, The correction method includes the following steps: 1) Analyzing the number of unique CDR3 sequences in an unknown sample containing cellular genomic DNA and denoting it as N. i,k The unique CDR3 is denoted as CDR3i; 2) Obtain the V gene and J gene corresponding to each CDR3i; 3) Adjust the baseline preference coefficient b for each VJ combination. k The number of sequences for each CDR3i is corrected according to the following formula: i, k, and c are all integers greater than or equal to 1, where i represents the corresponding CDR3 i, and Correction N. i,k This indicates the number of corrected sequences for CDR3i, where k represents the VJ combination k corresponding to CDR3i, and b... k c represents the baseline bias correction coefficient corresponding to the VJ combination k, b represents the number of cycles in multiplex PCR, and c represents the number of cycles in multiplex PCR. k c-1 This represents the preference correction coefficient corresponding to combination k in VJ during multiplex PCR cycle c. The b k The following steps were taken: 1) DNA was extracted from all M training samples and subjected to c... low 1) First multiplex PCR amplification with low cycle number; 2) c low After multiplex PCR amplification, all products were divided into two portions. One portion was designated as the low-cycle sample. The multiplex PCR amplification products from this low-cycle sample were then subjected to next-generation sequencing for library construction. The concentration of the constructed library was denoted as Con. low The other product was further subjected to multiplex PCR amplification and designated as the high-cycle sample. The total number of cycles for the two amplifications of the high-cycle sample was denoted as c. high The multiplex PCR amplification products of the high-cycling samples were subjected to next-generation sequencing for library construction. The concentration after library construction was denoted as Con. high 3) Sequencing of high- and low-cycle samples, obtaining at least 6 million sequences from each sample; 4) Sampling of the sequencing results from each high- and low-cycle sample, ensuring that the proportion of sampled data is proportional to the Con... high / Con low Consistent; 5) Perform bioinformatics analysis on the sequencing results and obtain the N values ​​of the two amplification products. i,k 6) Obtain the baseline preference correction coefficient b for each VJ combination k according to the following formula. k : m is an integer greater than 1 and less than M, representing one of the M training samples, b m,k b represents the preference correction coefficient for combination k of VJ in sample m. m,k It is obtained from the following formula: B m,k For c high -c low The total preference value for each round is obtained using the following formula: This represents the number of sequences of all CDR3i corresponding to VJ combination k in sample m in the low-cycle samples. B represents the number of sequences in the high-cycle samples corresponding to all CDR3 i of VJ combination k in sample m. m,k By analyzing vectors and B is obtained by performing a linear fit with a zero intercept. m,k Let be the correlation coefficient of the fitting function of VJ combination k in sample m.

2. The correction method as described in claim 1, characterized in that, The sampling method involves downsampling the sequencing sequences of the samples using seqtk(v1,4) to ensure comparability of the sequencing quantities of the samples. And / or, the linear fitting is performed using least squares approximation; And / or, the primers used in the multiplex PCR all contain the same adapter at their 5' ends; And / or, the concentration ratio of primer combinations in multiplex PCR is optimized using a combination of artificially synthesized TCR-β DNA sequences; And / or, the immune repertoire is a TCR-β immune repertoire; And / or, the unknown samples and / or the training samples are obtained by magnetic bead method or column method.

3. The correction method as described in claim 1 or 2, characterized in that, The bioinformatics analysis includes: 1) removal of data adapters after sampling; 2) splicing of paired-end sequencing data read1 and read2; 3) TCR-βV / J gene annotation and CDR3 region identification; 4) removal of primer dimers and multimers; and 5) CDR3 i sequence number counting.

4. The correction method as described in claim 3, characterized in that, The adapter removal is performed using trim_galore software; and / or, the splicing is performed using flash software; and / or, the TCR-βV / J gene annotation and CDR3 region identification are performed using igblast software; and / or, the CDR3 i sequence count is obtained through the following steps: 1) Using V and J region genes from the IMGT database as reference sequences, the splicing is performed and the sequence length is compared with the reference genome, while also referring to the length of the primer combinations used in multiplex PCR; 2) Spliced ​​sequences that are too short to align with the reference genome are removed.

5. The correction method as described in claim 2, characterized in that, The TCR-βDNA artificially synthesized sequence combination contains multiple nucleotide sequences, each of which contains the following components: (1) a partial sequence of a V gene, including the entire V region in the CDR3 region and the entire V gene region amplified by the multiple primer combination; (2) the complete sequence of a J gene, including the entire J region in the CDR3 region and the entire J gene region amplified by the multiple primer combination; (3) the complete sequence of a D gene; and (4) a tag sequence.

6. The correction method as described in claim 5, characterized in that, The TCR-β DNA artificially synthesized sequence assembly comprises two parts, one of which is a partial sequence of the same V gene at its 5' end, denoted as V. high frequency The 3' end of one part represents the complete sequence of all J genes; the other part is a synthetic DNA sequence whose 5' end represents partial sequences of all V genes, and whose 3' end represents the complete sequence of the same J gene, denoted as J. high frequency .

7. The correction method as described in claim 6, characterized in that, The V high frequency and J high frequency They are V28 and J2-7, respectively.

8. The correction method as described in claim 5, characterized in that, The tag sequence is an amino acid sequence as shown in SEQ ID NO:125, located between the V gene and the D gene.

9. The correction method as described in claim 6, characterized in that, The TCR-βDNA artificially synthesized sequence combination includes nucleotide sequences as shown in SEQ ID NO:63~123.

10. The correction method as described in claim 2, characterized in that, The primer combination includes the following nucleotide sequences as shown in SEQ ID NO:1 to 62.

11. The correction method as described in claim 10, characterized in that, The amounts of each primer in the primer combination are as follows: 。 12. A method for quantitative analysis of an immune repertoire, characterized in that, The quantitative analysis method includes performing next-generation sequencing on an immune repertoire library constructed based on multiplex PCR, and correcting the number of CDR3 sequencing sequences in the next-generation sequencing using the correction method described in any one of claims 1-11.

13. The quantitative analysis method as described in claim 12, characterized in that, The immune repertoire is a T-cell receptor β-chain immune repertoire from DNA samples.

14. The quantitative analysis method as described in claim 12, characterized in that, The primers used in the second-generation sequencing are primer combinations that include nucleotide sequences as shown in SEQ ID NO:1 to 62.

15. The quantitative analysis method as described in claim 14, characterized in that, The quantitative analysis method further includes using the Jurkat cell line to verify the results of the next-generation sequencing corrected by the correction method as described in any one of claims 1-11.

Citation Information

Patent Citations

  • Primer compositions for the amplification of T cell receptor beta chain CDR3 coding sequences and uses of the primer compositions

    CN103205420A

  • TCR libraries

    CN108602874A

  • Method for accurately detecting T cell immune repertoire based on high-throughput sequencing and primer system thereof

    CN113122618A

  • Coding PCR second-generation sequencing library establishing method, kit and detection method

    CN108504649A

  • Method for constructing immune repertoire high-throughput sequencing library for removing chimera sequence in sample, group of primers and kit

    CN113999891A