A method for quantifying TCR β based on high-throughput sequencing

By using external reference cells and optimized multiplex PCR primers and molecular barcode template sequences, the problems of amplification bias and sequencing errors in high-throughput sequencing were solved, and accurate quantification of TCR β and accurate TCR library distribution were achieved.

CN114561454BActive Publication Date: 2025-11-14ARMY MEDICAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210222259.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-03
Publication Date
2025-11-14
Estimated Expiration
2042-03-03

AI Technical Summary

Technical Problem

Existing high-throughput sequencing technologies suffer from amplification bias and sequencing errors in T-cell receptor library sequencing, leading to biases in the estimation of T-cell receptor population characteristics and affecting the accuracy of TCR diversity.

Method used

Using external reference cells (such as 2B4 hybridoma cells) and optimized multiplex PCR primers, combined with molecular barcode template sequences and high-throughput sequencing platforms, accurate quantification of TCR β was achieved through amplification bias correction and sequencing error correction.

Benefits of technology

This method effectively estimates the total number of cells in a sample, reduces interference from multiplex PCR primers, corrects and standardizes TCR high-throughput sequencing data, and obtains accurate TCR β library distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114561454B_ABST
    Figure CN114561454B_ABST
Patent Text Reader

Abstract

This invention discloses a method for quantifying TCRβ based on high-throughput sequencing, comprising the following steps: lysing samples using Trizol and adding external reference cell lysis buffer to the lysed samples; extracting total RNA from the samples and external reference cells; performing reverse transcription using C-terminal specific primers for TCRβ; adding a template to the reverse transcription product; constructing a high-throughput sequencing library of TCRβ using a set of multiplex PCR primers with optimized sequence composition and concentration, and performing sequencing; and quantifying TCRβ in the samples using two rounds of external reference data. This invention, by adding external reference cells, can effectively estimate the total number of T cells in the sample; by using multiplex PCR primers with optimized sequence composition and concentration, it can better reduce interference and amplification bias between multiplex primers; by adding a template, it can help correct and standardize TCR high-throughput sequencing data, finally obtaining an accurate and realistic TCRβ library distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biotechnology, and in particular to a method for quantifying TCR β based on high-throughput sequencing. Background Technology

[0002] T cell receptors (TCRs) are specific receptors on the surface of T cells, responsible for recognizing antigens presented by the major histocompatibility complex (MHC) and mediating immune responses. The vast majority of T cells are α,β-T cells (accounting for approximately 90%–95% of all T cells), composed of highly diverse α and β subunits, both derived from germline genes rearranged in the thymus. Mature T cells migrate to the periphery, forming an extremely diverse T cell receptor repertoire that determines the body's adaptive mechanisms and capabilities to environmental changes. The diversity of TCRs originates from the recombination of V(D)J gene segments. During recombination, random insertions or deletions of nucleotides (non-template) at the junctions of VJ, VD, and DJ frequently occur, greatly enhancing TCR diversity. The complementarities determining region 3 (CDR3) on the T cell receptor B subunit is a crucial region of the TCR. This region exhibits the strongest binding affinity for antigenic peptides and is also the most diverse, best representing the diversity of TCRs. Therefore, researchers often study the diversity of the T cell receptor β chain CDR3 (TCR β CDR3) to investigate the diversity of the T cell receptor repertoire. With the development of high-throughput sequencing technology, researchers have developed T cell receptor sequencing technology (TCR-seq), which can sequence and analyze all TCRs in a sample to obtain the genetic information of all T cell receptors, comprehensively revealing the complexity and diversity of the T cell receptor repertoire. However, due to amplification bias and sequencing errors in T cell receptor repertoire sequencing data, previous studies have generally had biases in estimating the population characteristics of T cell receptors. Moreover, we have found that sequencing errors have a significant impact on the estimation of T cell repertoire diversity during deep sequencing of T cell receptor repertoires. Summary of the Invention

[0003] In view of the shortcomings of the prior art described above, the purpose of this invention is to provide a method for quantifying TCR β based on high-throughput sequencing. This method can quantify TCR β in samples and is applicable to research in the biomedical field involving immunology.

[0004] To achieve the above and other related objectives, this invention provides a method for quantifying TCR β based on high-throughput sequencing, comprising the following steps:

[0005] (1) Use Trizol to lyse the sample and add a fixed number of external reference cell lysis buffer to the lysed sample;

[0006] (2) Extract total RNA from samples and external reference cells;

[0007] (3) Reverse transcription was performed using a C-terminal specific primer for TCR β;

[0008] (4) Add a fixed number of templates to the reverse transcription product;

[0009] (5) Construct a high-throughput sequencing library of TCR β using a set of multiplex PCR primers with optimized sequence composition and concentration; (or: perform multiplex PCR using the sequence shown in SEQ ID NO.3-25 as multiplex PCR primers with optimized concentration, and recover the PCR products by gel extraction)

[0010] (6) Sequencing was performed using a high-throughput sequencing platform;

[0011] (7) Quantify TCR β in the sample using two rounds of external parameter data.

[0012] Furthermore, in step (1), the external reference cells are cells whose TCR sequences are different from those in the sample; preferably, the external reference cells are 2B4 hybridoma cells, but not limited to 2B4 hybridoma cells. As long as their TCR sequences are different from those in the sample, they can be used as external reference cells; in this embodiment of the invention, the number of 2B4 hybridoma cells used is 200, and the specific number of external reference cells can be adjusted according to the number of T cells in the sample.

[0013] Furthermore, in step (2), the total RNA extraction method is the Trizol method, but is not limited to the Trizol method.

[0014] Furthermore, in step (3), the C-terminal specific primer of TCR β is TRBC, the sequence of which is shown in SEQ ID NO.1, but is not limited to the primer in SEQ ID NO.1. Those skilled in the art can design and use it according to actual needs.

[0015] Furthermore, in step (3), the reverse transcription using the C-terminal specific primer of TCR β is performed as follows:

[0016] ① Take 0.1ug of RNA from step (2), 1ul of primer TRBC (10uM), and the remainder of water to prepare a 12ul reaction system. Then incubate the system in a PCR instrument at 72℃ for 3min and immediately place it on ice for 5min.

[0017] ② Prepare a 20ul reaction system by mixing the product obtained in step ①, 4ul 5X first strand buffer, 2ul dNTPs, 1ul RNase inhibitor, and 1ul RevertAid reverse transcriptase. Then incubate the mixture in a PCR instrument at 42℃ for 60min and 70℃ for 10min.

[0018] Furthermore, in step (4), the template consists of 23 sequences, as shown in SEQ ID NO.26-48. The template sequence of this invention comprises the V gene, three 6-bit molecular barcodes (BCs), the D gene, the J gene, and the C gene, reflecting the sequence characteristics of TCR β. Specifically, the V and C genes contain sites for amplification primer binding. Since there are only 23 functional V genes, this invention designed and synthesized 23 template sequences using different V genes, each sequence being 366 bp in length. The length of the molecular barcodes is not limited to 6; those skilled in the art can adjust them according to actual needs.

[0019] Furthermore, in step (5), the primer sequences for the multiplex PCR are shown in SEQ ID NO.3-25, and the SEQ ID NO.3-25 sequence contains a high-throughput sequencing adapter;

[0020] The sequence of the reverse primer is shown in SEQ ID NO.2:

[0021] CCATCTCATCCCTGCGTGTCTCCGACTCAG <barcode>AGACCTTGGGTGGAGTCAC.

[0022] In the sequences SEQ ID NO.2-25, adapter sequence 1: CCATCTCATCCCTGCGTGTCTCCGACTCAG (SEQ ID NO.49) and adapter sequence 2: CCTTCTATGGGCAGTCGGTGAT (SEQ ID NO.50) are adapter sequences for the Ion PGM platform. Different adapter sequences can be selected according to different high-throughput sequencing platforms.

[0023] Furthermore, in step (5), the multiplex PCR reaction system consists of 50 μL and includes the following reaction components: 25 μL mPCR premix, 5 μL forward primer (FW-primer mix), 5 μL reverse primer (RW-primer), 1 μL template mix, 5 μL cDNA, and 9 μL water.

[0024] The forward primer (FW-primer mix) consists of the sequence shown in SEQ ID NO.3-25, and the ratio of the sequence shown in SEQ ID NO.3-25 is 1∶2∶6∶6∶2∶2∶6∶2∶6∶6∶1∶2∶2∶6∶6∶6∶6∶1∶1∶2∶1∶2∶2.

[0025] Optionally, the multiplex PCR reaction program is as follows: 95℃ pre-denaturation for 10 min; 95℃ denaturation for 30 s, 59℃ annealing for 90 s, 72℃ extension for 90 s, for 35 cycles; and finally 72℃ extension for 10 min.

[0026] Furthermore, in step (6), the high-throughput sequencing platform used is the Ion PGM platform, but it is not limited to this. Those skilled in the art can choose different high-throughput sequencing platforms according to their needs.

[0027] Furthermore, in step (7), the quantitative method includes the following steps:

[0028] (a) Analyze amplification bias patterns using the added template sequence;

[0029] (b) Correcting sequence errors caused by base mutation bias;

[0030] (c) Standardize sample sequencing data using external reference cells;

[0031] (d) Accurately quantify TCR β in the sample.

[0032] Furthermore, in step (a), the number of sequencing reads containing template sequences of different V genes is counted using molecular barcodes. The template number is used to examine the amplification bias pattern of the template sequences after sample contamination, and the amplification bias index is calculated. The formula for calculating the amplification bias index is as follows:

[0033]

[0034] i = 1…23, n = 23, Count(V i V is the template sequence obtained from sequencing. i The number; if N(s) is the frequency of the CDR3 sequence s, V i If the V gene type is s, then its corrected frequency N′(s) = N(s) × ABI(V) i ).

[0035] Furthermore, in step (b), the Dayhoff method is used to construct a substitution matrix to calculate the similarity between sequences in the complementarities determining region 3 (CDR3) of TCR β, in order to correct sequence errors caused by base mutation bias. The specific steps are as follows: the obtained substitution matrix is ​​used as a parameter for double sequence alignment, the similarity score between sequences is calculated, the similarity threshold between the original sequence and the erroneous sequence is determined, and the low-frequency erroneous sequence is merged into the high-frequency sequence based on this threshold to achieve sequencing error correction.

[0036] Furthermore, in step (c), the sequencing data of the samples is standardized using external reference cells: assuming the number of external reference cells added is n, the number of reads measured is m, and the number of reads for a certain CDR3 is k, then after standardization, the number p of cells corresponding to this CDR3 is...

[0037]

[0038] A second aspect of the present invention provides a composition for quantifying TCR β based on high-throughput sequencing, comprising: a reverse sequence as shown in SEQ ID NO.2; a set of multiplex PCR primers with optimized sequence composition and concentration as shown in SEQ ID NO.3-25, the sequences of which are supplemented with high-throughput sequencing adapters and can be used to construct a high-throughput sequencing library of TCR β; a set of template sequences with molecular barcodes, each template having three molecular barcodes (BC) of length 6 as shown in SEQ ID NO.26-48; and an external reference cell.

[0039] Furthermore, the external reference cell is a hybridoma 2B4 cell.

[0040] As described above, the method for quantifying TCR β based on high-throughput sequencing of the present invention has the following beneficial effects:

[0041] This invention, by adding external reference cells, specifically 2B4 hybridoma cells, can effectively estimate the total number of cells in a sample; by further using optimized multiplex PCR primers, it can better reduce interference and amplification bias between multiplex PCR primers; by adding a template sequence containing three molecular barcodes, it can help correct and standardize TCR high-throughput sequencing data, and finally obtain accurate and true TCR β library distribution.

[0042] This invention further optimizes the primer sequences and primer concentration ratios for multiplex PCR, designs and synthesizes 23 TCR β sequence templates with molecular barcodes, and introduces hybridoma cells 2B4 as external reference cells. The template sequences and external reference cells are added as external references to T cell receptor library sequencing samples, establishing a reliable and accurate TCR quantification method. Attached Figure Description

[0043] Figure 1 The diagram shown is a flowchart of the library construction and sequencing process in Embodiment 1 of the present invention.

[0044] Figure 2 The diagram shown is a data processing flowchart of Embodiment 2 of the present invention.

[0045] Figure 3 The diagram shows a substitution matrix for Embodiment 2 of the present invention.

[0046] Figure 4 The diagram shown is a schematic diagram of the sequence similarity threshold determination method in Embodiment 2 of the present invention.

[0047] Figure 5 The diagram shown is a schematic diagram of the sequence frequency threshold determination method in Embodiment 2 of the present invention.

[0048] Figure 6 The diagram shown is a schematic diagram of the sequencing error correction method of Embodiment 2 of the present invention.

[0049] Figure 7 The image shown is an example of data after correction and standardization in Embodiment 2 of the present invention. Detailed Implementation

[0050] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.

[0051] This invention provides a method for quantifying TCR β based on high-throughput sequencing.

[0052] This invention, by adding external reference cells, specifically 2B4 hybridoma cells, can effectively estimate the total number of cells in a sample; by further using optimized multiplex PCR primers, it can better reduce interference and amplification bias between multiplex PCR primers; by adding a template sequence containing three molecular barcodes, it can help correct and standardize TCR high-throughput sequencing data, and finally obtain accurate and true TCR β library distribution.

[0053] This invention further optimizes the primer sequences and primer concentration ratios for multiplex PCR, designs and synthesizes 23 TCR β sequence templates with molecular barcodes, and introduces hybridoma cells 2B4 as external reference cells. The template sequences and external reference cells are added as external references to T cell receptor library sequencing samples, establishing a reliable and accurate TCR quantification method.

[0054] The following examples use CD3+ T cells sorted from mouse spleens as an example to demonstrate library construction, sequencing, and data analysis.

[0055] The present invention will now be described in detail through specific embodiments.

[0056] Example 1

[0057] Construction and sequencing of a TCR β library from mouse spleen CD3+ T cells.

[0058] 1. RNA extraction

[0059] Strain 1,000,000 mouse spleen CD3+ T cells, add 800 μL Trizol (TRIzol Reagent, Invitrogen, 15596018), mix by pipetting, incubate at room temperature for 5 min, add 200 μL Trizol solution containing 200 2B4 hybridoma cells; add 200 μL chloroform, mix by inversion for 30 s, incubate at room temperature for 3 min; centrifuge at 12000 g for 15 min at 4 °C; transfer the supernatant to another EP tube; add isopropanol at 0.5 ml / ml Trizol, mix by inversion, incubate at room temperature for 10 min; centrifuge at 12000 g for 10 min at 4 °C; add 1 ml 75% ethanol / ml Trizol solution... Trizol was added to 75% ethanol, and the precipitate was gently shaken to suspend it. The precipitate was centrifuged at 8000g for 5 minutes at 4°C, and the supernatant was removed. The precipitate was air-dried at room temperature for 5-10 minutes. The precipitate was dissolved in an appropriate volume of RNase-free water to obtain total mouse RNA, and the concentration was determined and quality controlled.

[0060] The extracted RNA was analyzed using epoch measurement to determine its concentration and quality. The results showed that the OD260 / OD280 ratio was approximately 1.8–2.0.

[0061] Denaturing gel electrophoresis was used to detect the bands at 28s, 18s, and 5s. The results showed that the bands were distinct, with the 28s / 18s band ratio around 2.0.

[0062] The above test results indicate that the quality of the RNA extracted in this embodiment meets the requirements for library construction and can be used for subsequent library construction.

[0063] 2. Reverse transcription

[0064] The total RNA sample obtained in step 1 was reverse transcribed using the RevertAid First StrandcDNA Synthesis Kit (Thermo Scientific, K1622). The specific procedures were performed according to the kit's instructions. The detailed reverse transcription procedure is as follows:

[0065] ① Take 0.1 μg of RNA, prepare the reaction system as shown in Table 1, and then incubate it in a PCR instrument at 72°C for 3 min, followed by immediate freezing for 5 min.

[0066] Table 1

[0067] Reagent Volume(ul) TRBC primer (10µM) 1 water 11-x RNA 0.1ug (xul) total 12

[0068] The sequence of TRBC is: ACTGTGGACCTCCTTGCCA (SEQ ID NO.1).

[0069] ② Add the above products to the reaction system shown in Table 2, and then incubate in a PCR instrument at 42℃ for 60 min and 70℃ for 10 min to obtain the reverse transcription product cDNA.

[0070] Table 2

[0071] Reagent Volume(u1) 5X first strand buffer 4 dNTPs 2 RNase Inhibitor 1 RevertAid Reverse Transcriptase 1 Total 20

[0072] 3. Multiplex PCR

[0073] Prepare the reaction system as shown in Table 3:

[0074] Table 3

[0075] Reagent Volume(ul) mPCR premix 25 FW-primer mix 5 RW-primer 5 Template mix 1 cDNA 5 Water 9 Total volume 50

[0076] The sequence of the reverse primer is shown in SEQ ID NO.2:

[0077] CCATCTCATCCCTGCGTGTCTCCGACTCAG <barcode>AGACCTTGGGTGGAGTCAC.

[0078] The composition and ratio of the FW-primer mix in Table 3 are shown in Table 4:

[0079] Table 4 Composition and Proportion of FW-primer mix

[0080]

[0081] The components of Template mix in Table 3 are shown in Table 5:

[0082] Table 5. Template mix composition

[0083]

[0084]

[0085]

[0086]

[0087] The final concentration was 200 copies / Template / ul. The reaction program was as follows: 95℃ pre-denaturation for 10 min; 95℃ denaturation for 30 s, 59℃ annealing for 90 s, 72℃ extension for 90 s, for 35 cycles; and finally 72℃ extension for 10 min.

[0088] 4. Gel Recovery (QI Aquick Gel Extraction Kit):

[0089] Prepare a 3% TAE agarose gel (low melting point agarose gel) and electrophoresis at 50V for 3 hours. Then, under UV light, cut the gel containing the target band and place it in a 1.5ml EP tube. Add 1ml of QG solubilization buffer and incubate at 45℃ for 5-10 minutes until the gel is completely dissolved, then incubate on ice for 1-2 minutes. Add the dissolved sol to the adsorption column in 500μl increments, centrifuging at 13,000rpm for 1 minute. If the entire solution cannot be added at once, add it in multiple increments. After centrifugation, discard the waste liquid in the collection tube, return the adsorption column to the collection tube, add 300μl of QG solubilization buffer, and centrifuge at 13,000rpm for 1 minute. Discard the waste liquid in the collection tube, return the adsorption column to the collection tube, and centrifuge at 13,000rpm for 2 minutes. Place the adsorption column in a new 1.5ml EP tube. Add 30 μL of ultrapure water to the adsorption column in an EP tube, let stand for 2 min, and finally centrifuge at 13,000 rpm for 2 min. Collect the eluent, which is the sample for library construction.

[0090] 5. Document Quality Control

[0091] After library purification, the library concentration of the samples was determined using Qubit 2.0, and the fragment distribution of the libraries in different samples was detected using the Agilent FragmentAnalyzer fully automated capillary electrophoresis system.

[0092] 6. Sequencing

[0093] The obtained library was sequenced using the Ion PGM high-throughput sequencing platform.

[0094] Example 2

[0095] Quantitative analysis of TCR β in mouse spleen CD3+ T cells.

[0096] The number of TCRs measured in external reference cells was used as a reference for the number of T cells in the sample; sequencing errors were corrected using template sequences. Utilizing the assumption that "high-frequency sequences are more likely to be the original correct sequences," a stepwise extraction clustering method was employed to correct sequence errors. The template sequence was isolated from the sequencing sample using molecular barcodes within the template sequence. The number of sequencing reads for different V template sequences was counted, and the amplification bias pattern after contamination was examined to explore methods for correcting amplification bias and to correct sequence errors generated during the sequencing process.

[0097] This embodiment uses mouse spleen CD3+ T cell sequencing data (see Example 1) to study the TCR β repertoire characteristics of the mouse spleen. The data processing flowchart is shown below. Figure 1 As shown, the specific processing procedure is as follows:

[0098] 1. Sequencing data of TCR β sequence (i.e., sample), template sequence, and external reference cell sequence isolated from mouse spleen CD3+ T cells.

[0099] Template sequence sequencing data was separated from the sequencing data based on the template tag sequence, and reference cell 2B4 cell sequence sequencing data was separated from the sequencing data based on the 2B4 cell CDR3 sequence. The remaining data were mouse spleen TCR sequence sequencing data.

[0100] 2. Calculate the amplification bias index.

[0101] The sequences were assigned to 23 template sequences based on their template sequence tags, and the frequencies of each template sequence were counted. The amplification bias index was calculated using the following formula:

[0102]

[0103] i = 1...23, n = 23, Count(V i V is the template sequence obtained from sequencing. i The number.

[0104] 3. Calculate the replacement matrix.

[0105] The substitution matrix was calculated based on the actual and sequencing sequences of the 23 templates. Figure 2 The steps are as follows: Align the erroneous sequence and the original sequence (the sequence obtained from sequencing, i.e., the real sequence) using a double sequence alignment method; calculate the relative mutation rate mj of base j (mj refers to the number of times j is substituted by other bases); for each base pair i and j, calculate the number of times j is substituted by i; divide the number of substitutions by the relative mutation rate (mj); normalize j using the frequency of base occurrence; take the common logarithm to obtain the base substitution matrix.

[0106] 4. Calculate the similarity threshold and frequency ratio threshold.

[0107] The obtained substitution matrix is ​​used as a parameter for double sequence alignment. The similarity score between the erroneous sequence and the true sequence is calculated using the substitution matrix. The similarity threshold between the original sequence and the erroneous sequence is determined with a 95% confidence interval. Figure 3 The frequency ratio of erroneous sequences to true sequences is calculated, and the frequency ratio threshold is determined using the erroneous sequence with the highest frequency. Figure 4 ).

[0108] Low-frequency erroneous sequences are merged into the real sequences based on a frequency ratio threshold, thereby achieving sequencing error correction.

[0109] 5. Identify the CDR3 sequence of TCR β in mouse spleen CD3+ T cells.

[0110] Based on the characteristic sequences of the V and J genes of the TCR β chain CDR3, the frequency of each CDR3 sequence was counted from the clone sequences in the sequencing data, and the corresponding V gene, J gene, CDR3 nucleotide sequence and CDR3 amino acid sequence were determined.

[0111] 6. Correct amplification bias.

[0112] The amplification bias of the CDR3 sequence obtained in step 5 is corrected based on the amplification bias index obtained in step 2. Specifically, if N(s) is the frequency of the CDR3 sequence s, V i If the CDR3 sequence s represents the V gene type, then its corrected frequency N′(s) = N(s) × ABI(V) i ).

[0113] 7. Correct sequencing errors.

[0114] like Figure 5 As shown, the steps are as follows: a) Sort the CDR3 sequences from highest to lowest frequency; b) Select the CDR3 sequence with the highest frequency as the cluster center; c) For the remaining un-clustered CDR3 sequences, merge those with a similarity score greater than the similarity threshold, a frequency ratio less than the proportion threshold, and the same V and J genes into the cluster; d) Repeat steps b) and c) until all CDR3 sequences have been clustered.

[0115] 8. Standardize the sample using reference cell 2B4.

[0116] Assuming the number of 2B4 cells added is n, the number of reads measured is m, and the number of reads for a certain CDR3 sequence is k, then after standardization, the number of cells p corresponding to this CDR3 sequence is:

[0117]

[0118] 9. Output the corrected and standardized data.

[0119] The corrected and standardized TCR β data of mouse spleen CD3+ T cells obtained above were then processed according to... Figure 6 Save it in CSV format for later analysis.

[0120] Table 6. Comparison of sorted cell count and corrected cell count

[0121]

[0122] In summary, this invention, by adding 2B4 hybridoma cells as an external reference, can effectively estimate the total number of cells in a sample; by using optimized multiplex PCR primers, it can better reduce interference and amplification bias between multiplex PCR primers; by adding a template sequence containing three molecular barcodes, it can help correct and standardize TCR high-throughput sequencing data, and finally obtain an accurate and realistic distribution of the T cell receptor library.

[0123] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention. SEQUENCE LISTING <110> Army Medical University of the Chinese People's Liberation Army <120> A method for quantifying TCR β based on high-throughput sequencing <130> PCQLJ2110592-HZ <160> 50 <170> PatentIn version 3.5 <210> 1 <211> 19 <212> DNA <213> Artificial <220> <223> TRBC <400> 1 actgtggacc tccttgcca 19 <210> 2 <211> 49 <212> DNA <213> Artificial <220> <223> Reverse primer <400> 2 ccatctcatc cctgcgtgtc tccgactcag agaccttggg tggagtcac 49 <210> 3 <211> 49 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV01 <400> 3 cctctctatg ggcagtcggt gatcaaagag gtcaaatctc ttcccggtg 49 <210> 4 <211> 45 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV02 <400> 4 cctctctatg ggcagtcggt gatgcctcaa gtcgcttcca acctc 45 <210> 5 <211> 48 <212> DNA <213> Artificial <220> <223> trP1‑mmTRBV03 <400> 5 cctctctatg ggcagtcggt gatggtaaag tcatggagaa gtctaaac 48 <210> 6 <211> 46 <212> DNA <213> Artificial <220> <223> trP1‑mmTRBV04 <400> 6 cctctctatg ggcagtcggt gatgcaactc attgtaaacg aaacag 46 <210> 7 <211> 43 <212> DNA <213> Artificial <220> <223> trP1‑mmTRBV05 <400> 7 cctctctatg ggcagtcggt gatacggtgc ccagtcgttt tat 43 <210> 8 <211> 44 <212> DNA <213> Artificial <220> <223> trP1‑mmTRBV12‑1 <400> 8 cctctctatg ggcagtcggt gatggattcc tacccagcag attc 44 <210> 9 <211> 43 <212> DNA <213> Artificial <220> <223> trP1‐mmTRBV12‐2 <400> 9 cctctctatg ggcagtcggt gatggagaga gataaaggaa acc 43 <210> 10 <211> 45 <212> DNA <213> Artificial <220> <223> trP1‐mmTRBV13‐1 <400> 10 cctctctatg ggcagtcggt gattgctggc aaccttcgaa tagga 45 <210> 11 <211> 44 <212> DNA <213> Artificial <220> <223> trP1‐mmTRBV13‐2 <400> 11 cctctctatg ggcagtcggt gatcattatt catatggtgc tggc 44 <210> 12 <211> 46 <212> DNA <213> Artificial <220> <223> trP1‐mmTRBV13‐3 <400> 12 cctcttatg ggcagtcggt gatggctgat cattactca tatgtc 46 <210> 13 <211> 47 <212> DNA <213> Artificial <220> <223> trP1‐mmTRBV14 <400> 13 cctctctatg ggcagtcggt gataggccta aaggaactaa ctccact 47 <210> 14 <211> 45 <212> DNA <213> Artificial <220> <223> trP1‑mmTRBV15 <400> 14 cctctctatg ggcagtcggt gatgatggtg gggctttcaa ggatc 45 <210> 15 <211> 48 <212> DNA <213> Artificial <220> <223> trP1‑mmTRBV16 <400> 15 cctctctatg ggcagtcggt gatgcactca actctgaaga tccagagc 48 <210> 16 <211> 47 <212> DNA <213> Artificial <220> <223> trP1‑mmTRBV17 <400> 16 cctctctatg ggcagtcggt gattctctct acattggctc tgcaggc 47 <210> 17 <211> 44 <212> DNA <213> Artificial <220> <223> trP1‑mmTRBV19 <400> 17 cctctctatg ggcagtcggt gatctctcac tgtgacatct gccc 44 <210> 18 <211> 47 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV20 <400> 18 cctctctatg ggcagtcggt gatcccatca gtcatcccaa cttatcc 47 <210> 19 <211> 44 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV21 <400> 19 cctctctatg ggcagtcggt gatctgctaa gaaaccatgt acca 44 <210> 20 <211> 42 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV23 <400> 20 cctctctatg ggcagtcggt gatcagcctg ggaatcagaa cg 42 <210> 21 <211> 44 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV24 <400> 21 cctctcatag ggcagtcggt gatctaagtg ttcctcgaac tcac 44 <210> 22 <211> 45 <212> DNA <213> Artificial <220> <223> trP1‑mmTRBV26 <400> 22 cctctctatg ggcagtcggt gatccttgca gcctagaaat tcagt 45 <210> 23 <211> 45 <212> DNA <213> Artificial <220> <223> trP1‑mmTRBV29 <400> 23 cctctctatg ggcagtcggt gattacaggg tctcacggaa gaagc 45 <210> 24 <211> 47 <212> DNA <213> Artificial <220> <223> trP1‑mmTRBV30 <400> 24 cctctctatg ggcagtcggt gatcagccgg ccaaacctaa cattctc 47 <210> 25 <211> 45 <212> DNA <213> Artificial <220> <223> trP1‑mmTRBV31 <400> 25 cctctctatg ggcagtcggt gatacgacca attcatccta agcac 45 <210> 26 <211> 366 <212> DNA <213> Artificial <220> <223> TP1 <400> 26 aacacagcga cctcgggtgt gcttgccaaa agcaactaca gtggctgttc actctgcgga 60 gtcctgggga caagaggtc aaatctctc ccggtgctga ttacctggcc acacgggtca 120 ctgatacgga gctgaggctg caagtggcca acatgagcca gggcagaacc ttgtactgca 180 cctgcagtgc agatgcttgg ggacaggggg ctgcttgcaa accagaagt cttctttgt 240 aaaggaacca gactcacagt tgtagtgctt ggaggatctg agaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaagcaga gattgcaac aaaaaaagg ctaccctcgt 360 gtgctt 366 <210> 27 <211> 366 <212> DNA <213> Artificial <220> <223> TP2 <400> 27 aacacagcga cctcgggtgc acggtagcct ctagagttca tgttttccta cagctatcaa 60 aaacttatgg acatcagac tgcctcaagt cgcttccac ctcaagttc aaagaaaaac 120 catttagacc ttcagatcac agctctaaag cctgatgact cggccacata cttctgtgcc 180 agcagccaag acacggtggg actgggggg ccacggtcaa acccgacta caccttcggc 240 tcagggacca ggcttttggt aatagcacgg tgaggatctg aagaatgtga ctccacccaa ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt gtgctt 366 <210> 28 <211> 366 <212> DNA <213> Artificial <220> <223> TP3 <400> 28 aacacagcga cctcgggtgc ggtgtagatg gagtttctgg ttaatttcta caatggtaaa gtcatggaga agtctaaact gtttaaggat cagttttcag ttgaaagacc agatggttca 180. tatttcactc tgaaaatcca acccacagca ctggaggact cagctgtgta cttctgtgcc agcagcttag ccggtgtgggg acaggggcc ggtgtttctg gaatacgct ctattttgga gaggagcc ggctcattgt tgtagcggtg tgaggatctg agaaatgtga ctccacccaa ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt gtgctt 366 <210> 29 <211> 366 <212> DNA <213> Artificial <220> <223> TP4 <400> 29 aacacagcga cctcgggtga cggcgtgctg aagattatgt ttagctacaa taataagcaa 60 ctcattgtaa acgaaacagt tccaaggcgc ttctcacctc agtcttcaga taaagctcat 120 ttgaatcttc gaatcaagtc tgtagagccg gaggactctg ctgtgtatct ctgtgccagc 180 agctaagaac ggcggggact gggggggcac ggcgtttcca acgaaagatt atttttcggt 240 catggaacca agctgtctgt cttggacggc ggaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 30 <211> 366 <212> DNA <213> Artificial <220> <223> TP5 <400> 30 aacacagcga cctcgggtga tcgcaagaag ccgccagagc tcatgtttct ctacaatctt 60 aaacagttga ttcgaaatga gacggtgccc agtcgtttta tacctgaatg cccagacagc 120 tccaagctac tttacatat atctgccgtg gatccagaag actcagctgt ctatttttgt 180 gccagcagcc aagaatcgca gggacagggg gcatcgcact cctatgaaca gtacttcggt 240 cccggcacca ggctcacggt tttagatcgc agaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 31 <211> 366 <212> DNA <213> Artificial <220> <223> TP6 <400> 31 aacacagcga cctcgggtgt cattaaggaa ttaaagttcc ttattcagca tttgaaaag 60 gtggagagag acaaaggatt cctacccagc agattctcag tccaacagtt tgatgactat 120 cactctgaaa tgaacatgag tgccttggaa ctggaggact ctgctatgta cttctgtgcc 180 agctctctct cattagggac tggggggct cattataact atgctgagca gttcttcgga 240 ccagggacac gactcaccgt cctagtcatt agaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 32 <211> 366 <212> DNA <213> Artificial <220> <223> TP7 <400> 32 aacacagcga cctcgggtgt aagggcagga actaaagttc ttcattcagc attatgataa 60 aatggagaga gataaaggaa acctgcccag cagattctca gtccaacagt ttgatgacta 120 tcactctgag atgaacatga gtgccttgga gctagaggac tctgccgtgt acttctgtgc 180 cagctctctc taagggggga cagggggcta agggcaaaca ccgggcagct ctactttggt 240 gaaggctcaa agctgacagt gctggtaagg ggaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 33 <211> 366 <212> DNA <213> Artificial <220> <223> TP8 <400> 33 aacacagcga cctcgggtgg cacgaatggg ctgaggctga tccattactc atatggtgct 60 ggcaaccttc gaataggaga tgtccctgat gggtacaagg ccaccagaac aacgcaagaa 120 gacttcttcc tcctgctgga attggcttct ccctctcaga catctttgta cttctgtgcc 180 agcagtgatg gcacgagggga ctgggggggc gcacgaagtg cagaaacgct gtattttggc 240 tcaggaacca gactgactgt tctcggcacg agaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 34 <211> 366 <212> DNA <213> Artificial <220> <223> TP9 <400> 34 aacacagcga cctcgggtgc tagatatggg ctgaggctga tccattattc atatggtgct 60 ggcagcactg agaaaggaga tatccctgat ggatacaagg cctccagacc aagccaagag 120 aacttctccc tcattctgga gttggctacc ccctctcaga catcagtgta cttctgtgcc 180 agcggtgatg ctagatggga cagggggcct agatgttctg gaaatacgct ctattttgga 240 gaaggaagcc ggctcattgt tgtagctaga tgaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 35 <211> 366 <212> DNA <213> Artificial <220> <223> TP10 <400> 35 aacacagcga cctcgggtgc attggatggg ctgaggctga tccattactc atatgcgct 60 gacagcacgg agaaaggaga tatccctgat gggtacaagg cctccagacc aagccaagag 120 aatttctctc tcattctgga gttggcttc cttctcaga cagctgtata tttctgtgcc 180 agcagtgatg cattggggga ctggggggc cattggaacc aagacaccca gtactttggg 240 ccaggcactc ggctcctcgt gttagcattg ggaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 36 <211> 366 <212> DNA <213> Artificial <220> <223> TP11 <400> 36 aacacagcga cctcgggtgc atattcaggg gccccagctt ctagtttact ttcgggatga 60 ggctgtata gataattcac agttgccctc ggatcgattt tctgctgtga ggcctaaagg 120 aactaactcc actctcaaga tccagtctgc aaagcagggc gacacagcca cctatctctg 180 tgccagcagt ttctcatatt gggacagggg gccatattct cctatgaaca gtacttcggt 240 cccggcacca ggctcacggt tttagcatat tgaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 37 <211> 366 <212> DNA <213> Artificial <220> <223> TP12 <400> 37 aacacagcga cctcgggtgt tccaggacta gagttgctga gctacttccg cagcaagtct 60 cttatggaag atggtggggc tttcaaggat cgattcaaag ctgagatgct aaattcatcc 120 ttctccactc tgaagattca acctacagaa cccaaggact cagctgtgta tctgtgtgcc 180 agcagtttag cttccagggg acagggggct tccaggaacc aagacaccca gtactttggg 240 ccaggcactc ggctcctcgt gttagttcca ggaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 38 <211> 366 <212> DNA <213> Artificial <220> <223> TP13 <400> 38 aacacagcga cctcgggtgg gttatgggcc tggagttcct gacttacttt cgaaatcaag 60 ctcctataga tgattcaggg atgcccaagg aacgattctc agctcagatg cccaatcagt 120 cgcactcaac tctgaagatc cagagcacgc aaccccagga ctcagcggtg tatctttgtg 180 caagcagctt agaggttatg ggacaggggg cggttatcaa actccgacta caccttcggc 240 tcagggacca ggcttttggt aatagggtta tgaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 39 <211> 366 <212> DNA <213> Artificial <220> <223> TP14 <400> 39 aacacagcga cctcgggtga cacatgcacc aaagcttctt cttttctact atgataagat 60 tttgaacagg gaagctgaca cttttgagaa gttccaatcc agtcggccta acaattcttt 120 ctgctctc tacattggct ctgcaggcct agagtattct gccatgtacc tctgtgctag 180 cagtagagaa cacatgggac tggggggca cacatttctg gaatacgct ctattttgga 240 gaaggaagcc ggctcattgt tgtagacaca tgaggatctg agaatgtga ctccaccca 300 ggtctccttg tttgagccat caaagcaga gattgcaac aaaaaaagg ctaccctcgt 360 gtgctt 366 <210> 40 <211> 366 <212> DNA <213> Artificial <220> <223> TP15 <400> 40 aacacagcga cctcggggtga accggaagga ttgagactga tctactc ataactgaa 60 aacgatcttc aaaaaggcga tctatctgaa ggctatgatg cgtctcgaga gagaagtca 120 tctttctc tcactgtgac atctgcccag agaacgaga tggccgtttt tctctgtgcc 180 agcagtatag aaccggggga cagggggcaa ccggtttcca acgaaagatt atttttcggt 240 catggaacca agctgtctgt cttggaaccg ggaggatctg agaatgtga ctccaccca 300 ggtctccttg tttgagccat caaagcaga gattgcaac aaaaaaagg ctaccctcgt 360 gtgctt 366 <210> 41 <211> 366 <212> DNA <213> Artificial <220> <223> TP16 <400> 41 aacacagcga cctcgggtgc ccctttttga actgatagca ctttctactg tgaactcagc 60 aatcaaatat gaacaaaatt ttacccagga aaaatttccc atcagtcatc ccaacttatc 120 cttttcatct atgacagttt taaatgcata tcttgaagac agaggcttat atctctgtgg 180 tgctagggac cccttgggac tgggggggcc cccttaaca accaggctcc gctttttgga 240 gaggggactc gactctctgt tctagcccct tgaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 42 <211> 366 <212> DNA <213> Artificial <220> <223> TP17 <400> 42 aacacagcga cctcgggtgg agctgtcaaa ttttggtttt actttcagaa tgaagacatc 60 atcgacaaaa tagatatgat tggtaaaaac atttcagcaa aatgccctgc taagaaacca 120 tgtaccatag agatccagtc cagcaagcta acagattcag ctgtgtactt ctgtgctagc 180 agtcaatcga gctggggact gggggggcga gctggttctg gaatacgct ctattttgga 240 gaaggaagcc ggctcattgt tgtaggagct ggaggatctg agaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaagcaga gattgcaac aaaaaaagg ctaccctcgt 360 gtgctt 366 <210> 43 <211> 366 <212> DNA <213> Artificial <220> <223> TP18 <400> 43 aacacagcga cctcgggtga gcaactaaag ttcctgatt acttcagaa tcacagcct 60 cttgatcaaa tagacatggt CAggagaga ttctcagctg tgtgcccctc cagctcactc 120 tgcagcctgg gatcagaac gtgcgaagca gagactcag cactgtactt gtgctccagc 180 agtcaatcag caacgggact gggggggcag siaccaaca ccgggcagct ctactttggt 240 gaaggctcaa agctgacagt gctggagcaa cgaggatctg agaatgtga ctccaccca 300 ggtctccttg tttgagccat caaagcaga gattgcaac aaaaaaagg ctaccctcgt 360 gtgctt 366 <210> 44 <211> 366 <212> DNA <213> Artificial <220> <223> TP19 <400> 44 aacacagcga cctcgggtgc gccggaact tacatttttg attagctttc gaaatgaaga 60 aattatggaa caaacagact tggtcaagaa gagattctca gctaagtgtt cctcgaactc 120 acgctgcatc ctggaaatcc tatcctctga agaagacgac tcagcactgt acctctgtgc 180 cagcagtctg tacgcccggg gacagggggc cgcccgagtg cagaaacgct gtattttggc 240 tcaggaacca gactgactgt tctcgcgccc ggaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 45 <211> 366 <212> DNA <213> Artificial <220> <223> TP20 <400> 45 aacacagcga cctcgggtgg tatcaagtttt aaatttttga ttaactttca gaatcaagaa 60 gttcttcagc aatagacat gactgaaaaa cgattctctg ctgagtgtcc ttcaaactca 120 ccttgcagcc tagaaattca gtcctctgag gcaggagact cagcactgta cctctgtgcc 180 agcagtctgt cgtatcaggg acagggggcg tatcagagtg cagaaacgct gtattttggc 240 tcaggaacca gactgactgt tctcggtatc agaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 46 <211> 366 <212> DNA <213> Artificial <220> <223> TP21 <400> 46 aacacagcga cctcgggtgc gctaactggg gctacagctg attatatct catacgatgt 60 tgatagtaac agcgaaggag acatccctaa aggatacagg gtctcacgga agaagcggga 120 gcatttctcc ctgattctgg attctgctaa aacaaaccag acatctgtgt acttctgtgc 180 tagcagttta tccgctaagg gacaggggc cgctaaaacc aagacaccca gtactttggg 240 ccaggcactc ggctcctcgt gttagcgcta agaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 47 <211> 366 <212> DNA <213> Artificial <220> <223> TP22 <400> 47 aacacagcga cctcgggtc cttcaagctt gatgctcatg gcaactgcaa atgaaggctc 60 tgaagccaca tacgagagtg gattcaccaa ggacaagtttt ccaatcagcc ggccaaacct 120 aacattctca acgttgacag tgaacaatgc aaggcctgga gacagcagta tctatttctg 180 tagttctaga gaccttcagg gactgggggg gcccttcact cctatgaaca gtacttcggt 240 cccggcacca ggctcacggt tttagccttc agaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 48 <211> 366 <212> DNA <213> Artificial <220> <223> TP23 <400> 48 aacacagcga cctcgggtgg gaatacacag gaggcaccct ccagcaactc ttctactcta 60 ttactgttgg ccaggtagag tcggtggtgc aactgaacct ctcagcttcc aggccgaagg 120 acgaccaatt catcctaagc acggagaagc tgcttctcag ccactctggc ttctacctct 180 gtgcctggag tctggaatag ggacaggggg cggaatacaa acacagaagt cttctttggt 240 aaaggaacca gactcacagt tgtagggaat agaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 49 <211> 30 <212> DNA <213> Artificial <220> <223> Adapter sequence 1 <400> 49 ccatctcatc cctgcgtgtc tccgactcag 30 <210> 50 <211> 23 <212> DNA <213> Artificial <220> <223> Adapter sequence 2 <400> 50 cctctctatg ggcagtcggt gat 23< / barcode> < / barcode>

Claims

1. A method for quantifying TCR β based on high-throughput sequencing, characterized in that, Includes the following steps: (1) Use Trizol to lyse the sample and add a fixed number of external reference cell lysis buffer to the lysed sample; the external reference cells are 2B4 hybridoma cells; (2) Extract total RNA from the sample and the reference cells; (3) Reverse transcription was performed using C-terminal specific primers of TCR β; (4) A fixed number of templates are added to the reverse transcription product. The templates have 23 sequences as shown in SEQ ID NO.26-48. The template sequences consist of the V gene, three molecular barcodes of length 6, the D gene, the J gene, and the C gene. (5) A high-throughput sequencing library of TCR β was constructed using a set of multiplex PCR primers with optimized sequence composition and concentration; the sequences of the multiplex PCR primers are shown in SEQ ID NO.3-25, the sequences of SEQ ID NO.3-25 are filled with high-throughput sequencing adapters, and the reverse sequences are shown in SEQ ID NO.2; (6) Sequencing was performed using a high-throughput sequencing platform; (7) Quantify the TCR β in the sample using two rounds of external parameter data. The quantification method includes the following steps: (a) Analysis of amplification bias using the added template sequences: The number of sequencing reads containing template sequences of different V genes was counted using molecular barcodes. The amplification bias of the template sequences after sample contamination was examined using the template number, and the amplification bias index was calculated. The formula for calculating the amplification bias index is as follows: i=1…23, n=23, Count(V i V is the template sequence obtained from sequencing. i The number; if N(s) is the frequency of the CDR3 sequence s, V i If the V gene type is s, then its corrected frequency N'(s) = N(s) × ABI(V) i ); (b) Correction of sequence errors caused by base mutation bias: The Dayhoff method is used to construct a substitution matrix to calculate the similarity between sequences in the complementarity determination region 3 of TCRβ in order to correct the errors generated during sequencing. The specific steps are as follows: the obtained substitution matrix is ​​used as a parameter for double sequence alignment, the similarity score between sequences is calculated, the similarity threshold between the original sequence and the error sequence is determined, and the low-frequency error sequence is merged into the high-frequency sequence based on this threshold to achieve sequencing error correction. (c) Standardizing sample sequencing data using external reference cells: Assuming the number of external reference cells added is n, the number of reads measured is m, and the number of reads for a certain CDR3 is k, then after standardization, the number p of cells corresponding to this CDR3 is... ; (d) Accurately quantify TCR β in the sample.

2. The method according to claim 1, characterized in that: In step (2), the total RNA was extracted using the Trizol method.

3. The method according to claim 1, characterized in that: In step (3), the C-terminal specific primer for TCR β is TRBC, and its sequence is shown in SEQ ID NO.1; In step (3), the reverse transcription using the C-terminal specific primer of TCR β is performed as follows: ① Take 0.1ug of RNA from step (2), 1ul of 10uM primer TRBC, and the remainder of water to prepare a 12ul reaction system. Then incubate the system in a PCR instrument at 72℃ for 3min and immediately place it on ice for 5min. ② Prepare a 20ul reaction system by mixing the product obtained in step ①, 4ul 5X first strand buffer, 2ul dNTPs, 1ul RNase inhibitor, and 1ul RevertAid reverse transcriptase. Then incubate the mixture in a PCR instrument at 42℃ for 60min and 70℃ for 10min.

4. The method according to claim 1, characterized in that: In step (5), the multiplex PCR reaction system consists of 50 μL, including the following reaction components: 25 μL mPCR premix, 5 μL forward primer FW-primer mix, 5 μL reverse primer RW-primer, 1 μL template mix, 5 μL cDNA, and 9 μL water; The forward primer (FW-primer mix) consists of the sequence shown in SEQ ID NO.3-25, with the sequence ratio of SEQ ID NO.3-25 being 1:2:6:6:2:2:6:2:6:6:1:2:2:6:6:6:6:1:1:2:1:2:

2.

5. The method according to claim 1, characterized in that: In step (6), the high-throughput sequencing platform used is the IonPGM platform.

6. A composition for quantifying T-cell receptors based on high-throughput sequencing, characterized in that, The invention includes a reverse sequence as shown in SEQ ID NO.2; a set of multiplex PCR primers with optimized sequence composition and concentration as shown in SEQ ID NO.3-25, the sequences of which are supplemented with high-throughput sequencing adapters and can be used to construct high-throughput sequencing libraries for TCR; a set of template sequences with molecular barcodes, each template having three 6-bit molecular barcodes BC, the sequences of which are shown in SEQ ID NO.26-48; and an external reference cell, the external reference cell being hybridoma 2B4 cells.

Citation Information

Patent Citations

  • Immune repertoire standard substance sequence as well as design method and application thereof

    CN111850016A