Method for correcting and normalizing TCR β high-throughput sequencing data based on template sequences and reference cells
By incorporating ginseng cells and templates into the samples, amplification bias and correcting sequencing errors were solved, and the problems existing in the sequencing of T cell receptor library were achieved, precise quantification and standardization of TCRβ library were achieved, and accurate distribution of T cell receptor library was obtained.
Patent Information
- Application Number
- CN202210213693.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-03
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-03-03
AI Technical Summary
In the prior art, sequencing errors and amplification bias during deep sequencing of T cell receptor libraries seriously affect the estimation of the diversity of T cell libraries, and there is a lack of effective correction methods.
By incorporating a fixed number of ginseng cells and synthetic templates into the sample, the amplification bias pattern is analyzed using template sequences, sequencing errors are corrected, and the sample sequencing data is standardized using ginseng cells to achieve accurate quantification of the TCR β library.
The correction and standardization of TCRβ high-throughput sequencing data was achieved, and accurate and true T cell receptor library distribution was obtained.
Smart Images

Figure CN114596915B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and in particular to a method for correcting and standardizing TCRβ high-throughput sequencing data based on a template sequence and a reference cell. Background Art
[0002] The T cell receptor (TCR) is a specific receptor on the surface of T cells, which is responsible for recognizing antigens presented by the major histocompatibility complex (MHC) and mediating immune responses. Understanding the diversity composition of the T cell receptor repertoire helps us understand the immune status of the body, and further clarify the internal causes of the occurrence and development of immune diseases, providing assistance for the development of related vaccines and the treatment of diseases. The complementarity-determining region 3 (CDR3) on the β subunit of the T cell receptor is a very important region on the TCR receptor. This region has the strongest binding ability to antigen peptides, is also the region with the highest diversity, and best represents the diversity of TCR. Therefore, most researchers study the diversity of the T cell receptor immunoglobulin repertoire by studying the diversity of the CDR3 of the T cell receptor β chain (TCRβCDR3).
[0003] The research on the T cell receptor repertoire has gone through three main development stages technically, which is a process from rough to fine. The initial flow cytometry can only analyze the distribution and deletion of each T cell subfamily using monoclonal antibodies against each T cell subfamily, obtaining relatively rough results. Later, researchers proposed the immune scanning lineage analysis technology based on the TCR gene rearrangement rules and the characteristics of TCR gene family homology. Compared with flow cytometry, this technology can not only analyze the distribution of each T cell subfamily, but also analyze the distribution law of the CDR3 length in the TCR repertoire, but it still cannot analyze specific TCR sequences. With the development of high-throughput sequencing technology, researchers have developed the T cell receptor sequencing technology (TCR-seq), which can sequence and analyze all TCRs in a sample, obtain the genetic information of all T cell receptors, and comprehensively reveal the complexity and diversity of the T cell receptor repertoire. However, during the deep sequencing of the T cell receptor repertoire, sequencing errors seriously affect the estimation of the diversity of the T cell repertoire. Moreover, during the library construction process, multiple PCR amplification of the CDR3 sequences in the sample is required, and the interference between multiple primers and the different amplification efficiencies will cause amplification bias. It can be seen that at present, the correction problem of T cell receptor repertoire sequencing data has not been well solved. Therefore, it is necessary to establish an effective method to correct amplification bias, PCR and sequencing errors to promote the research of the T cell receptor repertoire. Summary of the Invention
[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a method for correcting and standardizing TCRβ high-throughput sequencing data based on a template sequence and exogenous reference cells. By introducing the template sequence and exogenous reference cells, the present invention develops a method for correcting and standardizing TCRβ library sequencing data and establishes a method for quantitative analysis of TCRβ libraries.
[0005] To achieve the above object and other related objects, the present invention provides a method for correcting and standardizing TCR high-throughput sequencing data based on a template sequence and reference cells, comprising the following steps:
[0006] (a) Incorporate a fixed number of exogenous reference cells and a fixed number of synthetic templates into the sample, construct a high-throughput sequencing library of TCRβ, and perform sequencing using a high-throughput sequencing platform;
[0007] (b) Analyze the amplification bias pattern using the added template sequence;
[0008] (c) Correct the sequencing errors generated during the sequencing process;
[0009] (d) Standardize the sample sequencing data using the exogenous reference cells;
[0010] (e) Accurately quantify TCRβ in the sample.
[0011] Furthermore, step (a) includes the following steps:
[0012] (1) Lyse the sample using Trizol, and add a fixed number of exogenous reference cell lysates to the lysed sample;
[0013] (2) Extract the total RNA of the sample and the exogenous reference cells;
[0014] (3) Perform reverse transcription using a C-terminal specific primer of TCRβ;
[0015] (4) Add a fixed number of templates to the reverse-transcribed sample;
[0016] (5) Construct a high-throughput sequencing library of TCRβ using a set of multiplex PCR primers with optimized sequence composition and usage concentration;
[0017] (6) Perform sequencing using a high-throughput sequencing platform.
[0018] Furthermore, in step (2), the method for total RNA extraction is the Trizol method, but is not limited to the Trizol method.
[0019] Further, in step (3), the C-terminal specific primer for TCRβ is TRBC, and its sequence is as shown in SEQ ID NO.1, but is not limited to the primer of SEQ ID NO.1. Those skilled in the art can design and use it according to actual needs.
[0020] Optionally, in step (3), the steps of reverse transcription using the C-terminal specific primer of TCRβ are as follows:
[0021] ① Take 0.1 μg of the RNA in step (2), 1 μl of primer TRBC (10 μM), and the rest is water, and prepare a 12 μl reaction system. Then incubate it at 72 °C for 3 min in a PCR instrument, and quickly place it on ice for 5 min;
[0022] ② Mix the product obtained in step ①, 4 μl of 5X first strand buffer, 2 μl of dNTPs, 1 μl of RNase inhibitor, and 1 μl of RevertAid reverse transcriptase to prepare a 20 μl reaction system. Then incubate it at 42 °C for 60 min and at 70 °C for 10 min in a PCR instrument.
[0023] Optionally, in step (4), there are 23 templates, and the sequences are as shown in SEQ ID NO.26 - 48. The template sequences of the present invention are composed of a V gene, three molecular barcodes (BCs) with a length of 6, a D gene, a J gene, and a C gene, which reflect the sequence characteristics of TCRβ. Specifically, primer binding sites are included in the V gene and the C gene. Since there are only 23 functional V genes, 23 template sequences are designed and synthesized with different V genes in the present invention, and the length of this sequence is 366 bp. The length of the molecular barcode is not limited to 6, and those skilled in the art can adjust it according to actual needs.
[0024] Optionally, in step (5), the multiplex PCR primers are primers optimized based on the invention patent with the application number CN201510027905, "Multiplex PCR Primers and Methods for Constructing a Mouse TCRB Library Based on High-Throughput Sequencing", and their use concentrations are also optimized to make the amplification of each TCRβ better reflect the true TCRβ situation in the sample.
[0025] Optionally, in step (5), the sequences of the multiplex PCR primers are as shown in SEQ ID NO.3 - 25, and the sequences of SEQ ID NO.3 - 25 are added with high-throughput sequencing adapters.
[0026] The reverse sequence is as shown in SEQ ID NO.2:
[0027] CCATCTCATCCCTGCGTGTCTCCGACTCAG<barcode>AGACCTTGGGTGGAGTCAC。
[0028] Among the sequences SEQ ID NO.2 - 25, adapter sequence 1: CCATCTCATCCCTGCGTGTCTCCGACTCAG (SEQ ID NO.49) and adapter sequence 2: CCTCTCTATGGGCAGTCGGTGAT (SEQ ID NO.50) are adapter sequences for the Ion PGM platform. Specifically, different adapter sequences can be selected according to different high - throughput sequencing platforms.
[0029] Furthermore, in step (5), the multiplex PCR reaction system is a total of 50 μl, including the following reaction components: 25 μl of mPCR premix, 5 μl of forward primer (FW - primer mix), 5 μl of reverse primer (RW - primer), 1 μl of Template mix, 5 μl of cDNA, and 9 μl of water.
[0030] The forward primer (FW - primer mix) consists of the sequences shown in SEQ ID NO.3 - 25, and the proportions of the sequences shown in SEQ ID NO.3 - 25 are 1∶2∶6∶6∶2∶2∶6∶2∶6∶6∶1∶2∶2∶6∶6∶6∶6∶1∶1∶2∶1∶2∶2 in sequence.
[0031] Optionally, the multiplex PCR reaction program is: pre - denaturation at 95°C for 10 min; denaturation at 95°C for 30 s, annealing at 59°C for 90 s, extension at 72°C for 90 s, for 35 cycles; and finally post - extension at 72°C for 10 min.
[0032] Furthermore, in step (6), the high - throughput sequencing platform used is the Ion PGM platform, but it is not limited to this platform. Those skilled in the art can select different high - throughput sequencing platforms according to requirements.
[0033] Furthermore, in step (a), the T - cell receptor sequence of the external reference cells is different from the T - cell receptor sequence in the sample; preferably, the external reference cells are 2B4 hybridoma cells, but not limited to 2B4 hybridoma cells. As long as their TCR sequences are different from the TCR sequences in the sample, they can be used as external reference cells; the number of 2B4 hybridoma cells used in the examples of the present invention is 200, and the specific number of external reference cells can be adjusted according to the number of T cells in the sample.
[0034] Further, in step (a), there are 23 templates, and the sequences are as shown in SEQ ID NO. 26-48. The template sequences are composed of a V gene, three molecular barcodes (BCs) with a length of 6, a D gene, a J gene, and a C gene, reflecting the sequence characteristics of TCRβ. Specifically, the sites for binding amplification primers are included in the V gene and the C gene. Since there are only 23 functional V genes, 23 template sequences were designed and synthesized using different V genes in the present invention, and the length of this sequence is 366 bp. The length of the molecular barcode is not limited to 6, and those skilled in the art can adjust it according to actual needs.
[0035] Further, in step (b), the number of sequencing reads of the template sequences containing different V genes is counted using the molecular barcodes, and the amplification bias law of the template sequences after mixing into the sample is investigated using the number of templates, and the amplification bias index is calculated. The formula for the amplification bias index is as follows:
[0036]
[0037] i = 1…23, n = 23, Count(V i ) is the number of the template sequence V i obtained by sequencing; if N(s) is the frequency of the CDR3 sequence s, and V i is the V gene type of s, then its corrected frequency N′(s) = N(s) × ABI(V i ).
[0038] Further, in step (c), the Dayhoff method is used to construct a substitution matrix for calculating the similarity between the complementarities determining region 3 (CDR3) sequences of TCRβ to correct the sequence errors generated during sequencing. The specific steps are as follows: The obtained substitution matrix is used as a parameter for pairwise sequence alignment to calculate the similarity score between sequences, determine the similarity threshold between the original sequence and the error sequence, and based on this threshold, the low-frequency error sequences are merged into the high-frequency sequences to achieve sequencing error correction.
[0039] Further, in step (d), the external reference cells are used to standardize the sample sequencing data: Assume that the number of added external reference cells is n, the number of measured reads is m, and the number of reads of a certain CDR3 is k. Then, after standardization, the number of cells p corresponding to this CDR3 is
[0040]
[0041] As described above, the method for TCRβ high-throughput sequencing data correction and standardization based on template sequences and reference cells of the present invention has the following beneficial effects:
[0042] The present invention uses a high-throughput sequencing method and a standardized process. Through amplification bias correction, sequencing error correction, and sample standardization, it can correct and standardize TCR high-throughput sequencing data, and finally obtain an accurate and real T cell receptor repertoire distribution. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It shows a data processing flow chart provided by an embodiment of the present invention.
[0044] Figure 2 It shows a schematic diagram of an alternative matrix provided by an embodiment of the present invention.
[0045] Figure 3 It shows a schematic diagram of a method for determining a sequence similarity threshold provided by an embodiment of the present invention.
[0046] Figure 4 It shows a schematic diagram of a method for determining a sequence frequency threshold provided by an embodiment of the present invention.
[0047] Figure 5 It shows a schematic diagram of a sequencing error correction method provided by an embodiment of the present invention.
[0048] Figure 6 It shows an example of data after correction and standardization. DETAILED DESCRIPTION OF THE INVENTION
[0049] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.
[0050] The following examples use mouse spleen CD3+ T cells as an example to perform library construction, sequencing, and data analysis. The present invention will be described in detail through specific examples below.
[0051] Example 1
[0052] Construction and sequencing of the TCRβ library of mouse spleen CD3+ T cells.
[0053] 1. RNA extraction
[0054] Sort 1,000,000 mouse spleen CD3+ T cells, add 800 μl of Trizol (TRIzol Reagent, Invitrogen, 15596018), pipette to mix well, let stand at room temperature for 5 min, and add 200 μl of Trizol solution containing 200 2B4 hybridoma cells; add 200 μl of chloroform, invert to mix well for 30 s, and let stand at room temperature for 3 min; centrifuge at 12,000 g for 15 min at 4°C; aspirate the upper aqueous phase into another EP tube; add isopropanol at a ratio of 0.5 ml of isopropanol / ml of Trizol, invert to mix well, and let stand at room temperature for 10 min; centrifuge at 12,000 g for 10 min at 4°C; add 75% ethanol at a ratio of 1 ml of 75% ethanol / ml of Trizol, gently shake to suspend the precipitate; centrifuge at 8,000 g for 5 min at 4°C, aspirate the supernatant; air-dry at room temperature for 5 - 10 min; dissolve with an appropriate volume of RNase-free water to obtain total mouse RNA, and perform concentration determination and quality control.
[0055] Measure the concentration and quality of the extracted RNA using epoch. The detection results show that the OD260 / OD280 is around 1.8 - 2.0.
[0056] Detect 28s, 18s, and 5s using denaturing gel electrophoresis. The detection results show that the bands are obvious and the 28s / 18s is around 2.0.
[0057] The above detection results indicate that the quality of the RNA extracted in this example meets the requirements for library construction and can be used for subsequent library construction.
[0058] 2. Reverse Transcription
[0059] Perform reverse transcription on the total RNA sample obtained in step 1. When performing reverse transcription, use the RevertAid First Strand cDNA Synthesis Kit (Thermo Scientific, K1622), and the specific operation is carried out according to the kit instructions. The specific reverse transcription operation is as follows:
[0060] ① Take 0.1 μg of RNA, prepare the reaction system as shown in Table 1, then incubate in a PCR instrument at 72°C for 3 min, and quickly place on ice for 5 min.
[0061] Table 1
[0062] Reagent Volume(ul) TRBC primer(10uM) 1 water 11-x RNA 0.1ug(xul) total 12
[0063] Among them, the sequence of TRBC is: ACTGTGGACCTCCTTGCCA (SEQ ID NO.1).
[0064] ② Add the above product to the reaction system shown in Table 2, and then incubate it in a PCR instrument at 42 °C for 60 min and at 70 °C for 10 min to obtain the reverse transcription product cDNA.
[0065] Table 2
[0066] Reagent Each well V(ul) 5X first strand buffer 4 dNTPs 2 RNase Inhibitor 1 RevertAid Reverse Transcriptase 1 Total 20
[0067] 3. Multiplex PCR
[0068] Prepare the reaction system as shown in Table 3:
[0069] Table 3
[0070] Reagent Volume(ul) mPCR premix 25 FW-primner mix 5 RW-primer 5 Template mix 1 cDNA 5 Water 9 Total volume 50
[0071] The sequence of the reverse primer is shown in SEQ ID NO.2:
[0072] CCATCTCATCCCTGCGTGTCTCCGACTCAG<barcode>AGACCTTGGGTGGAGTCAC.
[0073] The composition and proportion of FW-primer mix in Table 3 are shown in Table 4:
[0074] Composition and proportion of FW-primer mix
[0075]
[0076]
[0077] The composition of Template mix in Table 3 is shown in Table 5:
[0078] Composition of Template mix
[0079]
[0080]
[0081]
[0082] The final concentration is 200 copies / Template / ul. The reaction procedure is: pre-denaturation at 95 °C for 10 min; denaturation at 95 °C for 30 s, annealing at 59 °C for 90 s, extension at 72 °C for 90 s, for 35 cycles; finally, post-extension at 72 °C for 10 min.
[0083] 4. Gel extraction (QI Aquick Gel Extraction kit):
[0084] Prepare 3% TAE agarose gel (low melting point agarose gel), electrophorese at 50 V for 3 h; then, under ultraviolet light, cut out the gel containing the target band and place it in a 1.5 ml EP tube; add 1 ml of QG Solubilization buffer, incubate in a water bath at 45 °C for 5 - 10 min until the gel block is completely dissolved, and then ice-bath for 1 - 2 min; add the dissolved sol to the adsorption column, add 500 μl each time, centrifuge at 13,000 rpm for 1 min. If it cannot be added all at once, it can be added in multiple times; after centrifugation, pour out the waste liquid in the collection tube, put the adsorption column back into the collection tube, add 300 μl of QG Solubilization buffer to the adsorption column, centrifuge at 13,000 rpm for 1 min; pour out the waste liquid in the collection tube, put the adsorption column back into the collection tube, centrifuge at 13,000 rpm for 2 min; n; place the adsorption column in a new 1.5 ml EP tube, add 30 μl of ultrapure water to the adsorption column, let it stand for 2 min, and finally, centrifuge at 13,000 rpm for 2 min. Collect the eluate, which is the library sample for library construction.
[0085] 5. Library quality control
[0086] After the library purification is completed, use Qubit 2.0 to measure the library concentration of different samples, and use the Agilent Fragment Analyzer automated capillary electrophoresis system to detect the fragment distribution of the libraries of different samples.
[0087] 6. Sequencing
[0088] Sequence the obtained library through the Ion PGM high-throughput sequencing platform.
[0089] Example 2
[0090] Quantitative analysis of TCRβ of mouse spleen CD3+ T cells.
[0091] The number of TCRs measured by evaluating reference cells is used as a reference for the number of T cells in the sample; correct sequencing errors through the template sequence. Using the assumption that "high-frequency sequences are more likely to be the original correct sequences", use the stepwise extraction clustering method to correct sequencing errors. Use the molecular barcodes in the template sequence to separate the template sequence from the sequencing samples, count the number of sequencing reads of different V template sequences, investigate the law of amplification bias after mixing into the samples, and explore methods to correct amplification bias to correct sequencing errors caused by base mutation bias.
[0092] This example is based on the sequencing data of mouse spleen CD3+ T cells (see Example 1) to study the TCRβ repertoire characteristics of mouse spleen. The data processing flow chart is as Figure 1 shown, and the specific processing process is as follows:
[0093] 1. Isolate the sequencing data of the TCRβ sequences (i.e., samples), template sequences, and reference cell sequences of mouse spleen CD3+ T cells.
[0094] According to the tag sequences of the templates, the template sequence sequencing data is isolated from the sequencing data. According to the CDR3 sequences of 2B4 cells, the reference cell 2B4 cell sequence sequencing data is isolated from the sequencing data. The remaining is the mouse spleen TCR sequence sequencing data.
[0095] 2. Calculate the amplification bias index.
[0096] According to the template sequence tags, the sequences are assigned to 23 template sequences, and the frequencies of various template sequences are counted. Calculate the amplification bias index, and the calculation formula of the amplification bias index is as follows:
[0097]
[0098] i = 1…23, n = 23, Count(Vi) is the number of the template sequence Vi obtained by sequencing.
[0099] 3. Calculate the substitution matrix.
[0100] Calculate the substitution matrix ( Figure 2 ) according to the true sequences and sequencing sequences of 23 templates. The steps are as follows: Align the error sequences and the original sequences (refers to the sequences obtained by sequencing, that is, the true sequences) using the pairwise sequence alignment method; calculate the relative mutation rate mj of base j (rnj refers to the number of times j is replaced by other bases); for each base pair i and j, calculate the number of times j is replaced by i; divide the number of substitutions by the relative mutation rate (mj); standardize j using the frequency of the base appearance; take the common logarithm to obtain the base substitution matrix.
[0101] 4. Calculate the similarity threshold and the frequency proportion threshold.
[0102] Take the obtained substitution matrix as the parameter of pairwise sequence alignment, calculate the similarity score between the error sequence and the true sequence using the substitution matrix, and determine the similarity threshold ( Figure 3 ) between the original sequence and the error sequence with a 95% confidence interval. Count the frequency proportion of the error sequence and the true sequence, and determine the frequency proportion threshold ( Figure 4 ) using the error sequence with the highest frequency.
[0103] Based on the frequency proportion threshold, merge the low-frequency error sequences into the true sequences to achieve sequencing error correction.
[0104] 5. Identify the CDR3 sequence of TCRβ of mouse spleen CD3+ T cells.
[0105] Based on the V gene and J gene characteristic sequences of CDR3 of the TCRβ chain, identify the clone sequences in the sequencing data, and count the frequencies of each CDR3 sequence, the corresponding V gene, J gene, CDR3 nucleotide sequence, and CDR3 amino acid sequence.
[0106] 6. Correct amplification bias.
[0107] Correct the amplification bias of the CDR3 sequences obtained in step 5 according to the amplification bias index obtained in step 2. Specifically: If N(s) is the frequency of the CDR3 sequence s, and V i is the V gene type of the CDR3 sequence s, then its corrected frequency N′(s) = N(s) × ABI(V i ).
[0108] 7. Correct sequencing errors.
[0109] As Figure 5 shown, the steps are as follows: a) Arrange the CDR3 sequences in descending order according to the frequency, b) Take the CDR3 sequence with the highest frequency as the clustering center, c) For the remaining unclustered CDR3 sequences, merge the CDR3 sequences with a similarity score greater than the similarity threshold, a frequency ratio less than the ratio threshold, and the same V gene and J gene into this cluster. d) Repeat b) and c) until all CDR3 sequences have been clustered.
[0110] 8. Standardize the sample using the reference cell 2B4.
[0111] Assume that the number of 2B4 cells added is n, the number of reads measured is m, and the number of reads of a certain CDR3 sequence is k. Then, after standardization, the number of cells p corresponding to this CDR3 sequence is
[0112]
[0113] 9. Output the corrected and standardized data.
[0114] Save the TCRβ data of the corrected and standardized mouse spleen CD3+ T cells obtained above in the Figure 6 csv format for subsequent analysis.
[0115] Table 6 Comparison chart of sorted cell number and corrected cell number
[0116]
[0117] In summary, by adding the external reference cell 2B4 hybridoma cells, the total cell number of the sample can be effectively estimated; by using the optimized multiplex PCR primers, the interference and amplification bias between the multiplex PCR primers can be better reduced; by adding the template sequence containing three molecular barcodes, it can help correct and standardize the TCR high-throughput sequencing data, and finally obtain the accurate and real T cell receptor repertoire distribution.
[0118] The above embodiments merely illustrate the principles and effects of the present invention, rather than limiting the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention. SEQUENCE LISTING <110> Army Medical University of the Chinese People's Liberation Army <120> Method for Correcting and Standardizing TCR β High-Throughput Sequencing Data Based on Template Sequence and Reference Cells <130> PCQLJ2110591-HZ <160> 50 <170> PatentIn version 3.5 <210> 1 <211> 19 <212> DNA <213> Artificial <220> <223> TRBC <400> 1 actgtggacc tccttgcca 19 <210> 2 <211> 49 <212> DNA <213> Artificial <220> <223> Reverse primer <400> 2 ccatctcatc cctgcgtgtc tccgactcag agaccttggg tggagtcac 49 <210> 3 <211> 49 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV01 <400> 3 cctctctatg ggcagtcggt gatcaaagag gtcaaatctc ttcccggtg 49 <210> 4 <211> 45 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV02 <400> 4 cctctctatg ggcagtcggt gatgcctcaa gtcgcttcca acctc 45 <210> 5 <211> 48 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV03 <400> 5 cctctctatg ggcagtcggt gatggtaaag tcatggagaa gtctaaac 48 <210> 6 <211> 46 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV04 <400> 6 cctctctatg ggcagtcggt gatgcaactc attgtaaacg aaacag 46 <210> 7 <211> 43 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV05 <400> 7 cctctctatg ggcagtcggt gatacggtgc ccagtcgttt tat 43 <210> 8 <211> 44 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV12-1 <400> 8 cctctctatg ggcagtcggt gatggattcc tacccagcag attc 44 <210> 9 <211> 43 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV12-2 <400> 9 cctctctatg ggcagtcggt gatggagaga gataaaggaa acc 43 <210> 10 <211> 45 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV13-1 <400> 10 cctctctatg ggcagtcggt gattgctggc aaccttcgaa tagga 45 <210> 11 <211> 44 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV13-2 <400> 11 cctctctatg ggcagtcggt gatcattatt catatggtgc tggc 44 <210> 12 <211> 46 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV13-3 <400> 12 cctctctatg ggcagtcggt gatggctgat ccattactca tatgtc 46 <210> 13 <211> 47 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV14 <400> 13 cctctctatg ggcagtcggt gataggccta aaggaactaa ctccact 47 <210> 14 <211> 45 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV15 <400> 14 cctctctatg ggcagtcggt gatgatggtg gggctttcaa ggatc 45 <210> 15 <211> 48 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV16 <400> 15 cctctctatg ggcagtcggt gatgcactca actctgaaga tccagagc 48 <210> 16 <211> 47 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV17 <400> 16 cctctctatg ggcagtcggt gattctctct acattggctc tgcaggc 47 <210> 17 <211> 44 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV19 <400> 17 cctctctatg ggcagtcggt gatctctcac tgtgacatct gccc 44 <210> 18 <211> 47 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV20 <400> 18 cctctctatg ggcagtcggt gatcccatca gtcatcccaa cttatcc 47 <210> 19 <211> 44 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV21 <400> 19 cctctctatg ggcagtcggt gatctgctaa gaaaccatgt acca 44 <210> 20 <211> 42 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV23 <400> 20 cctctctatg ggcagtcggt gatcagcctg ggaatcagaa cg 42 <210> 21 <211> 44 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV24 <400> 21 cctctctatg ggcagtcggt gatctaagtg ttcctcgaac tcac 44 <210> 22 <211> 45 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV26 <400> 22 cctctctatg ggcagtcggt gatccttgca gcctagaaat tcagt 45 <210> 23 <211> 45 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV29 <400> 23 cctctctatg ggcagtcggt gattacaggg tctcacggaa gaagc 45 <210> 24 <211> 47 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV30 <400> 24 cctctctatg ggcagtcggt gatcagccgg ccaaacctaa cattctc 47 <210> 25 <211> 45 <212> DNA <213> Artificial <220> <223> trP1-mmTRBV31 <400> 25 cctctctatg ggcagtcggt gatacgacca attcatccta agcac 45 <210> 26 <211> 366 <212> DNA <213> Artificial <220> <223> TP1 <400> 26 aacacagcga cctcgggtgt gcttgccaaa agcaactaca gtggctgttc actctgcgga 60 gtcctgggga caaagaggtc aaatctcttc ccggtgctga ttacctggcc acacgggtca 120 ctgatacgga gctgaggctg caagtggcca acatgagcca gggcagaacc ttgtactgca 180 cctgcagtgc agatgcttgg ggacaggggg ctgcttgcaa acacagaagt cttctttggt 240 aaaggaacca gactcacagt tgtagtgctt ggaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 27 <211> 366 <212> DNA <213> Artificial <220> <223> TP2 <400> 27 aacacagcga cctcgggtgc acggtagcct ctagagttca tgttttccta cagctatcaa 60 aaacttatgg acaatcagac tgcctcaagt cgcttccaac ctcaaagttc aaagaaaaac 120 catttagacc ttcagatcac agctctaaag cctgatgact cggccacata cttctgtgcc 180 agcagccaag acacggtggg actggggggg ccacggtcaa actccgacta caccttcggc 240 tcagggacca ggcttttggt aatagcacgg tgaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 28 <211> 366 <212> DNA <213> Artificial <220> <223> TP3 <400> 28 aacacagcga cctcgggtgc ggtgtagatg gagtttctgg ttaatttcta caatggtaaa 60 gtcatggaga agtctaaact gtttaaggat cagttttcag ttgaaagacc agatggttca 120 tatttcactc tgaaaatcca acccacagca ctggaggact cagctgtgta cttctgtgcc 180 agcagcttag ccggtgtggg acagggggcc ggtgtttctg gaaatacgct ctattttgga 240 gaaggaagcc ggctcattgt tgtagcggtg tgaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 29 <211> 366 <212> DNA <213> Artificial <220> <223> TP4 <400> 29 aacacagcga cctcgggtga cggcgtgctg aagattatgt ttagctacaa taataagcaa 60 ctcattgtaa acgaaacagt tccaaggcgc ttctcacctc agtcttcaga taaagctcat 120 ttgaatcttc gaatcaagtc tgtagagccg gaggactctg ctgtgtatct ctgtgccagc 180 agctaagaac ggcggggact gggggggcac ggcgtttcca acgaaagatt atttttcggt 240 catggaacca agctgtctgt cttggacggc ggaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 30 <211> 366 <212> DNA <213> Artificial <220> <223> TP5 <400> 30 aacacagcga cctcgggtga tcgcaagaag ccgccagagc tcatgtttct ctacaatctt 60 aaacagttga ttcgaaatga gacggtgccc agtcgtttta tacctgaatg cccagacagc 120 tccaagctac ttttacatat atctgccgtg gatccagaag actcagctgt ctatttttgt 180 gccagcagcc aagaatcgca gggacagggg gcatcgcact cctatgaaca gtacttcggt 240 cccggcacca ggctcacggt tttagatcgc agaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 31 <211> 366 <212> DNA <213> Artificial <220> <223> TP6 <400> 31 aacacagcga cctcgggtgt cattaaggaa ttaaagttcc ttattcagca ttatgaaaag 60 gtggagagag acaaaggatt cctacccagc agattctcag tccaacagtt tgatgactat 120 cactctgaaa tgaacatgag tgccttggaa ctggaggact ctgctatgta cttctgtgcc 180 agctctctct cattagggac tgggggggct cattataact atgctgagca gttcttcgga 240 ccagggacac gactcaccgt cctagtcatt agaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 32 <211> 366 <212> DNA <213> Artificial <220> <223> TP7 <400> 32 aacacagcga cctcgggtgt aagggcagga actaaagttc ttcattcagc attatgataa 60 aatggagaga gataaaggaa acctgcccag cagattctca gtccaacagt ttgatgacta 120 tcactctgag atgaacatga gtgccttgga gctagaggac tctgccgtgt acttctgtgc 180 cagctctctc taagggggga cagggggcta agggcaaaca ccgggcagct ctactttggt 240 gaaggctcaa agctgacagt gctggtaagg ggaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 33 <211> 366 <212> DNA <213> Artificial <220> <223> TP8 <400> 33 aacacagcga cctcgggtgg cacgaatggg ctgaggctga tccattactc atatggtgct 60 ggcaaccttc gaataggaga tgtccctgat gggtacaagg ccaccagaac aacgcaagaa 120 gacttcttcc tcctgctgga attggcttct ccctctcaga catctttgta cttctgtgcc 180 agcagtgatg gcacgaggga ctgggggggc gcacgaagtg cagaaacgct gtattttggc 240 tcaggaacca gactgactgt tctcggcacg agaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 34 <211> 366 <212> DNA <213> Artificial <220> <223> TP9 <400> 34 aacacagcga cctcgggtgc tagatatggg ctgaggctga tccattattc atatggtgct 60 ggcagcactg agaaaggaga tatccctgat ggatacaagg cctccagacc aagccaagag 120 aacttctccc tcattctgga gttggctacc ccctctcaga catcagtgta cttctgtgcc 180 agcggtgatg ctagatggga cagggggcct agatgttctg gaaatacgct ctattttgga 240 gaaggaagcc ggctcattgt tgtagctaga tgaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 35 <211> 366 <212> DNA <213> Artificial <220> <223> TP10 <400> 35 aacacagcga cctcgggtgc attggatggg ctgaggctga tccattactc atatgtcgct 60 gacagcacgg agaaaggaga tatccctgat gggtacaagg cctccagacc aagccaagag 120 aatttctctc tcattctgga gttggcttcc ctttctcaga cagctgtata tttctgtgcc 180 agcagtgatg cattggggga ctgggggggc cattggaacc aagacaccca gtactttggg 240 ccaggcactc ggctcctcgt gttagcattg ggaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 36 <211> 366 <212> DNA <213> Artificial <220> <223> TP11 <400> 36 aacacagcga cctcgggtgc atattcaggg gccccagctt ctagtttact ttcgggatga 60 ggctgttata gataattcac agttgccctc ggatcgattt tctgctgtga ggcctaaagg 120 aactaactcc actctcaaga tccagtctgc aaagcagggc gacacagcca cctatctctg 180 tgccagcagt ttctcatatt gggacagggg gccatattct cctatgaaca gtacttcggt 240 cccggcacca ggctcacggt tttagcatat tgaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 37 <211> 366 <212> DNA <213> Artificial <220> <223> TP12 <400> 37 aacacagcga cctcgggtgt tccaggacta gagttgctga gctacttccg cagcaagtct 60 cttatggaag atggtggggc tttcaaggat cgattcaaag ctgagatgct aaattcatcc 120 ttctccactc tgaagattca acctacagaa cccaaggact cagctgtgta tctgtgtgcc 180 agcagtttag cttccagggg acagggggct tccaggaacc aagacaccca gtactttggg 240 ccaggcactc ggctcctcgt gttagttcca ggaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 38 <211> 366 <212> DNA <213> Artificial <220> <223> TP13 <400> 38 aacacagcga cctcgggtgg gttatgggcc tggagttcct gacttacttt cgaaatcaag 60 ctcctataga tgattcaggg atgcccaagg aacgattctc agctcagatg cccaatcagt 120 cgcactcaac tctgaagatc cagagcacgc aaccccagga ctcagcggtg tatctttgtg 180 caagcagctt agaggttatg ggacaggggg cggttatcaa actccgacta caccttcggc 240 tcagggacca ggcttttggt aatagggtta tgaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 39 <211> 366 <212> DNA <213> Artificial <220> <223> TP14 <400> 39 aacacagcga cctcgggtga cacatgcacc aaagcttctt cttttctact atgataagat 60 tttgaacagg gaagctgaca cttttgagaa gttccaatcc agtcggccta acaattcttt 120 ctgctctctc tacattggct ctgcaggcct agagtattct gccatgtacc tctgtgctag 180 cagtagagaa cacatgggac tgggggggca cacatttctg gaaatacgct ctattttgga 240 gaaggaagcc ggctcattgt tgtagacaca tgaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 40 <211> 366 <212> DNA <213> Artificial <220> <223> TP15 <400> 40 aacacagcga cctcgggtga accggaagga ttgagactga tctactattc aataactgaa 60 aacgatcttc aaaaaggcga tctatctgaa ggctatgatg cgtctcgaga gaagaagtca 120 tctttttctc tcactgtgac atctgcccag aagaacgaga tggccgtttt tctctgtgcc 180 agcagtatag aaccggggga cagggggcaa ccggtttcca acgaaagatt atttttcggt 240 catggaacca agctgtctgt cttggaaccg ggaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 41 <211> 366 <212> DNA <213> Artificial <220> <223> TP16 <400> 41 aacacagcga cctcgggtgc ccctttttga actgatagca ctttctactg tgaactcagc 60 aatcaaatat gaacaaaatt ttacccagga aaaatttccc atcagtcatc ccaacttatc 120 cttttcatct atgacagttt taaatgcata tcttgaagac agaggcttat atctctgtgg 180 tgctagggac cccttgggac tgggggggcc ccctttaaca accaggctcc gctttttgga 240 gaggggactc gactctctgt tctagcccct tgaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 42 <211> 366 <212> DNA <213> Artificial <220> <223> TP17 <400> 42 aacacagcga cctcgggtgg agctgtcaaa tttttggttt actttcagaa tgaagacatc 60 atcgacaaaa tagatatgat tggtaaaaac atttcagcaa aatgccctgc taagaaacca 120 tgtaccatag agatccagtc cagcaagcta acagattcag ctgtgtactt ctgtgctagc 180 agtcaatcga gctggggact gggggggcga gctggttctg gaaatacgct ctattttgga 240 gaaggaagcc ggctcattgt tgtaggagct ggaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 43 <211> 366 <212> DNA <213> Artificial <220> <223> TP18 <400> 43 aacacagcga cctcgggtga gcaactaaag ttcctgattt actttcagaa tcaacagcct 60 cttgatcaaa tagacatggt caaggagaga ttctcagctg tgtgcccctc cagctcactc 120 tgcagcctgg gaatcagaac gtgcgaagca gaagactcag cactgtactt gtgctccagc 180 agtcaatcag caacgggact gggggggcag caaccaaaca ccgggcagct ctactttggt 240 gaaggctcaa agctgacagt gctggagcaa cgaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 44 <211> 366 <212> DNA <213> Artificial <220> <223> TP19 <400> 44 aacacagcga cctcgggtgc gcccggaact tacatttttg attagctttc gaaatgaaga 60 aattatggaa caaacagact tggtcaagaa gagattctca gctaagtgtt cctcgaactc 120 acgctgcatc ctggaaatcc tatcctctga agaagacgac tcagcactgt acctctgtgc 180 cagcagtctg tacgcccggg gacagggggc cgcccgagtg cagaaacgct gtattttggc 240 tcaggaacca gactgactgt tctcgcgccc ggaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 45 <211> 366 <212> DNA <213> Artificial <220> <223> TP20 <400> 45 aacacagcga cctcgggtgg tatcaagttt aaatttttga ttaactttca gaatcaagaa 60 gttcttcagc aaatagacat gactgaaaaa cgattctctg ctgagtgtcc ttcaaactca 120 ccttgcagcc tagaaattca gtcctctgag gcaggagact cagcactgta cctctgtgcc 180 agcagtctgt cgtatcaggg acagggggcg tatcagagtg cagaaacgct gtattttggc 240 tcaggaacca gactgactgt tctcggtatc agaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 46 <211> 366 <212> DNA <213> Artificial <220> <223> TP21 <400> 46 aacacagcga cctcgggtgc gctaactggg gctacagctg atttatatct catacgatgt 60 tgatagtaac agcgaaggag acatccctaa aggatacagg gtctcacgga agaagcggga 120 gcatttctcc ctgattctgg attctgctaa aacaaaccag acatctgtgt acttctgtgc 180 tagcagttta tccgctaagg gacagggggc cgctaaaacc aagacaccca gtactttggg 240 ccaggcactc ggctcctcgt gttagcgcta agaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 47 <211> 366 <212> DNA <213> Artificial <220> <223> TP22 <400> 47 aacacagcga cctcgggtgc cttcaagctt gatgctcatg gcaactgcaa atgaaggctc 60 tgaagccaca tacgagagtg gattcaccaa ggacaagttt ccaatcagcc ggccaaacct 120 aacattctca acgttgacag tgaacaatgc aaggcctgga gacagcagta tctatttctg 180 tagttctaga gaccttcagg gactgggggg gcccttcact cctatgaaca gtacttcggt 240 cccggcacca ggctcacggt tttagccttc agaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 48 <211> 366 <212> DNA <213> Artificial <220> <223> TP23 <400> 48 aacacagcga cctcgggtgg gaatacacag gaggcaccct ccagcaactc ttctactcta 60 ttactgttgg ccaggtagag tcggtggtgc aactgaacct ctcagcttcc aggccgaagg 120 acgaccaatt catcctaagc acggagaagc tgcttctcag ccactctggc ttctacctct 180 gtgcctggag tctggaatag ggacaggggg cggaatacaa acacagaagt cttctttggt 240 aaaggaacca gactcacagt tgtagggaat agaggatctg agaaatgtga ctccacccaa 300 ggtctccttg tttgagccat caaaagcaga gattgcaaac aaacaaaagg ctaccctcgt 360 gtgctt 366 <210> 49 <211> 30 <212> DNA <213> Artificial <220> <223> Adapter sequence 1 <400> 49 ccatctcatc cctgcgtgtc tccgactcag 30 <210> 50 <211> 23 <212> DNA <213> Artificial <220> <223> Adapter sequence 2 <400> 50 cctctctatg ggcagtcggt gat 23
Claims
1. A method for correcting and normalizing TCR β high-throughput sequencing data based on a template sequence and a reference cell, characterized in that, It includes the following steps: (a) Incorporate a fixed number of exogenous reference cells and synthetic templates into the sample, construct a high-throughput sequencing library of TCR β using multiplex PCR primers, and perform sequencing using a high-throughput sequencing platform; the exogenous reference cells are 2B4 hybridoma cells, the template sequences are as shown in SEQ ID NO.26-48, and the multiplex PCR primer sequences are as shown in SEQ ID NO.3-25; (b) Use molecular barcodes to count the number of sequencing reads of template sequences containing different V genes, investigate the amplification bias law of template sequences after being mixed into the sample using the number of templates, and calculate the amplification bias index according to the following formula: i = 1…23, n = 23, Count(V i ) is the number of the template sequence V obtained by sequencing i ; if N(s) is the frequency of the CDR3 sequence s, and V i is the V gene type of s, then its corrected frequency N'(s) = N(s) × ABI(V i ); (c) Construct a substitution matrix using the Dayhoff method as a parameter for pairwise sequence alignment, calculate the similarity score between the complementarity-determining region 3 sequences of TCR β, determine the similarity threshold between the original sequence and the error sequence, and merge low-frequency error sequences into high-frequency sequences based on this threshold to correct sequencing errors. The steps are as follows: 1) Arrange the CDR3 sequences in descending order of frequency, 2) Take the CDR3 sequence with the highest frequency as the clustering center, 3) For the remaining unclustered CDR3 sequences, merge the CDR3 sequences with similarity scores greater than the similarity threshold, frequency ratios less than the ratio threshold, and the same V and J genes into this cluster; 4) Repeat 2) and 3) until all CDR3 sequences have been clustered; (d) Standardize the sample sequencing data using exogenous reference cells: Assume the number of exogenous reference cells added is n, the number of reads measured is m, and the number of reads of a certain CDR3 is k. After standardization, the number of cells p corresponding to this CDR3 is ; (e) Perform precise quantification of TCR β in the sample.
2. The method according to claim 1, characterized in that: Step (a) includes the following steps: (1) Lyse the sample using Trizol and add the lysate of a fixed number of exogenous reference cells to the lysed sample; (2) Extract the total RNA of the sample and exogenous reference cells; (3) Perform reverse transcription using the C-terminal specific primer of TCR β; (4) Add a fixed number of template sequences to the reverse-transcribed sample; (5) Construct a high-throughput sequencing library of TCR β using a set of multiplex PCR primers with optimized sequence composition and usage concentration; (6) Perform sequencing using a high-throughput sequencing platform.
3. The method according to claim 2, wherein: In step (2), the method for total RNA extraction is the Trizol method; and / or, in step (3), the C-terminal specific primer of TCR β is TRBC, and its sequence is as shown in SEQ ID NO.1; In step (5), the SEQ ID NO.3-25 sequences are added with high-throughput sequencing adapters, and the reverse sequences are as shown in SEQ ID NO.
2.
4. The method according to claim 3, characterized in that: In step (3), the steps for performing reverse transcription using the C-terminal specific primer of TCR β are as follows: ① Take 0.1 ug of the RNA from step (2), 1 ul of 10 uM primer TRBC, and the rest is water, prepare a 12 ul reaction system, then incubate at 72 °C for 3 min in a PCR instrument, and quickly place on ice for 5 min; ② Prepare a 20-μl reaction system by mixing the product obtained in step ①, 4 μl of 5X first strand buffer, 2 μl of dNTPs, 1 μl of RNase inhibitor, and 1 μl of RevertAid reverse transcriptase, and then incubate in a PCR instrument at 42 °C for 60 min and at 70 °C for 10 min.
5. The method according to claim 3, characterized in that: In step (5), the sequences of the multiplex PCR primers are as shown in SEQ ID NO. 3-25, and the SEQ ID NO. 3-25 sequences are added with high-throughput sequencing adapters, and the reverse sequences are as shown in SEQ ID NO. 2; Among them, the forward primer FW-primer mix consists of the sequences shown in SEQ ID NO. 3-25, and the proportions of the sequences shown in SEQ ID NO. 3-25 are 1:2:6:6:2:2:6:2:6:6:1:2:2:6:6:6:6:1:1:2:1:2:2 in sequence.
6. The method according to claim 2, wherein: In step (6), the high-throughput sequencing platform used is the Ion PGM platform.
Citation Information
Patent Citations
Multi-PCR (Polymerase Chain Reaction) primer and method for constructing mouse TCRB (T-Cell Receptor Beta) library based on high-throughput sequencing
CN104531698A
Preparation of gene sequencing calibration reference strain
CN108300700A
Methods using randomer-containing synthetic molecules
US20170292149A1