A method and kit for sequencing human immune repertoire

Through specific primer design and UDG/UNG enzyme digestion, it is ensured that each template nucleic acid corresponds to only one UMI, which solves the problems of RNA degradation, cumbersome operation and amplification errors in human immune repertoire sequencing, and realizes efficient and accurate TCR CDR3 region sequencing.

CN114774517BActive Publication Date: 2025-09-09SHENZHEN HAPLOX BIOTECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210381164.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-12
Publication Date
2025-09-09
Estimated Expiration
2042-04-12

AI Technical Summary

Technical Problem

Existing human immune repertoire sequencing technologies have problems such as easy degradation of RNA samples, cumbersome experimental operations, high difficulty in multiplex PCR amplification, incomplete coverage of amplification products, high amplification error rate, and false positives.

Method used

Specific primers were designed and the first, second, and third primers were added sequentially, combined with UDG/UNG enzyme digestion, to ensure that each template nucleic acid corresponded to only one unique identifier (UMI). PCR amplification and enrichment were performed to cover the CDR3 region of the human T cell receptor β chain, and the library was constructed using the Illumina high-throughput sequencing platform.

Benefits of technology

It improves capture efficiency, reduces experimental steps, reduces DNA loss, corrects amplification errors, realizes quantitative detection of target gene copy number, and ensures the accuracy and authenticity of sequencing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003591926050000101
    Figure BDA0003591926050000101
  • Figure BDA0003591926050000111
    Figure BDA0003591926050000111
Patent Text Reader

Abstract

The present application discloses a method and kit for sequencing the human immune repertoire. The method of the present application includes extending the template nucleic acid once with a first primer; the first primer includes a platform upstream primer binding region, a UMI and a target upstream primer sequence, and part of its T is replaced by dU; a second primer is added to the reaction system for one extension; the second primer includes a platform downstream primer binding region and a target downstream primer sequence; UDG / UNG enzyme is added to the reaction system to digest dU; after digestion, a third primer is added to the reaction system for PCR; the third primer is the platform upstream primer binding region sequence of the first primer; the target sequence of the first and second primers is a specific sequence designed for the human TCR gene. The method of the present application fully covers the CDR3 of the TCR of the human immune repertoire and has high capture efficiency; through the special design of the primers, three primers are added to the reaction system in sequence to achieve each original template nucleic acid corresponding to only one UMI, correct the amplification deviation of each target point, and realize quantitative detection of the target gene copy number.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of human immune repertoire sequencing, and in particular to a method and kit for human immune repertoire sequencing. Background Art

[0002] The human immune repertoire refers to the sum of all T and B lymphocyte clones with varying specificities circulating in an individual at any given time. The process by which T cells recognize and activate antigens involves: antigens are ingested by antigen-presenting cells (such as macrophages), processed, and transformed into peptide fragments. These fragments are then presented on the cell surface as antigen-peptide-MHC complexes. These fragments are then recognized by the T cell receptor (TCR) on the T cell surface, thereby activating the body's immune response. A richer immune repertoire indicates a more effective defense against pathogens such as bacteria and viruses; conversely, a lower repertoire increases susceptibility to disease. TCRs are heterodimers composed of two distinct peptide chains. Most TCRs, approximately 95%, are composed of α and β chains, while a smaller minority, approximately 5%, consist of γ and δ chains. The V regions of the α and β chains (Vα and Vβ) each contain three hypervariable regions: CDR1, CDR2, and CDR3. CDR3 is the most variable and directly determines the antigen-binding specificity of the TCR. The CDR3 of the TCR is encoded by three genes: V, D, and J. During lymphocyte maturation, these genes rearrange to form a variety of recombinant sequence segments. This explains why the human genome, as revealed by proteomics, encodes a nearly infinite number of proteins despite a limited number of genes. The CDR3 of the TCR is the primary site for antigen recognition by T cells. In a narrow sense, TCR CDR3 sequence analysis represents a representative area of ​​T cell repertoire research.

[0003] Generally, human immune repertoire sequencing (IR-seq) refers to the study of T / B lymphocytes, using 5'RACE technology or multiplex PCR to amplify the complementarity-determining regions (CDR regions) that determine the diversity of B cell receptors (BCR) or T cell receptors (TCR), and then combining high-throughput sequencing technology to comprehensively evaluate the diversity of the immune system and deeply explore the relationship between the immune repertoire and disease.

[0004] However, human immune repertoire sequencing technology itself still has some defects. 5'RACE technology is an RNA-based amplification technology. Its technical principle is that the universal primer sequence of the constant region of TCR / BCR is amplified to the variable region, and then the introduced linker sequence is subjected to a second non-biased PCR amplification. Its disadvantages are: (1) RNA samples are used as amplification templates, and RNA has more stringent extraction conditions than DNA, which has higher requirements for both the experimental environment and technicians; (2) RNA has the defect of being easily degraded, and strict conditions are imposed on the transportation and storage of original tissue samples; (3) The experimental operation is more cumbersome and complicated.

[0005] Although multiplex PCR amplification uses DNA as the original template and can be directly used for genomic DNA by designing specific primers based on the TCR V, D, and J regions, its defects are also quite obvious: (1) The design of multiplex PCR primers is difficult, and the amplified products are difficult to cover all subtypes; (2) The operation process of the existing technology is complicated. After the multiplex PCR amplification is completed and purified, end repair, A base addition and end modification are required to build a library, which is relatively cumbersome; (3) Multiplex PCR amplification will involve V, D, and J region fragments that have not undergone rearrangement or are undergoing rearrangement, resulting in PCR amplification products ranging from tens to thousands of bp in length. In this step, the fragments need to be sorted to retain smaller effective fragments. Usually, fragment sorting is performed by gel electrophoresis and gel recovery, which takes a lot of experimental time and causes a certain loss of effective fragments; the operation is also relatively cumbersome; (4) Whether the amplification of TCRβ chain CDR3 uses DNA or RNA as the template, the multiplex PCR specific amplification stage cannot avoid biased amplification. A large number of amplification cycles will also accumulate a large number of amplification errors, which is not conducive to diversity analysis and easily leads to misinterpretation.

[0006] Therefore, how to reduce or correct false positives caused by PCR amplification errors or human factors remains the research focus of human immune repertoire sequencing. Summary of the Invention

[0007] The purpose of this application is to provide an improved method and kit for sequencing the human immune repertoire.

[0008] In order to achieve the above objectives, this application adopts the following technical solutions:

[0009] One aspect of the present application discloses a method for sequencing a human immune repertoire, comprising the following steps:

[0010] A reaction system is prepared, wherein a first primer is used to extend the template nucleic acid once to obtain a complementary strand; the first primer includes, from the 5' end to the 3' end, a sequencing platform upstream primer binding region, a unique identifier, and a target-specific upstream primer sequence; furthermore, in the first primer, the base T in the sequencing platform upstream primer binding region and the target-specific upstream primer sequence is replaced with deoxyuracil, and the sequencing platform upstream primer binding region corresponds to the 3' end of the upstream sequencing primer of the sequencing platform;

[0011] After the first primer is extended, a second primer is added to the reaction system, and the complementary strand of the first primer extension is extended once using the second primer to obtain a product consisting of a sequencing platform upstream primer binding region, a unique identifier, a target sequence, and a sequencing platform downstream primer binding region; the second primer includes, from the 5' end to the 3' end, the sequencing platform downstream primer binding region and a target-specific downstream primer sequence, and the sequencing platform downstream primer binding region corresponds to the 3' end of the sequencing platform downstream sequencing primer;

[0012] After the second primer extension is completed, UDG / UNG enzyme is added to the reaction system to digest deoxyuracil, thereby digesting the first primer and the extended chain of the first primer;

[0013] After UDG / UNG enzyme digestion is completed, a third primer is added to the reaction system, and the product extended by the second primer is PCR amplified and enriched using the third primer and the second primer to obtain a product in which all amplicons of the template nucleic acid are added with the same unique identifier; the third primer is the entire or partial sequence of the sequencing platform upstream primer binding region of the first primer from the 5' end, and the base T in the third primer is not replaced by deoxyuracil;

[0014] All amplicons obtained by PCR amplification and enrichment are added with the same unique identifier for sequencing library construction and sequencing, thus completing the sequencing of the human immune repertoire;

[0015] Among them, the target-specific upstream primer sequence and the target-specific downstream primer sequence are specific primer sequences designed for the human T cell receptor encoding gene and fully cover the CDR3 region encoding gene sequence thereof.

[0016] In the methods of this application, "performing a single extension" means that after a primer anneals and hybridizes to the target sequence, only that primer is extended, without further denaturation or annealing. This ensures that each template nucleic acid strand is labeled with a unique UMI. Of course, after the second primer is added, although the second primer is designed to anneal, hybridize, and extend, the first primer will also anneal, hybridize, and extend. However, since the second primer anneals, hybridizes, and extends only once, the product consisting of the upstream primer binding region of the sequencing platform, the unique identifier, the target sequence, and the downstream primer binding region of the sequencing platform will also contain only the strand labeled with the UMI initially extended by the first primer. Finally, during exponential PCR amplification and enrichment with the third and second primers, only the amplicons of the strand labeled with the UMI initially extended by the first primer are exponentially enriched. Furthermore, before PCR amplification and enrichment with the third and second primers, the first primer is removed using UDG / UNG enzyme digestion, preventing the first primer from introducing new UMIs in the next round of PCR amplification. The template nucleic acid of the present application can be DNA or cDNA.

[0017] It should be noted that the method of this application, through the special design of the first, second, and third primers, and the sequential addition of the three primers to the reaction system, can achieve the addition of the same UMI to all amplicon strands of a given template nucleic acid parent strand. This is particularly important for mutation detection. For example, the UMI can be directly used to determine which amplicon strands amplified by the same specific primers are derived from mutations or non-mutations, thereby quantitatively detecting mutations and obtaining an accurate mutation rate.

[0018] It should also be noted that the target-specific upstream primer sequence and target-specific downstream primer sequence of the present application are specific primer sequences designed for the human T cell receptor encoding gene to fully cover its CDR3 region encoding gene sequence. They can be quickly captured and amplified based on human genomic DNA or cDNA, covering the CDR3 functional region of the human immune repertoire TCR, with high capture efficiency; and, all amplicons obtained by PCR amplification and enrichment are directly added with the same unique identifier for sequencing library construction and sequencing, that is, the human immune repertoire sequencing is completed, which reduces the experimental steps, reduces DNA loss, and improves the authenticity of the final result; The human immune repertoire sequencing method of the present application can efficiently integrate the immune repertoire with the library construction process of the Illumina high-throughput sequencing platform. Using the primers designed in the present application, a two-step PCR rapid library construction can be achieved, and the entire experimental time can be shortened to less than 4 hours. In addition, in the method of the present application, UMIs are added to the immune repertoire sequencing through specially designed primers, so that all amplicon chains of a certain template nucleic acid mother chain are added with the same UMI, which can correct the amplification deviation of each target site and correct the amplification errors of each target site, and realize the quantitative detection of the target gene copy number, thereby solving the false positive problem caused by PCR amplification errors or human factors.

[0019] It should also be noted that the key to this application lies in the design of the primer structure and the order in which the primers are added, so that the final amplified and enriched amplicons all carry the same UMI. The specific primer sequences can be determined based on the target sequence and the specific sequencing platform. For example, conventional primer design software can be used to design the target-specific upstream primer sequence of the first primer for the specific target sequence, and the sequencing platform upstream primer binding region of the first primer can be designed for the intended sequencing platform, thereby forming the first primer.

[0020] In one implementation of the present application, the target-specific upstream primer sequence is a specific primer sequence designed for the V gene of the human T cell receptor β chain; the target-specific downstream primer sequence is a specific primer sequence designed for the J gene of the human T cell receptor β chain; and the CDR3 region encoding gene of the T cell receptor β chain can be fully covered by amplification of the target-specific upstream primer sequence and the target-specific downstream primer sequence.

[0021] It can be understood that the CDR3 of the T cell receptor β chain is encoded by three genes: V, D, and J; therefore, the upstream primer is designed for the V gene and the downstream primer is designed for the J gene, which can better ensure full coverage of the CDR3 region encoding genes.

[0022] In one implementation of the present application, in the first primer, at least one deoxyuracil is inserted into the sequence of the unique identifier, and the number of consecutive bases of the unique identifier is made less than 5 by the insertion of the deoxyuracil.

[0023] It should be noted that the purpose of inserting deoxyuracil or replacing T with deoxyuracil in primers is to allow for digestion of the primers with UDG / UNG enzymes when not needed. Inserting deoxyuracil in the unique identifier also minimizes the potential for nonspecific amplification of random UMIs during subsequent amplification. The unique identifier sequence can optionally include one or more fixed deoxyuracils inserted into the middle of its base sequence, with the number of consecutive N bases on either side of the deoxyuracil less than 5 nt. This effectively prevents nonspecific amplification. Of course, if the possibility of nonspecific amplification is not a concern, deoxyuracil can be omitted from the unique identifier.

[0024] In one implementation of the present application, the number of amplification cycles of PCR amplification enrichment is greater than or equal to 5.

[0025] It should be noted that the PCR amplification and enrichment using the third and second primers is primarily intended to exponentially enrich the amplicons of the parent strand initially extended by the first primer and labeled with the UMI, thereby obtaining more amplicon strands derived from the same template nucleic acid and with the same UMI, facilitating subsequent library construction and sequencing.

[0026] In one implementation of the present application, the first primer consists of 40 primers having sequences shown from Seq ID No. 1 to Seq ID No. 40.

[0027] It should be noted that the 40 primers of the sequences shown in Seq ID No. 1 to Seq ID No. 40 are specific upstream primers designed independently by this application specifically for the V gene reference sequence of the human T lymphocyte receptor β chain, based on base complementary pairing and primer design principles, combined with the required sites; the 40 primers designed in this application include 42 V functional regions; when used, the 40 primers can be mixed in equal proportions according to molar mass. It is understood that the 40 primers in this application are only upstream specific primers that have been confirmed to be specifically usable in one implementation of this application. Under the inventive concept of this application, several bases can be added or subtracted based on the 40 primers in this application, or the primer sequences can be redesigned according to the primer design principles.

[0028] In one implementation of the present application, the second primer consists of 12 primers having sequences shown in Seq ID No. 41 to Seq ID No. 52.

[0029] It should be noted that the 12 primers of the sequences shown in Seq ID No. 41 to Seq ID No. 52 are specific downstream primers designed independently by this application specifically for the J gene reference sequence of the human T lymphocyte receptor β chain, based on base complementary pairing and primer design principles, in combination with the required sites; the 12 primers of this application include 6 J1 functional regions and 7 J2 functional regions; when used, the 12 primers can be mixed in equal proportions according to molar mass. It is understood that the 12 primers of this application are only specific downstream specific primers that have been confirmed to be useful in one implementation of this application. Under the inventive concept of this application, several bases can be added or subtracted based on the 12 primers of this application, or the primer sequences can be redesigned according to the primer design principles.

[0030] In one implementation of the present application, the third primer is the sequence shown in Seq ID No.53.

[0031] It should be noted that the third primer in this application is actually the entire or partial sequence of the upstream primer binding region of the sequencing platform starting from the 5' end of the first primer, that is, a primer designed for the 3' end of the sequencing platform primer sequence. For example, the primer designed with reference to the 19 nt sequence 3' to the P5 end of the Illumina NovaSeq6000 sequencing platform is the sequence shown in Seq ID No. 53.

[0032] In one implementation of the present application, sequencing library construction includes the following steps:

[0033] All amplicons obtained by PCR amplification enrichment are added with the same unique identifier and purified to obtain a purified product; the purified product is amplified using the fourth primer and the fifth primer to obtain a sequencing library; the fourth primer is an upstream sequencing primer of the sequencing platform with a sequencing adapter and a barcode, and the fifth primer is a downstream sequencing primer of the sequencing platform with a sequencing adapter and a barcode.

[0034] It should be noted that the sequencing library construction method of the present application is actually to amplify and construct the library of the PCR amplification enrichment products of the third primer and the second primer; that is, the target product is amplified and enriched again using the upstream sequencing primer of the sequencing platform and the downstream sequencing primer of the sequencing platform.

[0035] In one implementation of the present application, the purification is at least one of magnetic bead purification, column purification, and gel purification.

[0036] In one implementation of the present application, the fourth primer is the sequence shown in Seq ID No.54.

[0037] In one implementation of the present application, the fifth primer is the sequence shown in Seq ID No.55.

[0038] It should be noted that in the fourth primer of the sequence shown in Seq ID No. 54 and the fifth primer of the sequence shown in Seq ID No. 55, "NNNNNN" refers to an index with a length of 6-10 nt, i.e., a barcode. For example, in one implementation of the present application, the "NNNNNN" of the fourth primer of the sequence shown in Seq ID No. 54 is specifically "TGCGTAAT", and the "NNNNNN" of the fifth primer of the sequence shown in Seq ID No. 55 is specifically "CCTAACCT".

[0039] Another aspect of the present application discloses a kit for sequencing a human immune repertoire, comprising a first primer, a second primer, a third primer and a UDG / UNG enzyme; the first primer comprises, from the 5' end to the 3' end, a sequencing platform upstream primer binding region, a unique identifier and a target-specific upstream primer sequence; and, in the first primer, the base T in the sequencing platform upstream primer binding region and the target-specific upstream primer sequence is replaced with deoxyuracil, and the sequencing platform upstream primer binding region corresponds to the 3' end of the upstream sequencing primer of the sequencing platform; the second primer comprises, from the 5' end to the 3' end, a sequencing platform upstream primer binding region, a unique identifier and a target-specific upstream primer sequence; and The first primer comprises a sequencing platform downstream primer binding region and a target-specific downstream primer sequence in sequence, and the sequencing platform downstream primer binding region corresponds to the 3' end of the downstream sequencing primer of the sequencing platform; the third primer is the entire or partial sequence of the sequencing platform upstream primer binding region from the 5' end of the first primer, and the base T in the third primer is not replaced by deoxyuracil; the target-specific upstream primer sequence and the target-specific downstream primer sequence are specific primer sequences designed for human T cell receptor encoding genes and fully cover the CDR3 region encoding gene sequence thereof.

[0040] It should be noted that the human immune repertoire sequencing kit of the present application is actually a kit that assembles the first primer, second primer, third primer and UDG / UNG enzyme used in the method of immune repertoire sequencing of the present applicant into a kit to facilitate the implementation of the human immune repertoire sequencing method of the present application. Therefore, the definition of the first primer, second primer and third primer in the kit can refer to the human immune repertoire sequencing method of the present application. For example, the target-specific upstream primer sequence is a specific primer sequence designed for the V gene of the human T cell receptor β chain; the target-specific downstream primer sequence is a specific primer sequence designed for the J gene of the human T cell receptor β chain; the CDR3 region encoding gene of the T cell receptor β chain can be fully covered by amplification of the target-specific upstream primer sequence and the target-specific downstream primer sequence. For another example, in the first primer, at least one deoxyuracil is inserted into the sequence of the unique identifier, and the number of consecutive bases of the unique identifier is less than 5 by inserting the separation of deoxyuracil.

[0041] It should also be noted that one of the key aspects of this application lies in the design of the primer structure. The specific primer sequence can be determined based on the target sequence and the specific sequencing platform. For example, conventional primer design software can be used to design the target-specific upstream primer sequence of the first primer for the specific target sequence, and the sequencing platform upstream primer binding region of the first primer can be designed for the intended sequencing platform, thereby forming the first primer.

[0042] In one implementation of the present application, the first primer in the kit consists of 40 primers with sequences shown in Seq ID No. 1 to Seq ID No. 40.

[0043] In one implementation of the present application, the second primer in the kit consists of 12 primers with sequences shown in Seq ID No. 41 to Seq ID No. 52.

[0044] In one implementation of the present application, the third primer in the kit is the sequence shown in Seq ID No.53.

[0045] In one implementation of the present application, the kit of the present application also includes a fourth primer and a fifth primer; the fourth primer is an upstream sequencing primer of the sequencing platform with a sequencing adapter and a barcode, and the fifth primer is a downstream sequencing primer of the sequencing platform with a sequencing adapter and a barcode.

[0046] It should be noted that the fourth primer and the fifth primer of the present application are actually designed for the upstream sequencing primer and the downstream sequencing primer of the sequencing platform, or the upstream and downstream primers of the sequencing platform can be directly used, as long as the sequencing adapter and barcode they carry are consistent with the present application. Therefore, the fourth primer and the fifth primer can be selectively added to the test kit as needed. Of course, for ease of use, the fourth primer and the fifth primer are included in the test kit of the present application.

[0047] In one implementation of the present application, the fourth primer in the kit is the sequence shown in Seq ID No.54.

[0048] In one implementation of the present application, the fifth primer in the kit is the sequence shown in Seq ID No.55.

[0049] In one implementation of the present application, the kit of the present application further includes a PCR amplification reagent.

[0050] It is understood that the PCR amplification reagents can be assembled into the kit of the present application according to needs, or conventionally used PCR amplification reagents, such as PCR reaction buffer, enzyme, etc., can be purchased separately.

[0051] Due to the adoption of the above technical solution, the beneficial effects of this application are:

[0052] The human immune repertoire sequencing method and kit of the present application designs specific primers for the human T cell receptor, fully covering the CDR3 region of the human immune repertoire TCR, with high capture efficiency. In addition, by specially designing the first primer, the second primer, and the third primer, the three primers are added to the reaction system in sequence, so that each original template nucleic acid corresponds to only one UMI label, thereby correcting the amplification deviation of each target point, correcting PCR amplification errors, and correcting amplification errors introduced artificially during the library construction process. The method and kit of the present application can label each original template nucleic acid, thereby achieving quantitative detection of the copy number of the target gene. DETAILED DESCRIPTION

[0053] The human immune repertoire sequencing method of the present application is improved based on the multiplex PCR amplification method of IR-seq, and amplification is performed using DNA or cDNA as a template. For example, by carefully designing the V region and J region primers upstream and downstream of the CDR3 of the TCRβ chain, the human TCR CDR3 region is specifically captured and enriched, covering 42 functional regions, 2 D functional regions, 6 J1 functional regions, and 7 J2 functional regions of the CDR3 of the human TCRβ chain. On this basis, the library construction process is highly optimized, reducing the conventional tedious experimental operations, and using the amplicon library construction method. At the same time, when designing the second-step amplification primers, a double-ended Index tag is used to design the amplification primers to ensure the accuracy of the data. Most importantly, the present application creatively uses specially designed first primers, second primers, and third primers, and sequentially adds three primers to the reaction system to insert a unique UMI tag into the complementary chain of each parent chain molecule, so that each original template nucleic acid corresponds to only one UMI, corrects the amplification deviation of each target point, and realizes quantitative detection of the target gene copy number.

[0054] The method for sequencing the human immune repertoire of the present application comprises the following steps:

[0055] A reaction system is prepared, wherein a first primer is used to extend the template nucleic acid once to obtain a complementary strand; the first primer includes, from the 5' end to the 3' end, a sequencing platform upstream primer binding region, a unique identifier, and a target-specific upstream primer sequence; furthermore, in the first primer, the base T in the sequencing platform upstream primer binding region and the target-specific upstream primer sequence is replaced with deoxyuracil, and the sequencing platform upstream primer binding region corresponds to the 3' end of the upstream sequencing primer of the sequencing platform;

[0056] After the first primer is extended, a second primer is added to the reaction system, and the complementary strand of the first primer extension is extended once using the second primer to obtain a product consisting of a sequencing platform upstream primer binding region, a unique identifier, a target sequence, and a sequencing platform downstream primer binding region; the second primer includes, from the 5' end to the 3' end, the sequencing platform downstream primer binding region and a target-specific downstream primer sequence, and the sequencing platform downstream primer binding region corresponds to the 3' end of the sequencing platform downstream sequencing primer;

[0057] After the second primer is extended, UDG / UNG enzyme is added to the reaction system to digest deoxyuracil, thereby digesting the first primer and the extended chain of the first primer;

[0058] After UDG / UNG enzyme digestion is completed, a third primer is added to the reaction system, and the product extended by the second primer is PCR amplified and enriched using the third primer and the second primer to obtain a product in which all amplicons of the template nucleic acid are added with the same unique identifier; the third primer is the entire or partial sequence of the sequencing platform upstream primer binding region of the first primer from the 5' end, and the base T in the third primer is not replaced by deoxyuracil;

[0059] The same unique identifier is added to all amplicons obtained by PCR amplification and enrichment, and the sequencing library is constructed and sequenced, thus completing the sequencing of the human immune repertoire;

[0060] The target-specific upstream primer sequence and the target-specific downstream primer sequence are specific primer sequences designed for human T cell receptor encoding genes and fully cover the CDR3 region encoding gene sequence thereof.

[0061] The human immune repertoire sequencing method of this application generally works as follows: A specific UMI sequence is designed, for example, by targeting the human T cell receptor β chain target region sequence, extending the complementary strand and simultaneously incorporating the UMI. Unused UMI sequences in the UDG / UNG enzyme digestion system are then used to ensure unique UMIs for each extended strand of the template molecule. This complementary strand is then used to amplify and enrich the targeted region, corresponding sequencing primers are designed, and barcodes / indexes and sequencing adapters are added to the enriched product to complete library construction.

[0062] The human immune repertoire sequencing method presented in this application is highly versatile and suitable for TCR sequencing using DNA as a starting template or multiplex PCR amplification using cDNA synthesized from RNA as a starting template. It exhibits high specificity and sensitivity. Enzymatic digestion of redundant UMI tags ensures that molecular chains with the same UMI originate from the same template, resulting in unique molecular tags. This can correct for false positives caused by PCR errors or human factors, as well as biases caused by uneven amplification efficiency in multiplex PCR, ensuring data authenticity.

[0063] The human immune repertoire sequencing method of this application has the following advantages:

[0064] 1. Independently design relevant primers, which can quickly capture and amplify based on human genomic DNA or cDNA, for example, covering the CDR3 functional region of the TCRβ chain of the human immune repertoire with high capture efficiency.

[0065] 2. Currently, some TCR products on the market use agarose gel electrophoresis to separate and purify the capture region. This operation is cumbersome and prone to errors, and electrophoresis purification is prone to nucleic acid loss, resulting in distorted experimental results. In one implementation of the present application, all amplicons obtained by PCR amplification and enrichment are directly purified, and sequencing libraries are constructed and sequenced. For example, a specific ratio of AMPure magnetic beads is used to perform DNA purification and fragment sorting in one step, reducing experimental steps, reducing DNA loss, and improving the authenticity of the final results.

[0066] 3. In terms of methods, this application efficiently integrates the human immune repertoire and the library construction process of the Illumina high-throughput sequencing platform, independently designs relevant primers, and can realize rapid two-step PCR library construction; the experimental time can be shortened to less than 4 hours.

[0067] 4. Specially designed primers and the addition of UMIs to immune library sequencing can correct the amplification deviation of each target site, correct the amplification errors of each target site, and achieve quantitative detection of the target gene copy number.

[0068] In the present application method, the design ideas of the first primer, the second primer, the third primer, the fourth primer and the fifth primer are as follows:

[0069] The first primer is a UMI sequence consisting of three parts: the first part is a fixed sequence of 15-25 nt, corresponding to the 3' end of the upstream sequencing primer on the sequencing platform; the second part is a random 6-8 N-base sequence, the UMI; and the third part is the target-specific upstream primer sequence. The sequence connection order is: 5'-first part-second part-third part-3'.

[0070] In the first and third sequences, dU (deoxyuridine) bases are used to replace T (thymine) bases. The first 15-25nt fixed sequence can refer to the complete 3' end sequence of the upstream adapter of different sequencing platforms. For example, the 19nt sequence at the 3' end of P5 of the Illumina NovaSeq6000 sequencing platform is designed as the sequence shown in Seq ID No.56.

[0071] Seq ID No. 56: 5'-CACGACGCUCUUCCGAUCU-3'.

[0072] Furthermore, to minimize the possibility of nonspecific amplification of random UMIs during subsequent amplification, one or more fixed deoxyuracils can be inserted into the middle of the random N-base sequence at positions 6-8 in the second sequence. The number of consecutive N-bases on either side of the deoxyuracil should be less than 5 nt. Of course, if the possibility of nonspecific amplification is not a concern, deoxyuracils can be omitted from the second sequence. For example, the second sequence can be 5'-NNNNUNNNN-3', 5'-NNNUNNNUNN-3', or 5'-NNNNNNNN-3'.

[0073] Furthermore, the third part of the sequence is a target-specific upstream primer sequence. The target gene can be searched according to authoritative databases such as NCBI, and the upstream primer can be independently designed in combination with the desired site according to the base complementary pairing and primer design principles. A plurality of PCR primers with strong specificity can be designed according to multiple targets and used in combination. For example, in one implementation of the present application, the V gene reference sequence of the human T lymphocyte receptor β chain is searched with reference to the NCBI and IMGT standard databases; the upstream primer is independently designed in combination with the desired site according to the base complementary pairing and primer design principles; a total of 40 target-specific upstream primer sequences are designed in the present application, including a total of 42 V functional regions; all functional region primers are mixed and used in equal proportions according to molar mass.

[0074] The second primer consists of two parts: the first part is the downstream specific primer sequence, and the second part is a fixed sequence of 15-25 nt corresponding to the 3' end complementary sequence of the downstream sequencing primer on the sequencing platform. The sequence connection order is: 5'-second part-first part-3'.

[0075] Furthermore, the first part of the sequence of the second primer is a target-specific downstream primer sequence. The target gene can be searched according to authoritative databases such as NCBI, and the downstream primer can be independently designed in combination with the desired site according to the base complementary pairing and primer design principles. A plurality of PCR primers with strong specificity can be designed according to multiple targets and used in combination. For example, in one implementation of the present application, the J gene reference sequence of the human T lymphocyte receptor β chain is searched with reference to the NCBI and IMGT standard databases; the downstream primer is independently designed in combination with the desired site according to the base complementary pairing and primer design principles; a total of 12 target-specific upstream primer sequences are designed in the present application, including 6 J1 functional regions and 7 J2 functional regions; a total of 13 functional region primers are mixed and used in equal proportions according to molar mass.

[0076] The second part of the second primer, a fixed sequence of 15-25 nt, can refer to the complete 3' end sequence of the downstream adapter of different sequencing platforms. For example, the complementary sequence of the 21 nt sequence at the 3' end of P7 of the Illumina NovaSeq6000 sequencing platform can be designed as the sequence shown in Seq ID No. 57.

[0077] Seq ID No. 57: 5'-AGACTGTGCTCTTCCGATCT-3'.

[0078] The third primer is the same as the fixed sequence of 15-25 nt in the first primer sequence. It should be noted that the T base in this sequence cannot be replaced by deoxyuracil. For example, the 19 nt sequence at the 3' end of P5 on the Illumina NovaSeq6000 sequencing platform is designed as the sequence shown in Seq ID No.53.

[0079] Seq ID No. 53: 5'-CACGACGCTCTCCGATCT-3'.

[0080] The fourth primer sequence is: the complete upstream sequencing adapter sequence with a barcode. The index length can be 6-10 nt. For example, the P5 sequence of the Illumina NovaSeq6000 sequencing platform is designed to be the sequence shown in Seq ID No. 54.

[0081] Seq ID No.54:

[0082] 5'-AATGATACGGCGACCACCGAGATCTACACNNNNNNACACTCTTTCCCTACACGACGCTCTTCCGATCT-3'.

[0083] "NNNNNN" in the fourth primer of the sequence shown in Seq ID No. 54 is the Index sequence.

[0084] The fifth primer sequence is the complementary sequence of the downstream sequencing adapter sequence with the barcode. The index length can be 6-10 nt. For example, referring to the P7 sequence of the Illumina NovaSeq6000 sequencing platform, the design is the sequence shown in Seq ID No. 55. Seq ID No. 55:

[0085] 5'-CAAGCAGAAGACGGCATACGAGATNNNNNNGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT-3'.

[0086] "NNNNNN" in the fifth primer of the sequence shown in Seq ID No. 55 is the Index sequence.

[0087] Based on the method for rapidly adding UMIs and the complete library preparation technology of the first to fifth primers of this application, the process includes: 1. Using the first primer to extend the template nucleic acid to obtain a complementary chain and introduce a UMI tag. 2. Adding the second primer to completely extend the complementary chain in step 1. 3. Using UDG / UNG enzyme to digest deoxyuracil and digest the first primer. 4. Adding the third primer to amplify and enrich the product of step 2; and purify the amplified product to remove system, excess primers and genomic contamination. 5. Using the fourth and fifth primers to amplify the library, add the barcode and sequencing adapter to complete the library construction.

[0088] Specifically, the technical process is described in detail as follows:

[0089] Step 1: Introduce UMI tags and extend the template complementary chain:

[0090] Take a template nucleic acid containing the target region, such as DNA / cDNA, with a total amount of 1-100 ng, preferably genomic DNA as the template, and the preferred nucleic acid amount is 100 ng.

[0091] Prepare the extension system: a commercially available PCR amplification kit or a self-developed PCR amplification kit can be used, the main components of which may include but are not limited to: DNA polymerase, Mg ions, dNTPs, and a buffer system.

[0092] If the experiment is designed as a multiplex PCR reaction, it is preferred to choose a commercially available or self-developed multiplex PCR amplification kit.

[0093] The first primer is added to the prepared extension system, and its working concentration can be 50-500 mM; preferably, the working concentration is set to 200 mM, that is, the concentration of each primer is 200 mM.

[0094] Add the prepared template DNA / cDNA to the prepared extension system, mix thoroughly, and then perform the template complementary strand extension reaction. Reaction parameters should be set according to the PCR amplification kit instructions. It should be noted that the extension time should be adjusted to be greater than the ratio of "target region length / extension speed" to ensure that the target region is fully and completely extended. Not setting the number of PCR cycles or setting the number of cycles to 1 means that only one extension reaction is performed, without repeated deformation, annealing, and extension.

[0095] After the reaction is completed, the resulting product chain is the template complementary chain to which the UMI has been added.

[0096] Step 2: Template complementary chain complementary extension

[0097] Take the first step reaction product and add the second primer, the working concentration of which can be 50-500mM; preferably, the working concentration is 200mM, that is, the concentration of each primer is 200mM.

[0098] After mixing evenly, put it into the PCR program, and the PCR program remains exactly the same as the first step.

[0099] After the reaction is completed, the resulting product chain is the library fragment with the UMI added, and its sequence is consistent with the target fragment on the template.

[0100] Step 3: Digestion of deoxyuridine using UDG / UNG enzyme

[0101] Prepare thermosensitive UDG / UNG enzyme, which can be commercially available or homemade.

[0102] Remove the second-step reaction product and add the prepared thermosensitive UDG / UNG enzyme. Adjust the enzyme dosage based on enzyme activity and digestion efficiency. Generally, if the enzyme activity is greater than 1U / μL, add 1μL. After thorough mixing, digest all deoxyuridine-containing sequences in the system according to the enzyme's optimal reaction temperature and conditions.

[0103] The purpose of this step is to digest any excess first primer in the system and the first primer extension sequence involved in the first step. Ultimately, the resulting product contains only the initial DNA template and the extension product chain with unique UMI tags from the second step.

[0104] Step 4: Specific amplification and enrichment

[0105] The third primer is added to the mixed third step product at a working concentration of 50-500 mM; preferably, the working concentration is 200 mM.

[0106] After thorough mixing, perform a template-complementary strand extension reaction. Reaction parameters should be set according to the PCR amplification kit instructions. The number of PCR cycles can be customized based on project requirements and kit performance, preferably 5 or more cycles.

[0107] After amplification is complete, the product is removed and purified. This results in a highly purified product free of impurities for library construction and amplification. Purification methods can include, but are not limited to, magnetic bead-based methods, column chromatography, or gel electrophoresis.

[0108] Step 5: Library construction and amplification

[0109] The purified product obtained in step 4 is added to the amplification system, and the fourth primer and the fifth primer are mixed thoroughly. The working concentration of the fourth primer and the fifth primer can be 200-2000 mM, preferably 1500 mM.

[0110] Reaction parameters should be set according to the PCR amplification kit instructions. The number of PCR cycles can be customized based on project requirements and kit performance.

[0111] After the reaction is complete, the product is a ready-to-use library with complete adapter information. Nucleic acid purification is performed to obtain a highly pure library. After quality control and quantification, the library is ready for sequencing.

[0112] The present invention is further described in detail below through specific examples. The following examples are only used to further illustrate the present invention and should not be construed as limiting the present invention.

[0113] Example

[0114] Following the above methods and ideas, this example designed the first to fifth primers for human immune repertoire sequencing for testing. This example incorporated UMIs into human immune repertoire sequencing to correct for biased amplification and amplification errors. Using genomic DNA as a template, 40 upstream V region primers and 12 downstream J region primers were designed to cover all subtypes in this region. The first and second primers designed in this example performed multiplex PCR amplification of the TCRβ chain CDR3 region, eliminating the effects of biased amplification and accumulated amplification errors, resulting in more realistic results. The details are as follows:

[0115] A genomic DNA sample extracted from human peripheral blood was collected for future use. The genomic DNA sample was provided and stored by Shenzhen Hypros Biotechnology Co., Ltd.

[0116] According to the above ideas, the first to fifth primers were designed:

[0117] For the first primer, in this example, reference sequences for the V gene of the human T lymphocyte receptor β chain were retrieved from the NCBI and IMGT standard databases. Upstream primers were independently designed based on base pairing and primer design principles, incorporating the desired sites. A total of 40 target-specific upstream primer sequences were designed, encompassing 42 V functional regions. All functional region primers were mixed in equal proportions based on molar mass. The first primers in this example consisted of 40 primers with sequences shown in Seq ID No. 1 to Seq ID No. 40, as shown in Table 1.

[0118] For the second primers, in this example, reference sequences for the J gene of the human T lymphocyte receptor β chain were searched using standard databases such as NCBI and IMGT. Downstream primers were independently designed based on base pairing and primer design principles, incorporating the desired sites. In this example, 12 target-specific upstream primer sequences were designed, encompassing six J1 functional regions and seven J2 functional regions. All functional region primers were mixed in equal proportions based on molar mass. The second primers in this example consisted of 12 primers with sequences represented by Seq ID No. 41 to Seq ID No. 52, as shown in Table 1.

[0119] Table 1 First primer and second primer

[0120]

[0121]

[0122] The third primer is the sequence shown in Seq ID No.53,

[0123] Seq ID No. 53: 5'-CACGACGCTCTCCGATCT-3'.

[0124] The fourth primer is the sequence shown in Seq ID No.54,

[0125] Seq ID No.54:

[0126] 5'-AATGATACGGCGACCACCGAGATCTACACNNNNNNACACTCTTTCCCTACACGAGCGCTTCCGATCT-3';

[0127] In the fourth primer of the sequence shown in Seq ID No. 54, "NNNNNN" is specifically "TGCGTAAT".

[0128] The fifth primer is the sequence shown in Seq ID No.55,

[0129] Seq ID No.55:

[0130] 5'-CAAGCAGAAGACGGCATACGAGATNNNNNNGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT-3';

[0131] In the fifth primer of the sequence shown in Seq ID No. 55, "NNNNNN" is specifically "CCTAACCT".

[0132] After designing and synthesizing the above primers, the primers were diluted with TE buffer. The first primer, the second primer, and the third primer were diluted to a concentration of 5 μM, and the fourth primer and the fifth primer were diluted to a concentration of 30 μM.

[0133] Take 100 ng of the genomic DNA of the sample mentioned above and perform amplification experiment using the QIAgen Multiplex PCR kit.

[0134] Prepare the reaction system and add it to a new 0.2 mL tube in sequence. The reaction system is: 4.5 μL of the first primer, 100 ng of genomic DNA, 25 μL of PCR Master mix, 5 μL of Q-solution, and add NF water to 45 μL.

[0135] After gently mixing the 0.2 mL sample tube containing the sample and reagent, place it in a BIORAD T100 PCR instrument for PCR reaction. The reaction program was as follows: denaturation at 95°C for 15 min, followed by 94°C for 30 s, 60°C for 90 s, and 72°C for 90 s, and finally extension at 72°C for 5 min, and standby at 4°C.

[0136] After the reaction is completed, the reaction product is taken out, 5 μL of the second primer is added, and the mixture is placed in a BIORAD T100PCR instrument for PCR reaction. The reaction program is: denaturation at 95°C for 15 min, then 94°C for 30 s, 60°C for 90 s, 72°C for 90 s, and finally extension at 72°C for 5 min, and standby at 4°C.

[0137] After the reaction was completed, the reaction product was taken out, 1 μL of heat-labile UDG enzyme (Heat-labile UDG, Vazyme) was added, mixed and placed in the following reaction program: digestion at 25°C for 10 min, inactivation at 55°C for 5 min, 95°C for 5 min, and standby at 4°C.

[0138] After the reaction is completed, the reaction product is taken out, 5 μL of the third primer is added, and the mixture is placed in a BIORAD T100PCR instrument for PCR reaction. The reaction program is: denaturation at 95°C for 15 min, followed by 30 cycles of: 94°C for 30 s, 56°C for 90 s, and 72°C for 90 s. After the cycle, extension is performed at 72°C for 5 min, and then standby at 4°C.

[0139] After the reaction is completed, the reaction product is taken out and purified using magnetic beads. The detailed steps are as follows:

[0140] 1. Purify multiplex PCR products using 1.2× AMpure XP beads: Add 50 μL of the multiplex PCR product and 60 μL of a uniform mixture of AMpure XP beads to a new 1.5 mL sample tube. Vortex to mix thoroughly. Incubate at room temperature for 10 minutes to allow the DNA to fully bind to the beads. Place the 1.5 mL tube on a magnetic stand and allow the beads to adsorb until the solution is clear. Carefully remove the supernatant.

[0141] 2. Add 500 μL of 80% ethanol and rotate the sample tube 180 degrees to allow the magnetic beads to pass through the solution and be absorbed to the other side of the tube wall. Rotate 2-3 times, let it stand for 15 seconds, and discard the supernatant.

[0142] 3. Repeat step 2 once;

[0143] 4. Allow the 1.5mL sample tube to stand naturally until all the alcohol has evaporated. Then, add 20μL of nuclease-free water to the 1.5mL tube and mix thoroughly. Place the 1.5mL tube on a magnetic stand and allow the beads to adsorb until the solution is clear. Carefully aspirate the supernatant and transfer it to a new 0.2μL tube to obtain the purified product.

[0144] For sequencing library construction, the fourth and fifth primers, amplification reagents, and purified product were mixed for library construction and amplification. KAPA HiFi Hotstart Ready Mix was used for amplification in the following proportions: 25 μL of 2× KAPA HiFi Hotstart Ready Mix, 2.5 μL of the fourth primer, 2.5 μL of the fifth primer, and 20 μL of the purified product. The mixture was then filled to 50 μL with NF water.

[0145] After mixing evenly, place the reaction in the following program: denaturation at 98°C for 45 seconds, followed by 5 cycles of: 98°C for 15 seconds, 60°C for 30 seconds, and 72°C for 30 seconds. After the cycle, extend at 72°C for 1 minute and stand by at 4°C.

[0146] After the program is completed, 50 μL of the library amplification PCR product is obtained and purified using 1× AMpure XP magnetic beads for multiplex PCR product purification:

[0147] In a new 1.5mL sample tube, add 50μL of the multiplex PCR product and 50μL of a uniformly mixed AMpure XP magnetic beads. Vortex to mix thoroughly and incubate at room temperature for 10 minutes to allow the DNA to fully bind to the beads. Place the 1.5mL tube on a magnetic stand and allow the beads to adsorb until the solution is clear. Carefully remove the supernatant.

[0148] Add 500 μL of 80% ethanol, rotate the sample tube 180 degrees to allow the magnetic beads to pass through the solution and be absorbed to the other side of the tube wall, rotate 2-3 times, let it stand for 15 seconds, and discard the supernatant; repeat this step once.

[0149] Allow the 1.5mL sample tube to stand naturally until the alcohol has completely evaporated. Then, add 20μL of Nuclease-Free Water to the tube and mix thoroughly. Place the tube on a magnetic stand and allow the beads to adsorb until the solution is clear. Carefully aspirate the supernatant, label, and store. This completes the amplicon library construction.

[0150] High-throughput sequencing: The Illumina NovaSeq 6000 high-throughput sequencing system was used for sequencing, and the sequencing mode was PE151+8+8+151.

[0151] result:

[0152] After standard quality control and sequencing depth threshold filtering, the sequencing data were analyzed for UMIs. Two analysis methods were used in this case. The first method used comparative analysis data, ignoring the presence of UMIs and directly analyzing the proportion of functional regions in the upstream V region and downstream J region, as well as diversity, clonal index, and related metrics. The second method used UMI filtering. Sequences with identical UMIs and CDR3 sequences were considered to originate from the same parent template and counted as one copy. Furthermore, all CDR3 sequences with the same UMI were compared to analyze reads for single or several occasional base differences. If any, these differences were filtered and corrected for amplification and experimental errors. This analysis was then used to analyze the proportion of functional regions in the upstream V region and downstream J region, as well as diversity, clonal index, and related metrics. The results are shown in Table 2.

[0153] Table 2 Comparative analysis of the impact of UMI sequences on human immune repertoire sequencing results

[0154] Analytical methods Do not filter UMI Filter UMI Total number of reads 12503452 12503452 Percentage of incorrect bases (%) 98 98 Total clone reads (strips) 6559378 6559378 Percentage of crawled reads (%) 98 98 Uncaptured Reads (%) 2 2 Randomly grab raw_read (items) 1000000 1000000 Number of unique cloned amino acid sequences 30456 20903 The highest frequency amino acid sequence percentage (%) 0.13 0.04 Shannon Index 11.2 10.2

[0155] In Table 2, "No UMI Filtering" refers to the first analysis method, "Filtering UMI" refers to the second analysis method, "Total Reads" refers to the total number of reads contained in the raw data, "Error Base Percentage" refers to the percentage of bases with an error rate of less than 0.1%, "Captured Reads Percentage" refers to the percentage of reads that captured the CDR3 region, "Uncaptured Reads" refers to the percentage of reads that did not capture the CDR3 region, "Randomly Captured Raw Reads" refers to reads randomly captured for analysis, "Unique Cloned Amino Acid Sequence Number" refers to the number of unique cloned amino acid sequences, and the higher the Shannon Index value, the higher the diversity of the immune repertoire.

[0156] The results show that without UMI filtering, the proportion of the highest frequency amino acid sequence is artificially high because biased amplification cannot be removed. Using UMI filtering can remove biased amplification and more accurately reflect the highest CDR3 proportion. In addition, when the data is not filtered by UMI, the number of unique cloned amino acid sequences is higher than the Shannon index. This is because a large number of amplification errors accumulate under high amplification cycle numbers. In the absence of UMI, it is difficult to identify whether these sequences are template sequences themselves or experimentally introduced errors. However, erroneous sequences can also increase the uniqueness and diversity of CDR3 and enter the analysis process, resulting in artificially high indicators. Data filtered by UMI is closer to the actual situation, correcting inaccurate quantification caused by amplification bias and correcting the artificially high analytical indicators caused by amplification errors or experimentally introduced errors, thereby improving data credibility.

[0157] Therefore, the human immune repertoire sequencing method in this example has a good correction effect on sequencing errors, which can correct the inaccurate quantification caused by amplification bias and correct the inflated positive rate / false positive caused by amplification errors or errors introduced by the experiment.

[0158] The above content is a further detailed description of the present application in conjunction with specific implementation methods, and the specific implementation of the present application cannot be considered to be limited to these descriptions. For ordinary technicians in the technical field to which the present application belongs, several simple deductions or substitutions can be made without departing from the concept of the present application. SEQUENCE LISTING <110> Shenzhen Hypros Medical Laboratory <120> A method and kit for sequencing human immune repertoire <130> 22I33618 <160> 57 <170> PatentIn version 3.3 <210> 1 <211> 94 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 1 cacgacgcuc uuccgaucun nnnunnnncg gcauacgaac ugucauuaug cggcauacga 60 acugucauua ugcggcauac gaacugucau uaug 94 <210> 2 <211> 76 <212> RNA <213> artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 2 cacgacgcuc uuccgaucun nnnunnncc cacguuaacc cugauuaua cacccacguu 60 aacccuagau uauca 76 <210> 3 <211> 72 <212> RNA <213> artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 3 cacgacgcuc uuccgaucun nnnunnnnac cguagauuau acaaugccug accguagauu 60 auacaaugcc ug 72 <210> 4 <211> 74 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 4 cacgacgcuc uuccgaucun nnnunnnnga acguuauauc cauauagaua ugaacguuau 60 auccauauag auau 743] <210> 5 <211> 70 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 5 cacgacgcuc uuccgaucun nnnunnnncg cugccugugc gugaauuggc gcugccugug 60 cgugaauugg 70 <210> 6 <211> 64 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 6 cacgacgcuc uuccgaucun nnnunnnnug ccugugccag aauuggugcc ugugccagaa 60 uugg 64 <210> 7 <211> 64 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 7 cacgacgcuc uuccgaucun nnnunnnngu gccagaauug uugauggugc cagaauuguu 60 gaug 64 <210> 8 <211> 66 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 8 cacgacgcuc uuccgaucun nnnunnnnau gauaaauaaa cgcacuaaug auaaauaaac 60 gcacua 66 <210> 9 <211> 66 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 9 cacgacgcuc uuccgaucun nnnunnnnau gauaaauaaa cccacuaaug auaaauaaac 60 ccacua 66 <210> 10 <211> 68 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 10 cacgacgcuc uuccgaucun nnnunnnnca uuaugugaac gacgugcaca uuaugugaac 60 gacgugca 68 <210> 11 <211> 74 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 11 cacgacgcuc uuccgaucun nnnunnnncg cauuauauuu cauuauguga acgcauuaua 60 uuucauuaug ugaa 74 <210> 12 <211> 64 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 12 cacgacgcuc uuccgaucun nnnunnnnug ccguuaacga gacacaugcc guuaacgaga 60 caca 64 <210> 13 <211> 66 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 13 cacgacgcuc uuccgaucun nnnunnnnua acauaugcag cuaacgauaa cauaugcagc 60 uaacga 66 <210> 14 <211> 72 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28)[[ID=5�]] <223> n is a, c, g, or u <400> 14 cacgacgcuc uuccgaucun nnnunnnnua uaucauaugc cacuaacgag uauaucauau 60 gccacuaacg ag 72 <210> 15 <211> 68 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 15 cacgacgcuc uuccgaucun nnnunnnnca cgcgacgggg gcauacgaca cgcgacgggg 60 gcauacga 68 <210> 16 <211> 70 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 16<000059�>cacgacgcuc uuccgaucun nnnunnnnua acuauaacau augcagcuuu aacuauaaca 60 uaugcagcuu 70 <210> 17 <211> 64 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 17 cacgacgcuc uuccgaucun nnnunnnncg gcauaagaag ugucuacggc auaagaagug 60 ucua 64 <210> 18 <211> 64 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 18 cacgacgcuc uuccgaucun nnnunnnnuu ggcgcuaacg agacacuugg cgcuaacgag 60 acac 64 <210> 19 <211> 64 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221] misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 19 cacgacgcuc uuccgaucun nnnunnnnuu augugaacga cgagucuuau gugaacgacg 60 aguc 64 <210> 20 <211> 66 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 20 cacgacgcuc uuccgaucun nnnunnnnau gaauaauaaa cgcacuaaug aauaauaaac 60 gcacua 66 <210> 21 <211> 62 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23)<ocke000665><223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 21 cacgacgcuc uuccgaucun nnnunnnca aagauaaacg cacuacaaag auaaacgcac 60 62 years old <210> 22 <211> 72 <212> RNA <213> artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 22 cacgacgcuc uuccgaucun nnnunnngu gucauuaugu gaacuacgug gugucauuau 60 gugaacuacg ug 72 <210> 23 <211> 64 <212> RNA <213> artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 23 cacgacgcuc uuccgaucun nnnunnngg gcauaggaac ugucuagggc auaggaacug 60 ucua 64 <210> 24 <211> 64 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 24 cacgacgcuc uuccgaucun nnnunnnncg cuaacgaaug acccgacgcu aacgaaugac 60 ccga 64 <210> 25 <211> 72 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 25 cacgacgcuc uuccgaucun nnnunnnnca uuaugugaac gacguuucga cauuauguga 60 acgacguuuc ga 72 <210> 26 <211> 68 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 26 cacgacgcuc uuccgaucun nnnunnnngg gcauacgaaa ugucauuagg gcauacgaaa 60 ugucauua 68 <210> 27 <211> 72 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220><000​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​<221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 28 cacgacgcuc uuccgaucun nnnunnnncu aacggcggaa gccaccuaac ggcggaagcc 60 ac 62 <210> 29 <211> 70 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 29 cacgacgcuc uuccgaucun nnnunnnnga uugcggcuua cgacgugucg auugcggcuu 60 acgacguguc 70 <210> 30 <211> 68 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <22​​​​​​​cacgacgcuc uuccgaucun nnnunnnngg gcauacgaag ugucuauagg gcauacgaag 60 ugucuaua 68 <210> 31 <211> 66 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 31 cacgacgcuc uuccgaucun nnnunnnncg acggucguga gcagcgccga cggucgugag 60 cagcgc 66<00008​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​ <211> 68 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 33 cacgacgcuc uuccgaucun nnnunnnngc acgugcaacg ugaaacuagc acgugcaacg 60 ugaaacua 68 <210> 34 <211> 62 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 34 cacgacgcuc uuccgaucun nnnunnnnga auggaccuaa uguaagaaug gaccuaaugu 60 aa 62 [[ID=�9]]<210> 35 <211> 62 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 35 cacgacgcuc uuccgaucun nnnunnnnua ugaacuaaaa ccaacuauga acuaaaacca 60 ac 62 <210> 36 <211> 66 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 36 cacgacgcuc uuccgaucun nnnunnnnaa uaugugaacg acgugccaau augugaacga 60 cgugcc 66 <210> 37 <211> 72 <212> RNA <213> Artificial sequence <220> <221> misc_feature[[ID=6o]] <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature It should be noted that there may be an error in the original text where the "6o" in the ID "6o" should probably be "60". The above translation is based on the original content as provided. <222> (25)..(28) <223> n is a, c, g, or u <400> 37 cacgacgcuc uuccgaucun nnnunnnnca uuaugugaac gacguuucua cauuauguga 60 acgacguuuc ua 72 <210> 38 <211> 70 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 38 cacgacgcuc uuccgaucun nnnunnnngu caauauguga acgacguuug ucaauaugug 60 aacgacguuu 70 <210> 39 <211> 70 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 39 cacgacgcuc uuccgaucun nnnunnnncc acugcgaaau gccgcacgac cacugcgaaa 60 ugccgcacga 70 <210> 40 <211> 74 <212> RNA <213> Artificial sequence <220> <221> misc_feature <222> (20)..(23) <223> n is a, c, g, or u <220> <221> misc_feature <222> (25)..(28) <223> n is a, c, g, or u <400> 40 60 auauagauua ucua 74 <210> 41 <211> 40 <212> DNA <213> Artificial sequence <400> 41 agacgtgtgc tcttccgatc ttgcgtatgg tgaattgtaa 40 <210> 42 <211> 37 <212> DNA <213> Artificial sequence <400> 42 agacgtgtgc tcttccgatc ttaaaagcca agccggt 37 <210> 43 <211> 42 <212> DNA <213> Artificial sequence <400> 43 agacgtgtgc tcttccgatc ttcaccacgt gcgaaccatt aa 42 <210> 44 <211> 40 <212> DNA <213> Artificial sequence <400> 44 agacgtgtgc tcttccgatc tcgtaaacta caacccctga 40 <210> 45 <211> 39 <212> DNA <213> Artificial sequence <400> 45 agacgtgtgc tcttccgatc ttggtaaact taaacccgt 39 <210> 46 <211> 42 <212> DNA <213> Artificial sequence <400> 46 agacgtgtgc tcttccgatc taattattca atcgacaggt gc 42 <210> 47 <211> 41 <212> DNA <213> Artificial sequence <400> 47 agacgtgtgc tcttccgatc tgtacgaatc gcgaattata a 41 <210> 48 <211> 37 <212> DNA <213> Artificial sequence <400> 48 agacgtgtgc tcttccgatc tggtgaatgg gaaccc 37 <210> 49 <211> 37 <212> DNA <213> Artificial sequence <400> 49 agacgtgtgc tcttccgatc tacggcgaac atagaga 37 <210> 50 <211> 38 <212> DNA <213> Artificial sequence <400> 50 agacgtgtgc tcttccgatc taaaagcccg tccggcag 38 <210> 51 <211> 41 <212> DNA <213> Artificial sequence <400> 51 agacgtgtgc tcttccgatc ttgtgcaagt gcgaatggtg a 41 <210> 52 <211> 41 <212> DNA <213> Artificial sequence <400> 52 agacgtgtgc tcttccgatc ttgtgcaagt gcgaatggtg a 41 <210> 53 <211> 19 <212> DNA <213> Artificial sequence <400> 53 cacgacgctc ttccgatct 19 <210> 54 <211> 68 <212> DNA <213> Artificial sequence <220> <221> misc_feature <222> (30)..(35) <223> n is a, c, g, or t <400> 54 aatgatacgg cgaccaccga gatctacacn nnnnnacact ctttccctac acgacgctct 60 tccgatct 68 <210> 55 <211> 64 <212> DNA <213> Artificial sequence <220> <221> misc_feature <222> (25)..(30) <223> n is a, c, g, or t <400> 55 caagcagaag acggcatacg agatnnnnnn gtgactggag ttcagacgtg tgctcttccg 60 atct 64 <210> 56 <211> 19 <212> RNA <213> Artificial sequence <400> 56 cacgacgcuc uuccgaucu 19 <210> 57 <211> twenty one <212> DNA <213> Artificial sequence <400> 57 agacgtgtgc tcttccgatc t 21

Claims

1. A method for sequencing a human immune repertoire, characterized in that: The following steps are included: A reaction system is prepared, wherein a first primer is used to extend the template nucleic acid once to obtain a complementary strand; the first primer includes, from the 5' end to the 3' end, a sequencing platform upstream primer binding region, a unique identifier, and a target-specific upstream primer sequence; furthermore, in the first primer, the base T in the sequencing platform upstream primer binding region and the target-specific upstream primer sequence is replaced with deoxyuracil, and the sequencing platform upstream primer binding region corresponds to the 3' end of the upstream sequencing primer of the sequencing platform; After the first primer is extended, a second primer is added to the reaction system, and the complementary strand of the first primer extension is extended once using the second primer to obtain a product consisting of a sequencing platform upstream primer binding region, a unique identifier, a target sequence, and a sequencing platform downstream primer binding region; the second primer includes, from the 5' end to the 3' end, the sequencing platform downstream primer binding region and a target-specific downstream primer sequence, and the sequencing platform downstream primer binding region corresponds to the 3' end of the sequencing platform downstream sequencing primer; After the second primer is extended, UDG / UNG enzyme is added to the reaction system to digest deoxyuracil, thereby digesting the first primer and the extended chain of the first primer; After the UDG / UNG enzyme digestion is completed, a third primer is added to the reaction system, and the product extended by the second primer is PCR amplified and enriched using the third primer and the second primer to obtain a product in which all amplicons of the template nucleic acid are added with the same unique identifier; The third primer is the entire or partial sequence of the sequencing platform upstream primer binding region of the first primer from the 5' end, and the base T in the third primer is not replaced by deoxyuracil; The same unique identifier is added to all amplicons obtained by PCR amplification and enrichment, and the sequencing library is constructed and sequenced, thus completing the sequencing of the human immune repertoire; The target-specific upstream primer sequence and the target-specific downstream primer sequence are specific primer sequences designed for human T cell receptor encoding genes and fully cover the CDR3 region encoding gene sequence thereof; The first primers consist of 40 primers with sequences shown in Seq ID No.1 to Seq ID No.40; The second primers consist of 12 primers with sequences shown in Seq ID No.41 to Seq ID No.52; The third primer is the sequence shown in Seq ID No.

53.

2. The method according to claim 1, wherein: The target-specific upstream primer sequence is a specific primer sequence designed for the V gene of the human T cell receptor β chain; The target-specific downstream primer sequence is a specific primer sequence designed for the J gene of the human T cell receptor β chain; Amplification using the target-specific upstream primer sequence and the target-specific downstream primer sequence can fully cover the gene encoding the CDR3 region of the T cell receptor β chain.

3. The method according to claim 1, wherein: In the first primer, at least one deoxyuracil is inserted into the sequence of the unique identifier, and the number of consecutive bases of the unique identifier is less than 5 by the insertion of the deoxyuracil.

4. The method according to claim 1, wherein: The number of amplification cycles of the PCR amplification enrichment is greater than or equal to 5.

5. The method according to any one of claims 1 to 4, characterized in that: The sequencing library construction includes the following steps: Purifying products to which the same unique identifier is added to all amplicons obtained by the PCR amplification enrichment to obtain purified products; The purified product was amplified using the fourth primer and the fifth primer to obtain a sequencing library; the fourth primer was an upstream sequencing primer of the sequencing platform with a sequencing adapter and a barcode, and the fifth primer was a downstream sequencing primer of the sequencing platform with a sequencing adapter and a barcode.

6. The method according to claim 5, characterized in that: The fourth primer is the sequence shown in Seq ID No.

54.

7. The method according to claim 5, characterized in that: The fifth primer is the sequence shown in Seq ID No.

55.

8. A kit for sequencing the human immune repertoire, characterized in that: It includes a first primer, a second primer, a third primer and UDG / UNG enzyme; The first primer includes, from the 5' end to the 3' end, a sequencing platform upstream primer binding region, a unique identifier, and a target-specific upstream primer sequence; and, in the first primer, the base T in the sequencing platform upstream primer binding region and the target-specific upstream primer sequence is replaced with deoxyuracil, and the sequencing platform upstream primer binding region corresponds to the 3' end of the upstream sequencing primer of the sequencing platform; The second primer includes a sequencing platform downstream primer binding region and a target-specific downstream primer sequence in sequence from the 5' end to the 3' end, and the sequencing platform downstream primer binding region corresponds to the 3' end of the downstream sequencing primer of the sequencing platform; The third primer is the entire or partial sequence of the sequencing platform upstream primer binding region of the first primer from the 5' end, and the base T in the third primer is not replaced by deoxyuracil; The target-specific upstream primer sequence and the target-specific downstream primer sequence are specific primer sequences designed for human T cell receptor encoding genes and fully cover the CDR3 region encoding gene sequence thereof; The first primers consist of 40 primers with sequences shown in Seq ID No.1 to Seq ID No.40; The second primers consist of 12 primers with sequences shown in Seq ID No.41 to Seq ID No.52; The third primer is the sequence shown in Seq ID No.

53.

9. The kit according to claim 8, wherein: The target-specific upstream primer sequence is a specific primer sequence designed for the V gene of the human T cell receptor β chain; The target-specific downstream primer sequence is a specific primer sequence designed for the J gene of the human T cell receptor β chain; Amplification using the target-specific upstream primer sequence and the target-specific downstream primer sequence can fully cover the gene encoding the CDR3 region of the T cell receptor β chain.

10. The kit according to claim 8, wherein: In the first primer, at least one deoxyuracil is inserted into the sequence of the unique identifier, and the number of consecutive bases of the unique identifier is less than 5 by the insertion of the deoxyuracil.

11. The kit according to any one of claims 8 to 10, characterized in that: Also included are a fourth primer and a fifth primer; The fourth primer is an upstream sequencing primer of the sequencing platform with a sequencing adapter and a barcode, and the fifth primer is a downstream sequencing primer of the sequencing platform with a sequencing adapter and a barcode.

12. The kit according to claim 11, characterized in that: The fourth primer is the sequence shown in Seq ID No.

54.

13. The kit according to claim 11, wherein: The fifth primer is the sequence shown in Seq ID No.55.

Citation Information

Patent Citations

  • Method for accurately detecting T cell immune repertoire based on high-throughput sequencing and primer system thereof

    CN113122618A

  • Method for adding unique identifier in amplicon sequencing and application

    CN114277114A