Method for three-generation long-reading sequencing typing of 16 KIR gene full-length sequences

By designing specific PCR primer sets and using long-distance PCR amplification method, combined with three generations of long-read and long-sequencing technology, the problem of full-length sequence sequencing of KIR genes in the existing technology is solved, and high-throughput, low-consumption KIR genotyping is achieved, providing more accurate and efficient typing results.

CN120118992APending Publication Date: 2025-06-10SHENZHEN BLOOD CENT +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510438090.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing KIR gene sequencing and typing technology cannot effectively detect the full-length sequence of the KIR gene, and there are problems such as long fragments of the amplification product, long amplification time, non-specific amplification and ambiguous results, making it difficult to achieve high-throughput and high-resolution KIR genotyping.

Method used

A PCR amplification primer set of 16 KIR genes full-length sequences was designed, and the full-length sequences of all 16 KIR genes in humans were amplified by long-distance PCR method, and the typing was combined with three generations of long-read and long-term sequencing technology. High-efficiency and low-consuming typing of KIR genes was achieved through library construction and on-machine sequencing.

Benefits of technology

High-throughput and low-consumption typing of the full-length sequence of the KIR gene is achieved, and ambiguous results are avoided. The amplification and library construction operation is simple and low cost is low. It can provide good application prospects in the fields of KIR population genetics, disease associations and hematopoietic stem cell transplantation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120118992A_ABST
    Figure CN120118992A_ABST
Patent Text Reader

Abstract

The invention provides a three-generation long-read-long sequencing and typing method for 16 KIR gene full-length sequences, and relates to the field of DNA sequencing and typing. Based on the structure of a KIR gene full-length sequence, single nucleotide polymorphism distribution and KIR allele polymorphism of Chinese population, aiming at all 16 KIR genes in a KIR gene family, three pairs of specific primers are designed, and a long-distance PCR (Polymerase Chain Reaction) method is adopted to amplify the full-length sequences of all 16 KIR genes of human beings. After the concentration and purity of amplification products are measured, the amplification products are mixed into one tube according to a certain mass ratio, and the KIR allele carried by a subject can be judged through library construction, on-machine sequencing and bioinformatics analysis. The high-throughput KIR gene three-generation long-reading sequencing typing method established for the first time has the characteristics of simplicity and convenience in operation of an amplification library building experiment, low cost, no ambiguous result and the like, and has a wide application prospect in the fields of population genetics, disease association, bone marrow transplantation and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of DNA sequencing and genotyping, and particularly to a method for genotyping the full-length sequences of 16 KIR genes by third-generation long-read sequencing. Background Art

[0002] Natural killer (NK) cells are an important part of the body's innate immunity. Killer cell immunoglobulin-like receptors (KIRs), which are expressed on the surface of NK cells and some activated T cells, interact with human leukocyte antigen (HLA) class I molecules on the surface of target cells to form the KIR-HLA receptor-ligand complex, which conducts activation or inhibition signals and jointly regulates the activity of NK cells, playing an important role in tumor immunity, fetal-maternal immunity, and the body's anti-infection.

[0003] KIR receptors can be divided into KIR2D and KIR3D according to the number of extracellular domains; according to the length of the cytoplasmic region and the presence or absence of immunoreceptor tyrosine-based inhibitory motifs (ITIMs), KIRs are divided into inhibitory (L-type) and activating (S-type).

[0004] The KIR gene family encoding KIR receptors is located on human chromosome 19 and contains 14 functional KIR genes (2DL1, 2DL2, 2DL3, 2DL4, 2DL5, 2DS1, 2DS2, 2DS3, 2DS4, 2DS5, 3DL1, 3DL2, 3DL3, 3DS1) and two pseudogenes (2DP1 and 3DP1). [1] In the international IPD-KIR database [2] (https: / / www.ebi.ac.uk / ipd / kir / ), the genomic full-length sequences of the 14 functional KIR genes that are included range from 9901 to 17009 bp, and the genomic full-length sequences of the two KIR pseudogenes 2DP1 and 3DP1 are 13126 bp and 5713 bp, respectively. The full-length coding sequence (CDS) of each functional KIR gene is only 915 to 1368 bp, while the non-coding region sequence is as long as 8773 to 15641 bp (see Table 1). In-depth analysis of the fragment sizes of the full-length sequences of each KIR gene is crucial for designing specific primers to amplify amplicons with similar target fragment sizes to ensure successful multiplex amplicon pooled sequencing.

[0005] Table 1 Structure of the coding region of each KIR gene and the lengths of the coding region, non-coding region, and full-length sequence

[0006] Note: 2DP1 and 3DP1 in this table are pseudogenes; Del: deletion of exons; Pseudoexon: pseudoexon.

[0007] The structure of KIR genes is complex. Among KIR2D genes, the third exon of 8 KIR genes such as KIR2DL1-3 and 2DS1-5 is a pseudoexon and cannot encode the corresponding D0 extracellular domain; KIR2DL4 and 2DL5 both lack the fourth exon and cannot encode the D1 extracellular domain. Among functional KIR3D genes, except for KIR3DL3 lacking the sixth exon, the remaining functional KIR3D genes (3DL1, 3DL2, 3DS1) all contain 9 exons, 8 introns, as well as the 5'-promoter region and 3'-UTR region.

[0008] KIR genes are inherited in haplotypes and have complex molecular genetic polymorphisms, which are reflected in: (1) There are differences in the composition and copy number of KIR genes on the KIR haplotype: Generally, except for 4 structural genes (3DL3, 3DP1, 2DL4, and 3DL2), other KIR genes may be present or absent in an individual, and the copy number of each gene is 0-3; (2) At the allelic level: As of December 2024, the total number of KIR alleles published in the international IPD-KIR Database (Release 2.13) has reached 2219 (including 47 non-expressing alleles); each KIR gene has 33-428 different alleles (see Table 2 for details). The homology between any two different KIR allele sequences is as high as 85%-98%. [3] The genetic complexity of the KIR gene family and the high homology of allelic sequences pose great challenges to KIR gene sequencing and typing.

[0009] Table 2 KIR alleles published in the IPD-KIR database (Release 2.13)

[0010] Identifying KIR alleles has functional significance. The expression levels of different KIR genes are different. [4], there are also significant differences in the expression levels, affinities, and mediated activation / inhibition abilities among different alleles of the same KIR gene. For example, the expression levels of 3DL1 alleles on the NK cell membrane from high to low are: KIR3DL1*01502 > *020 > *001 > *007 > *005, and the strengths of the mediated inhibitory effects are: KIR3DL1*001 > *005 > *01502 > *020 > *007 [5] .

[0011] Identifying KIR alleles has clinical significance. Boudreau et al. analyzed 1328 patients with acute myeloid leukemia who received hematopoietic stem cell transplantation from HLA 9-10 / 10-matched unrelated donors. Donors carrying 3DL1 non-expressing (null) or low-expressing alleles can produce an anti-leukemia effect and improve the overall survival rate [6] . Dubreuil et al. found through research that individuals with the centromeric KIR AA haplotype combination (cenAA) tend to carry the highly expressed 2DL1*003 and 2DL3*001 alleles, and their NK cell clones express higher levels of the corresponding KIR and have stronger NK cell activity; donor-derived cenAA can significantly reduce the recurrence of myeloid leukemia [7] .

[0012] Since identifying KIR alleles has functional and clinical significance, it is particularly important to create a high-throughput and high-resolution full-length KIR gene sequence sequencing and typing technology and conduct KIR application research at the allele level. However, the existing KIR gene polymorphism detection methods at home and abroad are not applicable to the sequencing and typing of the full-length KIR gene sequence. The reasons are as follows: (1) The sequence-specific primer-PCR (PCR-SSP), flow cytometry magnetic bead-reverse sequence-specific oligonucleotide probe (Flow-rSSO), and quantitative PCR (Q-PCR) methods widely used clinically. Commercial kits can only identify the presence or absence of each KIR gene at a low resolution level, and cannot reach the fine allele level and distinguish all non-expressing KIR alleles and variants [8, 9] .

[0013] (2) Based on the Sanger first-generation sequencing and typing technology, the applicant analyzed the problems existing in the KIR gene sequencing and typing (sequencing-based typing, PCR-SBT) methods reported at home and abroad

[10] : (1) Only partial exons of the KIR gene were sequenced and typed [11, 12, 13] . (2) The amplified product fragments and the amplification time required are too long

[14] . (3) There is non-specific amplification. For example, if the subject carries the 2DL1 and 2DS1 genes, when amplifying the sequences of exons 1 to 5 of 2DL1, the sequences of the highly homologous 2DS1 gene are also amplified simultaneously.

[15] . (4) PCR synchronous amplification and synchronous sequencing cannot be carried out under the same cycling parameters. [14, 15, 16] , making it impossible to achieve high throughput, and the experimental operation is time-consuming and laborious.

[10] .

[0014] To solve the above problems, the applicant once formulated a scientific PCR amplification strategy based on the structure of the full-length genomic sequences of each KIR gene, the distribution characteristics of single nucleotide polymorphisms in the coding region, and the lengths of the introns on both wings of the exons, designed and adopted KIR gene-specific PCR primers with similar annealing temperatures for PCR synchronous amplification, and created a method for synchronous sequencing and genotyping of 14 functional killer cell immunoglobulin-like receptor KIRs genes. [17, 18] .

[0015] However, due to the limitations of the Sanger sequencing technology itself, the existing first-generation Sanger sequencing and genotyping methods for KIR genes at home and abroad [11-18] have the following deficiencies: First, there are 16 KIR genes, and each exon of each KIR gene carried by the subject needs to be sequenced forward and backward one by one, which is time-consuming and laborious; second, the sequencing and genotyping can only cover the exon regions of each KIR gene and cannot involve the non-coding regions; third, the more alleles there are, the more ambiguous the genotyping results are; fourth, although KIR alleles can be identified at a high-resolution level, the copy number of KIR alleles cannot be determined.

[0016] (3) KIR gene NGS (Next generation sequencing) sequencing methods based on capture methods or PCR amplification [19, 20] , due to the short read lengths of next-generation sequencing (100 - 300 bp), after library construction and sequencing, there is a risk of difficulty in splicing and failure of sequencing and genotyping for the full-length sequences of highly homologous KIR genes (the full-length sequences of functional KIR genes in the IPD-KIR database are between 9901 and 17009 bp).

[20] ; It is necessary to rely on the cumbersome and complex expectation-maximization algorithm, KIR allele linkage inheritance, and bioinformatics analysis to deduce the KIR genotype; however, the result accuracy is difficult to guarantee. So far, there is still no commercial NGS KIR gene sequencing and genotyping kit at home and abroad.

[0017] The third-generation sequencing technology is a single-molecule sequencing technology with the advantage of long sequencing read length (10-100 kb). It is applied to the full-length sequence determination and typing of KIR genes. The amplification library construction method does not require fragmentation of PCR amplification products, and there is no ambiguity in the results. Compared with Sanger first-generation sequencing and NGS KIR gene sequencing typing technology, it has significant technical advantages. However, to date, there is no mature third-generation long-read sequencing typing technology for KIR genes in the world, nor is there any related commercial kit.

[0018] With the in-depth development of the biological functions, disease associations, and hematopoietic stem cell transplantation of KIR molecules, the application research field of KIR is becoming increasingly extensive. For the KIR gene family with complex structure, genetic polymorphism, and highly homologous sequences, the establishment of efficient, low-cost, high-throughput third-generation long-read sequencing typing technology and its commercialization and industrialization will help break through the technical bottlenecks in KIR gene sequencing typing at home and abroad, which is a key technical problem that needs to be solved in the current field of molecular immunogenetics and transplant immunology.

[0019] References [1] Vierra-Green C, Roe D, Hou L, et al. Allele-Level HaplotypeFrequencies and Pairwise Linkage Disequilibrium for 14 KIR Loci in 506European-American Individuals. PLoS One, 2012, 7(11): e47491. [2] Robinson J, Mistry K, McWilliam H, et al. IPD - the immunopolymorphism database. Nucleic Acids Res. 2010, 38: D863-869. [3] Jiang W, Johnson C, Simecek N, et al. qKAT: a high-throughput qPCRmethod for KIR gene copy number and haplotype determination. Genome Med.2016, 8(1):99. [4] McErlean C, Gonzalez AA, Cunningham R, et al. Differential RNAexpression of KIR alleles. Immunogenetics, 2010, 62(7): 431-440. [5] Yawata M, Yawata N, Draghi M, et al. Roles for HLA and KIRpolymorphisms in natural killer cell repertoire selection and modulation ofeffector function. J Exp Med, 2006, 203(3): 633-645. [6] Boudreau JE, Giglio F, Gooley TA, et al. KIR3DL1 / HLA-B SubtypesGovern Acute Myelogenous Leukemia Relapse After Hematopoietic CellTransplantation. J Clin Oncol. 2017, 35(20):2268-2278. [7] Dubreuil L, Maniangou B, Chevallier P, et al. Centromeric KIR AAIndividuals Harbor Particular KIR Alleles Conferring Beneficial NK CellFeatures with Implications in Haplo-Identical Hematopoietic Stem CellTransplantation. Cancers (Basel). 2020, 12(12):3595. [8] Deng ZH, Zhen JX, Zhang G, Yang ZC, Yu Q, Chen H. Establishment of a PCR-SSP method for simultaneous amplification and identification of the presence or absence of KIR genes in the Chinese population. Chinese Journal of Medical Genetics. 2023, 40(7): 881-886. [9] Li YN, Zhen JX, Liang S, Yu Q, Deng ZH. Establishment of a method for qualitative detection of the presence or absence of KIR genes by Q-PCR. Chinese Journal of Blood Transfusion. 2024, 37(6): 660-665.

[10] Zhen Jianxin, Yu Qiong. Research progress on molecular genetic polymorphism and sequencing typing technology of KIR. Chinese Journal of Medical Genetics. 2016, 33(6):867-870.

[11] Lebedeva TV, Ohashi M, Zannelli G, et al. Comprehensive approach to high-resolution KIR typing. Hum Immunol, 2007, 68(9):789-796.

[12] Yan LX, Zhu FM, Jiang K, et al. Diversity of the killer cell immunoglobulin-like receptor gene KIR2DS4 in the Chinese population. Tissue Antigens, 2007, 69(2): 133-138.

[13] Buhler S, Di Cristofaro J, Frassati C, et al. High levels of molecular polymorphism at the KIR2DL4 locus in French and Congolese populations: impact for anthropology and clinical studies. Hum Immunol. 2009, 70(11):953-959.

[14] Belle I, Hou L, Chen M, et al. Investigation of killer cell immunogloublin-like receptor gene diversity in KIR3DL1 and KIR3DS1 in a transplant population, Tissue Antigens. 2008, 71(5): 434-439.

[15] Hou L, Chen M, Steiner N, et al. Killer cell immunoglobulin-like receptors (KIR) typing by DNA sequencing. Methods Mol Biol, 2012, 882: 431-468.

[16] Meenagh A, Gonzalez A, Sleator C, et al. Investigation of killer cell immunoglobulin-like receptor gene diversity, KIR2DL1 and KIR2DS1. Tissue Antigens. 2008, 72(4):383-391.

[17] Deng Zhihui, Zhen Jianxin, Zhang Guobin. Method for Simultaneous Sequence-based Typing of 14 Functional Killer Cell Immunoglobulin-like Receptor KIRs Genes, 2018-3-20, Chinese Invention Patent, ZL 201710284545.3.

[18] Zhihui Deng, Jianxin Zhen, Guobin Zhang. Method for Simultaneous Sequence-based Typing of 14 Functional Killer Cell Immunoglobulin-like Receptor (KIR) Genes, 2019-4-23, Patent No., US 10266877 B2.

[19] Norman PJ, Hollenbach JA, Nemat-Gorgani N, et al. Defining KIR and HLA Class I Genotypes at Highest Resolution via High-Throughput Sequencing. Am J Hum Genet. 2016, 99(2):375-391.

[20] Wagner I, Schefzyk D, Pruschke J, et al. Allele-LevelKIRGenotyping of More Than a Million Samples: Workflow, Algorithm, andObservations. Front Immunol. 2018; 9:2843. Summary of the Invention The present invention aims to break through the technical bottlenecks existing in KIR gene sequencing and genotyping at home and abroad, and for the first time provides a method for genotyping all 16 killer cell immunoglobulin-like receptor (KIR) genes by third-generation long-read sequencing, which can be applied to research fields such as KIR population genetics and molecular evolution, hematopoietic stem cell transplantation donor-recipient tissue typing, and disease association, and lays a foundation for the commercialization and industrialization of KIR gene sequencing and genotyping reagents.

[0020] The present invention provides a PCR amplification primer set for the full-length sequences of 16 KIR genes, including 3 pairs of specific PCR primers, and the 3 pairs of specific PCR primers include the first pair of PCR primers, the second pair of PCR primers, and the third pair of PCR primers; The sequences of the first pair of PCR primers are shown as SEQ ID NO.1~2, the sequences of the second pair of PCR primers are shown as SEQ ID NO.3~4, and the sequences of the third pair of PCR primers are shown as SEQ ID NO.5~6; The first pair of PCR primers is used for specific multiplex PCR amplification of the full-length sequences of genes 2DL1, 2DL2, 2DL3, 2DL5, 2DS1, 2DS2, 2DS3, 2DS4, 2DS5, 3DL1, 3DL2, 3DL3, 3DS1, and 2DP1. The second pair of PCR primers is used for specific amplification of the full-length sequence of the 2DL4 gene. The third pair of PCR primers is used for specific amplification of the full-length sequence of the 3DP1 gene and a partial sequence of a KIR gene closest to the 3DP1 gene upstream of its flank.

[0021] In one embodiment, the third pair of PCR primers is used for specific amplification of the full-length sequence of the 3DP1 gene and a partial sequence of a KIR gene closest to the 3DP1 gene upstream of its flank: The KIR gene closest to the 3DP1 gene upstream of its flank includes any one of the genes 2DL1, 2DL2, 2DL3, 2DL5, 2DS2, 2DS3, 2DS5, 3DL3, and 2DP1; The partial sequence of a KIR gene closest to the 3DP1 gene includes the sequence from the fifth intron of the KIR gene closest to the 3DP1 gene to its 3'-untranslated region.

[0022] In one embodiment, the PCR amplification primer set is used for amplifying the full-length sequences of all 16 human KIR genes and library construction in the third-generation long-read sequencing genotyping of KIR genes by the long-distance PCR method; The lengths of the amplification products of the first pair of PCR primers, the second pair of PCR primers, and the third pair of PCR primers are all between 9.2 and 16.7 kb, covering all exons and introns of each KIR gene.

[0023] The present invention also provides a PCR amplification kit for the full-length sequences of 16 KIR genes, including the above-mentioned PCR amplification primer set for the full-length sequences of 16 KIR genes and a PCR amplification detection reagent.

[0024] The present invention also provides a third-generation long-read sequencing genotyping method for 16 KIR genes, including the following steps: 1) Mix the first pair of PCR primers, the second pair of PCR primers, and the third pair of PCR primers with the genomic DNA of the sample to be tested, a long-fragment rapid amplification DNA polymerase mixture, and ddH 2 O to prepare a PCR amplification reaction system to obtain a first PCR amplification reaction system, a second PCR amplification reaction system, and a third PCR amplification reaction system. The sequences of the first pair of PCR primers are as shown in SEQ ID NO.1-2, the sequences of the second pair of PCR primers are as shown in SEQ ID NO.3-4, and the sequences of the third pair of PCR primer pairs are as shown in SEQ ID NO.5-6; 2) The first PCR amplification reaction system, the second PCR amplification reaction system, and the third PCR amplification reaction system are respectively amplified under corresponding PCR conditions to obtain a first amplification product, a second amplification product, and a third amplification product; these three amplification products are mixed in a mass ratio of 30:2:3 into one tube to obtain a mixed detection PCR amplification product; 3) The mixed detection PCR amplification product is subjected to magnetic bead purification, concentration and fragment size detection, and after library construction, on-machine library preparation, on-machine sequencing, and bioinformatics analysis, the KIR genotype of the sample to be tested can be determined.

[0025] In one embodiment, the PCR amplification reaction systems are respectively: The first PCR amplification reaction system is 25.0 μL and includes the following components: 2.0 μL of 10 μM first pair of PCR primers, 2.0 μL of 50 - 100 ng / μL genomic DNA, 12.5 μL of long - fragment rapid amplification DNA polymerase mixture, and 8.5 μL of ddH 2 O; The second PCR amplification reaction system is 25.0 μL and includes the following components: 2.0 μL of 10 μM second pair of PCR primers, 2.0 μL of 50 - 100 ng / μL genomic DNA, 12.5 μL of long - fragment rapid amplification DNA polymerase mixture, and 8.5 μL of ddH 2 O; The third PCR amplification reaction system is 25.0 μL and includes the following components: 6.0 μL of 1.0 μM third pair of PCR primers, 2.0 μL of 50 - 100 ng / μL genomic DNA, 12.5 μL of long - fragment rapid amplification DNA polymerase mixture, and 4.5 μL of ddH 2 O;

[0026] In one embodiment, the conditions for PCR amplification are as follows: The reaction conditions for the first PCR amplification reaction system are: 98°C for 2 min; 98°C for 10 Sec; 59°C for 10 Sec; 72°C for 4 min, 35 cycles; 72°C for 5 min; The reaction conditions for the second PCR amplification reaction system are: 98°C for 2 min; 98°C for 10 Sec; 56°C for 10 Sec; 72°C for 4 min, 35 cycles; 72°C for 5 min; The reaction conditions for the third PCR amplification reaction system are: 98°C for 2 min; 98°C for 10 Sec; 65°C for 10 Sec; 72°C for 5 min, 35 cycles; 72°C for 5 min.

[0027] In one embodiment, the lengths of the amplification products of the first pair of PCR primers, the second pair of PCR primers, and the third pair of PCR primers are all between 9.2 - 16.7 kb, covering all exons and introns of each KIR gene.

[0028] The contribution of the present invention lies in that for the first time, a high-throughput third-generation long-read sequencing typing method for KIR genes has been established. According to the KIR allele sequence polymorphisms in the Chinese population and the structures and single nucleotide polymorphism distributions of the full-length KIR gene sequences included in the international IPD-KIR database, 3 pairs of specific primers are designed for all 16 KIR genes in the KIR gene family, and the long-range PCR (long-range PCR) method is used to amplify the full-length sequences of all 16 KIR genes in humans, as Figure 1 shown. After measuring the concentration and purity of the amplification products, they are mixed into 1 tube according to a certain mass ratio, and through library construction, on-machine sequencing (PacBio platform), and bioinformatics analysis, the KIR alleles carried by the tested subject can be determined.

[0029] The technical advantages of the present invention are as follows: Only 3 independent PCR reactions are required to amplify the full-length sequences of all 16 KIR genes, and then they are mixed into 1 tube according to a certain mass ratio for third-generation sequencing. It has the characteristics of simple experimental operation for amplification and library construction, low cost, typing covering the full-length sequences of each KIR gene, no ambiguous results, and high quality of the obtained sequences. It has significant technical advantages over the Sanger first-generation and NGS methods for KIR gene sequencing, can lay a foundation for the next-step commercialization and industrialization of this technology, and has good application prospects in the fields of KIR population genetics, disease association, hematopoietic stem cell transplantation, etc. Description of the Drawings

[0030] Figure 1 is the PCR amplification strategy for the full-length sequences of 16 KIR genes.

[0031] Figure 2 is the agarose gel electrophoresis pattern of the PCR amplification products of the sample in Example 1 (number: 13013519, KIR AA1 gene combination type).

[0032] Figure 3 is the visualization diagram of the alignment of the target sequences of all 9 KIR genes carried by the sample in Example 1 on the reference genome.

[0033] Figure 4 is the agarose gel electrophoresis pattern of the PCR amplification products of the sample in Example 2 (number: 13013484, KIR AB8 gene combination type).

[0034] Figure 5 is the visualization diagram of the alignment of the target sequences of all 13 KIR genes carried by the sample in Example 2 on the reference genome. Detailed Embodiments

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be described clearly and completely below. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0036] Since the identification of KIR alleles has functional and clinical significance, it is particularly important to create a high-throughput and high-resolution full-length KIR gene sequencing and typing technology and conduct KIR application research at the allele level. However, the existing KIR gene polymorphism detection methods at home and abroad are not applicable to the sequencing and typing of the full-length KIR gene sequence.

[0037] Third-generation sequencing technology, which is single-molecule sequencing, has the advantage of long sequencing read lengths (10 - 100 kb). When applied to the determination and typing of the full-length KIR gene sequence, the amplification library construction method does not require fragmenting the PCR amplification products and has no ambiguous results. Compared with Sanger first-generation sequencing and NGS KIR gene sequencing and typing technologies, it has significant technical advantages. However, so far, there is no mature third-generation long-read KIR gene sequencing and typing technology internationally, nor are there related commercial kits.

[0038] With the in-depth development of research on the biological functions, disease associations, and hematopoietic stem cell transplantation of KIR molecules, the application research field of KIR is becoming increasingly extensive. For the KIR gene family with complex structures, genetic polymorphisms, and highly homologous sequences, establishing an efficient, low-cost, and high-throughput third-generation long-read sequencing and typing technology and realizing commercialization and industrialization are conducive to breaking through the technical bottlenecks in KIR gene sequencing and typing at home and abroad, which is a key technical problem that urgently needs to be solved in the current fields of molecular immunogenetics and transplantation immunology.

[0039] In view of this, the present invention provides a "third-generation long-read sequencing genotyping method for 16 KIR genes". The specific PCR amplification primer set consists of 3 pairs. The sequences of each pair of PCR primers, the positions of the primers in the full-length KIR genomic sequence, and the lengths of the amplified target fragments and other characteristics are as described in Tables 3, 4, and 5. The long-range PCR method is used to amplify the full-length sequences of all 16 KIR genes in humans. The first pair of PCR primers specifically and multiplexly amplifies the full-length sequences of 14 KIR genes except 2DL4 and 3DP1 (including 2DL1, 2DL2, 2DL3, 2DL5, 2DS1, 2DS2, 2DS3, 2DS4, 2DS5, 3DL1, 3DL2, 3DL3, 3DS1, and 2DP1). The second pair of PCR primers specifically amplifies the full-length sequence of the 2DL4 gene. The third pair of PCR primers specifically amplifies the full-length sequence of the 3DP1 gene and a partial sequence of a KIR gene closest to the 3DP1 gene upstream of its flank. The PCR amplification strategy is detailed in Figure 1 .

[0040] Table 3 Positions where the first pair of PCR primers bind to each KIR gene, the regions where the primers are located, and the lengths of the amplification products

[0041] Table 4 Positions where the second pair of PCR primers bind to the 2DL4 gene, the regions where the primers are located, and the lengths of the amplification products

[0042] Table 5 Positions where the third pair of PCR primers bind to each KIR gene, the regions where the primers are located, and the lengths of the amplification products

[0043] The sequences of the 3 pairs of specific PCR primers, the positions of the primers in the full-length KIR genomic sequence, the lengths of the amplified target fragments, and the covered regions specifically include: 1) The first pair of PCR primers Forward primer sequence: 5'-AAKAGCCTGCRTACKTCAYCCTC-3' (SEQ ID NO.1); Reverse primer sequence: 5'-GTTGGAGAGGTGGGCAGGGGTCA-3' (SEQ ID NO.2).

[0044] Amplify the full-length sequences of 14 KIR genes except 2DL4 and 3DP1 (including 2DL1, 2DL2, 2DL3, 2DL5, 2DS1, 2DS2, 2DS3, 2DS4, 2DS5, 3DL1, 3DL2, 3DL3, 3DS1, and 2DP1). The positions of the primer pairs in the full-length sequences of the 14 KIR genes, as well as the lengths of the amplified target fragments and other characteristics are as follows: ⅰ) For the 2DL1 gene: The forward primer is located at nt-129 to nt-107 in the full-length sequence of the 2DL1 gene, in the 5'-UTR region; the reverse primer is located at nt14228 to nt14206 in the full-length sequence of the 2DL1 gene, in the 3'-UTR region; the length of the target amplified fragment is 14357 bp, covering all exons and all intron sequences of the 2DL1 gene.

[0045] ⅱ) For the 2DL2 gene: The forward primer is located at nt-129 to nt-107 in the full-length sequence of the 2DL2 gene, in the 5'-UTR region; the reverse primer is located at nt14270 to nt14248 in the full-length sequence of the 2DL2 gene, in the 3'-UTR region; the length of the target amplified fragment is 14399 bp, covering all exons and all intron sequences of the 2DL2 gene.

[0046] ⅲ) For the 2DL3 gene: The forward primer is located at nt-129 to nt-107 in the full-length sequence of the 2DL3 gene, in the 5'-UTR region; the reverse primer is located at nt14251 to nt14229 in the full-length sequence of the 2DL3 gene, in the 3'-UTR region; the length of the target amplified fragment is 14380 bp, covering all exons and all intron sequences of the 2DL3 gene.

[0047] ⅳ) For the 2DL5 gene: The forward primer is located at nt-129 to nt-107 in the full-length sequence of the 2DL5 gene, in the 5'-UTR region; the reverse primer is located at nt9168 to nt9146 in the full-length sequence of the 2DL5 gene, in the 3'-UTR region; the length of the target amplified fragment is 9297 bp, covering all exons and all intron sequences of the 2DL5 gene.

[0048] ⅴ) For the 2DS1 gene: The forward primer is located at nt-129 to nt-107 in the full-length sequence of the 2DS1 gene, in the 5'-UTR region; the reverse primer is located at nt14232 to nt14210 in the full-length sequence of the 2DS1 gene, in the 3'-UTR region; the length of the target amplified fragment is 14361 bp, covering all exons and all intron sequences of the 2DS1 gene.

[0049] vi) For the 2DS2 gene: The forward primer is located at positions nt-129 to nt-107 in the full-length 2DS2 genomic sequence, within the 5'-UTR region; the reverse primer is located at positions nt14056 to nt14034 in the full-length 2DS2 genomic sequence, within the 3'-UTR region; the length of the target amplification fragment is 14185 bp, covering all exons and all intron sequences of the 2DS2 gene.

[0050] vii) For the 2DS3 gene: The forward primer is located at positions nt-130 to nt-108 in the full-length 2DS3 genomic sequence, within the 5'-UTR region; the reverse primer is located at positions nt14562 to nt14540 in the full-length 2DS3 genomic sequence, within the 3'-UTR region; the length of the target amplification fragment is 14692 bp, covering all exons and all intron sequences of the 2DS3 gene.

[0051] viii) For the 2DS4 gene: The forward primer is located at positions nt-129 to nt-107 in the full-length 2DS4 genomic sequence, within the 5'-UTR region; the reverse primer is located at positions nt15604 to nt15582 in the full-length 2DS4 genomic sequence, within the 3'-UTR region; the length of the target amplification fragment is 15733 bp, covering all exons and all intron sequences of the 2DS4 gene.

[0052] ix) For the 2DS5 gene: The forward primer is located at positions nt-130 to nt-108 in the full-length 2DS5 genomic sequence, within the 5'-UTR region; the reverse primer is located at positions nt14738 to nt14716 in the full-length 2DS5 genomic sequence, within the 3'-UTR region; the length of the target amplification fragment is 14868 bp, covering all exons and all intron sequences of the 2DS5 gene.

[0053] x) For the 3DL1 gene: The forward primer is located at positions nt-129 to nt-107 in the full-length 3DL1 genomic sequence, within the 5'-UTR region; the reverse primer is located at positions nt14039 to nt14017 in the full-length 3DL1 genomic sequence, within the 3'-UTR region; the length of the target amplification fragment is 14168 bp, covering all exons and all intron sequences of the 3DL1 gene.

[0054] xi) For the 3DL2 gene: The forward primer is located at positions nt-129 to nt-107 in the full-length 3DL2 genomic sequence, which is in the 5'-UTR region; the reverse primer is located at positions nt16491 to nt16469 in the full-length 3DL2 genomic sequence, which is in the 3'-UTR region; the length of the target amplification fragment is 16620 bp, covering all exons and all intron sequences of the 3DL2 gene.

[0055] xii) For the 3DL3 gene: The forward primer is located at positions nt-124 to nt-102 in the full-length 3DL3 genomic sequence, which is in the 5'-UTR region; the reverse primer is located at positions nt11873 to nt11851 in the full-length 3DL3 genomic sequence, which is in the 3'-UTR region; the length of the target amplification fragment is 11997 bp, covering all exons and all intron sequences of the 3DL3 gene.

[0056] xiii) For the 3DS1 gene: The forward primer is located at positions nt-129 to nt-107 in the full-length 3DS1 genomic sequence, which is in the 5'-UTR region; the reverse primer is located at positions nt14423 to nt14401 in the full-length 3DS1 genomic sequence, which is in the 3'-UTR region; the length of the target amplification fragment is 14552 bp, covering all exons and all intron sequences of the 3DS1 gene.

[0057] xiv) For the 2DP1 gene: The forward primer is located at positions nt-129 to nt-107 in the full-length 2DP1 genomic sequence, which is in the 5'-UTR region; the reverse primer is located at positions nt12619 to nt12597 in the full-length 2DP1 genomic sequence, which is in the 3'-UTR region; the length of the target amplification fragment is 12748 bp, covering the full-length sequence from the 5'-UTR region to the 3'-UTR region of the 2DP1 gene.

[0058] 2) The second pair of PCR primers Forward primer sequence: 5'-CRTTGCGCATGATGTGAASTGAC-3' (SEQ ID NO.3); Reverse primer sequence: 5'-ATGRRGGSAGACATGTTTAYTTGAA-3' (SEQ ID NO.4).

[0059] Specifically amplify the 2DL4 gene. The forward primer is located at nt-228 to nt-206 in the full-length genomic sequence of 2DL4, within the 5'-UTR region; the reverse primer is located at nt10910 to nt10886 in the full-length genomic sequence of 2DL4, within the 3'-UTR region; the length of the target amplified fragment is 11138 bp, covering all exons and all intron sequences of the 2DL4 gene.

[0060] The forward primer sequence of the third pair of PCR primers: 5'-ACASRGCACCTCSAARCCCTCYT-3' (SEQ ID NO.5); The reverse primer sequence: 5'-ACAAGCAGTGGGTCACTAAGGTC-3' (SEQ ID NO.6).

[0061] Specifically amplify the full-length sequence of the 3DP1 gene and a partial sequence of the first KIR gene upstream of its flank. The full length of the PCR amplification product is between 13932 and 16518 bp (see Table 5 for details). The purpose of such amplification is to avoid the preferential library construction of the full-length sequence of the 3DP1 gene (5713 bp) due to its short fragment length, which may affect the library construction and sequencing of the target fragments of other KIR genes.

[0062] The positions of this primer pair in the full-length sequence of each KIR gene it binds to, the length of the amplified target fragment, and the covered regions are specifically as follows: ⅰ) 2DL1-3DP1: If the first KIR gene upstream of the 3DP1 gene flank is the 2DL1 gene, the forward primer is located at nt5753 to nt5775 in the full-length genomic sequence of 2DL1, within the 5th intron; the reverse primer is located at nt5667 to nt5645 in the full-length genomic sequence of 3DP1, at the end of the 5th exon; the length of the target amplified fragment is 16405 bp, covering a partial sequence of the 5th intron and the sequence from the 6th exon to the 3'-UTR region of the 2DL1 gene, the sequence between the 2DL1 and 3DP1 genes, and the full-length sequence of the 3DP1 gene.

[0063] ⅱ) 2DL2-3DP1: If the first KIR gene upstream of the 3DP1 gene flank is the 2DL2 gene, the forward primer is located at nt5670 to nt5692 in the full-length genomic sequence of 2DL2, within the 5th intron; the reverse primer is located at nt5667 to nt5645 in the full-length genomic sequence of 3DP1, at the end of the 5th exon; the length of the target amplified fragment is 16518 bp, covering a partial sequence of the 5th intron and the sequence from the 6th exon to the 3'-UTR region of the 2DL2 gene, the sequence between the 2DL2 and 3DP1 genes, and the full-length sequence of the 3DP1 gene.

[0064] iii) 2DL3-3DP1: If the first KIR gene upstream of the 3DP1 gene flank is the 2DL3 gene, the forward primer is located at nt5672 - nt5694 in the full-length 2DL3 gene sequence, which is in the 5th intron; the reverse primer is located at nt5667 - nt5645 in the full-length 3DP1 gene sequence, which is at the end of the 5th exon; the length of the target amplification fragment is 16509 bp, covering part of the 5th intron sequence of the 2DL3 gene, the sequence from the 6th exon to the 3'-UTR region, the sequence between the 2DL3 and 3DP1 genes, and the full-length sequence of the 3DP1 gene.

[0065] iv) 2DL5-3DP1: If the first KIR gene upstream of the 3DP1 gene flank is the 2DL5 gene, the forward primer is located at nt3166 - nt3188 in the full-length 2DL5 gene sequence, which is in the 5th intron; the reverse primer is located at nt5667 - nt5645 in the full-length 3DP1 gene sequence, which is at the end of the 5th exon; the length of the target amplification fragment is 13932 bp, covering part of the 5th intron sequence of the 2DL5 gene, the sequence from the 6th exon to the 3'-UTR region, the sequence between the 2DL5 and 3DP1 genes, and the full-length sequence of the 3DP1 gene.

[0066] v) 2DS2-3DP1: If the first KIR gene upstream of the 3DP1 gene flank is the 2DS2 gene, the forward primer is located at nt5558 - nt5580 in the full-length 2DS2 gene sequence, which is in the 5th intron; the reverse primer is located at nt5667 - nt5645 in the full-length 3DP1 gene sequence, which is at the end of the 5th exon; the length of the target amplification fragment is 16407 bp, covering part of the 5th intron sequence of the 2DS2 gene, the sequence from the 6th exon to the 3'-UTR region, the sequence between the 2DS2 and 3DP1 genes, and the full-length sequence of the 3DP1 gene.

[0067] vi) 2DS3-3DP1: If the first KIR gene upstream of the 3DP1 gene flank is the 2DS3 gene, the forward primer is located at nt6066 - nt6088 in the full-length 2DS3 gene sequence, which is in the 5th intron; the reverse primer is located at nt5667 - nt5645 in the full-length 3DP1 gene sequence, which is at the end of the 5th exon; the length of the target amplification fragment is 16426 bp, covering part of the 5th intron sequence of the 2DS3 gene, the sequence from the 6th exon to the 3'-UTR region, the sequence between the 2DS3 and 3DP1 genes, and the full-length sequence of the 3DP1 gene.

[0068] ⅶ) 2DS5-3DP1: If the first KIR gene upstream of the 3DP1 gene flank is the 2DS5 gene, the forward primer is located at nt6260 - nt6282 in the full-length 2DS5 gene sequence, in the 5th intron; the reverse primer is located at nt5667 - nt5645 in the full-length 3DP1 gene sequence, at the end of the 5th exon; the length of the target amplified fragment is 16408 bp, covering part of the 5th intron sequence of the 2DS5 gene, the sequence from the 6th exon to the 3'-UTR region, the sequence between the 2DS5 and 3DP1 genes, and the full-length sequence of the 3DP1 gene.

[0069] ⅷ) 3DL3-3DP1: If the first KIR gene upstream of the 3DP1 gene flank is the 3DL3 gene, the forward primer is located at nt5397 - nt5419 in the full-length 3DL3 gene sequence, in the 5th intron; the reverse primer is located at nt5667 - nt5645 in the full-length 3DP1 gene sequence, at the end of the 5th exon; the length of the target amplified fragment is 14385 bp, covering part of the 5th / 6th intron sequence of the 3DL3 gene, the sequence from the 7th exon to the 3'-UTR region, the sequence between the 3DL3 and 3DP1 genes, and the full-length sequence of the 3DP1 gene.

[0070] ⅸ) 2DP1-3DP1: If the first KIR gene upstream of the 3DP1 gene flank is the 2DP1 gene, the forward primer is located at nt5851 - nt5873 in the full-length 2DP1 gene sequence, in the 5th intron; the reverse primer is located at nt5667 - nt5645 in the full-length 3DP1 gene sequence, at the end of the 5th exon; the length of the target amplified fragment is 14698 bp, covering part of the 5th intron sequence of the 2DP1 gene, the sequence from the 6th exon to the 3'-UTR region, the sequence between the 2DP1 and 3DP1 genes, and the full-length sequence of the 3DP1 gene.

[0071] The lengths of the amplification products of the first pair of PCR primers, the second pair of PCR primers, and the third pair of PCR primers are all between 9.2 - 16.7 kb.

[0072] The PCR amplification primer set for the full-length sequences of 16 KIR genes in the present invention (including 3 pairs of KIR gene-specific primers) is designed based on the KIR allele sequence polymorphisms in the Chinese population and the structure and single nucleotide polymorphism distribution of the full-length KIR gene sequences included in the international IPD-KIR database, and is suitable for the Chinese population; in the third-generation long-read sequencing typing of KIR genes, it is suitable for specifically amplifying the full-length sequences of 16 KIR genes and library construction.

[0073] In the present invention, the PCR amplification detection reagent for the full-length sequences of 16 KIR genes includes: specific PCR primer pairs, a long-fragment rapid amplification DNA polymerase mixture, genomic DNA, and ddH 2 O.

[0074] The present invention provides a third-generation long-read sequencing genotyping method for 16 KIR genes, preferably including the following steps: 1) Mix the first pair of PCR primers, the second pair of PCR primers, and the third pair of PCR primers with the genomic DNA of the sample to be tested, the long-fragment rapid amplification DNA polymerase mixture, and ddH 2 O respectively to prepare a PCR amplification reaction system, obtaining the first PCR amplification reaction system, the second PCR amplification reaction system, and the third PCR amplification reaction system; 2) The first PCR amplification reaction system, the second PCR amplification reaction system, and the third PCR amplification reaction system are respectively amplified under corresponding PCR conditions to obtain the first amplification product, the second amplification product, and the third amplification product; these three amplification products are mixed into one tube according to a certain mass ratio (30:2:3) to obtain a mixed PCR amplification product; 3) The mixed PCR amplification product is subjected to magnetic bead purification, concentration and fragment size detection, and through library construction, on-machine library preparation, on-machine sequencing, and bioinformatics analysis, the KIR genotype of the sample to be tested can be determined.

[0075] The technical advantages of the present invention are as follows: Only 3 independent PCR reactions are required to amplify the full-length sequences of all 16 KIR genes, and then they are mixed into 1 tube according to a certain mass ratio for third-generation sequencing. It has the characteristics of simple experimental operation for amplification and library construction, low cost, the genotype covering the full-length sequences of each KIR gene, no ambiguous results, and high quality of the obtained sequences. It has significant technical advantages compared with the Sanger first-generation and NGS methods for KIR gene sequencing.

[0076] The technical solution of the present invention will be further described in detail below in conjunction with specific embodiments. It should be understood that the following embodiments are only used to explain the present invention and are not used to limit the present invention.

[0077] Implementation materials 1. Long-fragment rapid amplification DNA polymerase mixture: Purchased from Jiangsu Kangwei Century Biotechnology Co., Ltd., reagent production batch number: 30424.

[0078] 2. Samples with known KIR AA1 and KIR AB8 gene combinations: Derived from unpaid volunteer blood donors in Shenzhen, these are the remaining blood samples (5 mL per person) from the peripheral blood samples of blood donors after laboratory tests, anticoagulated with 5% EDTA, and collected during the applicant's previous undertaking of a general project of the National Natural Science Foundation of China (Project No.: 81373158, already completed). The applicant had previously used the Q-PCR method for qualitative detection of the presence or absence of KIR genes at a low-resolution level.

[0079] 3. Samples of 20 healthy unrelated individuals: Derived from unpaid volunteer blood donors in Shenzhen, these are the remaining blood samples (5 mL per person) from the peripheral blood samples of blood donors after laboratory tests, anticoagulated with 5% EDTA, and collected during the applicant's previous undertaking of a general project of the National Natural Science Foundation of China (Project No.: 81373158, already completed). The applicant had previously performed whole-genome sequencing on these samples and then carried out KIR gene typing after extracting KIR-related sequences.

[0080] Example 1 A method for third-generation long-read sequencing typing of the full-length sequences of 16 KIR genes suitable for the Chinese population In previous work, the applicant used the "Method for Synchronous Sequencing Typing of 14 Functional Killer Cell Immunoglobulin-like Receptor KIRs Genes" established through independent innovation (Chinese invention patent authorization patent number ZL 201710284545.3, US invention patent authorization patent number US 10266877 B2, see references

[17] and

[18] in the "Background Technology" section for details) to determine the sequences of the entire coding regions of 14 functional KIR genes in the Chinese population (n = 306) and detected 116 KIR alleles. [21, 22, 23] . As of January 2022, the applicant had previously discovered and identified 75 new alleles recognized by the World Health Organization (WHO) (for information on KIR new alleles, see Table 6). Based on the KIR allele sequence polymorphisms in the Chinese population and with reference to the full-length sequences of KIR genes included in the international IPD-KIR database, for all 16 KIR genes in the KIR gene family, 3 pairs of specific PCR primers were designed, and the long-range PCR method was used to amplify the full-length sequences of all 16 KIR genes in humans. The length of the amplified product fragments ranged from 9.2 to 16.7 kb. The sequences, positions of each PCR primer, and the length of the amplified target fragment and other characteristics are shown in Tables 3 to 5.

[0081] Table 6: 75 KIR new alleles discovered and identified by the applicant previously and recognized by the WHO

[0082]

[0083]

[0084]

[0085] The PCR amplification detection reagent for the full-length sequences of the 16 KIR genes described above includes: specific PCR primer pairs, a long-fragment rapid amplification DNA polymerase mixture, genomic DNA, and ddH 2 O.

[0086] In this example, a sample randomly selected for the detection of the presence or absence of KIR genes at the low-resolution level by Q-PCR method, with the detection result being the KIR AA1 gene combination type (carrying a total of 9 KIR genes, namely 2DL1, 2DL3, 2DL4, 2DS4, 3DL1, 3DL2, 3DL3, 2DP1, and 3DP1), was used to perform third-generation long-read sequencing typing on all KIR genes carried by this sample to verify the implementation effect of the present invention.

[0087] ⅰ) First, the present invention respectively used 3 pairs of KIR gene-specific primers to amplify the full-length sequences of all KIR genes carried by this sample. The amplification reaction was carried out on an ABI Veriti PCR instrument, and the system compositions of the PCR amplification reactions were respectively: The first pair of PCR primers was used for specific multiplex PCR amplification of 14 KIR genes except 2DL4 and 3DP1, and the second pair of PCR primers was used to amplify the 2DL4 gene. The system compositions of their PCR amplification reactions were both: 12.5 μL of long-fragment rapid amplification DNA polymerase mixture; 1.0 μL of each 10 μM PCR primer; 2.0 μL of genomic DNA at 50 - 100 ng / μL; Add ddH2O to 25.0 μL.

[0088] The third pair of PCR primers was used for specific amplification of the full-length sequence of the 3DP1 gene and a partial sequence of the first KIR gene upstream of its flank. The system composition of its PCR amplification reaction was: 12.5 μL of long-fragment rapid amplification DNA polymerase mixture; 3.0 μL of each 1.0 μM PCR primer; 2.0 μL of genomic DNA at 50 - 100 ng / μL; Add ddH2O to 25.0 μL.

[0089] The PCR amplification conditions were respectively: The first pair of PCR primers specifically amplifies 14 KIR genes except 2DL4 and 3DP1, and the amplification conditions are: 98°C for 2 min; 98°C for 10 Sec; 59°C for 10 Sec; 72°C for 4 min, 35 cycles; 72°C for 5 min; 4°C Infinite.

[0090] The second pair of PCR primers specifically amplifies the 2DL4 gene, and the amplification conditions are: 98°C for 2 min; 98°C for 10 Sec; 56°C for 10 Sec; 72°C for 4 min, 35 cycles; 72°C for 5 min; 4°C Infinite.

[0091] The third pair of PCR primers specifically amplifies the 3DP1 gene and a partial sequence of the first KIR gene upstream of its flank, and the amplification conditions are: 98°C for 2 min; 98°C for 10 Sec; 65°C for 10 Sec; 72°C for 5 min, 35 cycles; 72°C for 5 min; 4°C Infinite.

[0092] ii) Electrophoresis of PCR amplification products: Take 10.0 μL of PCR products, stain them with nucleic acid dye, and perform electrophoresis in 0.8% agarose gel. Referring to the Takara DL15000 DNA molecular marker, observe the bands of specific PCR amplification products in the gel imaging system. Clear and bright target bands are present in the products amplified by the first, second, and third pairs of PCR primers. See the PCR product electrophoresis diagram in Figure 2 。 Figure 2 In which M: DL15000 Marker; A1: Amplicon 1, the full-length sequences of 14 KIR genes specifically amplified by multiplex PCR with the first pair of PCR primers except 2DL4 and 3DP1; A2: Amplicon 2, the full-length sequence of the 2DL4 gene specifically amplified by the second pair of PCR primers; A3: the full-length sequence of the 3DP1 gene specifically amplified by the third pair of PCR primers and a partial sequence of the first KIR gene upstream of its flank.

[0093] iii) Quality inspection of PCR amplification products: Take 1.0 μL of the products amplified by the first, second, and third pairs of PCR primers respectively, and detect the purity by Nanodrop and the concentration by Qubit 4.0 instrument.

[0094] iv) Mixing and purification of the amplification products: The products amplified by the first, second, and third pairs of PCR primers are mixed into one tube according to a certain mass ratio (30:2:3); after mixing, 1.0 μg of the amplification product is taken, and the total volume is supplemented to 100 μL with deionized water, and purified twice by the AMpure magnetic bead method; the concentration of the purified amplification product is detected by Qubit 4.0.

[0095] v) Library construction: Take 300 - 500 ng of the amplification product and construct a library using the PacBio SMRTEBLL @ PREP KIT 3.0 kit.

[0096] First, perform end repair and A-tailing on the amplification product. The reaction system contains 8.0 μL of Repair buffer, 4.0 μL of End Repair Mix, 2.0 μL of DNA Repair mix, 40.0 μL of the product, and 6.0 μL of RNase-free water, for a total of 60.0 μL of reaction mixture; incubate at 37°C for 30 minutes and at 65°C for 5 minutes; make the target amplification product form a sticky end with a dA tail, which is convenient for ligation with the adapter with a T tail.

[0097] Adapter ligation: The reaction system contains 4.0 μL of SMRTbell adapter, 30.0 μL of Ligation mix, 1.0 μL of Ligation enhancer, and 60.0 μL of the product after end repair and A-tailing, for a total of 95.0 μL of reaction mixture; incubate at 20°C for 30 minutes. After adding 95.0 μL of SMRTbell cleanup beads to the adapter ligation product for purification, add 40.0 μL of buffer for elution.

[0098] The third step is nuclease treatment: The reaction system contains 5.0 μL of Nuclease buffer, 5.0 μL of Nuclease mix, and 40.0 μL of the purified adapter ligation product, for a total of 50.0 μL of reaction mixture; incubate at 37°C for 15 minutes. After adding 50.0 µl of AMpure magnetic beads to the digested product for purification, add 15.0 μL of buffer for elution.

[0099] Finally, perform library quality inspection and detect the library concentration using a Qubit instrument.

[0100] vi) Library processing before loading: Mix the library of the samples in this example with the libraries of other samples at an equal amount in ng, add AMpure magnetic beads with a volume 0.6 times that of the sampling volume of the mixed library for purification, and recover all target large fragments.

[0101] vii) Library preparation for loading: Take out approximately 70 ng of the mixed amplified library, and use the SEQUEL ® II BINDING KIT 3.2 kit to prepare the library for loading. Add the sequencing primer in this kit and incubate at 20 °C for 30 min to allow the sequencing primer to bind to the adapter sequence position of the library; then add the sequencing enzyme in this kit and incubate at 30 °C for 30 min to fix the sequencing enzyme to the sequencing primer; the prepared product is purified using the magnetic beads in the kit to obtain the library for loading.

[0102] viii) Sequencing on the machine: Take 100 pmol / L of the product prepared for loading and place it on the PacBio Sequel II high-throughput gene sequencer for sequencing. Use the SEQUEL II SEQUENCING KIT 2.0 kit and the PacBio SMRT Cell 8M chip to complete the sequencing reaction process on the machine.

[0103] ix) Bioinformatics analysis and KIR gene typing: The first step is data preprocessing: After sequencing, the system (SMRTLink, v13.0, Pacific Bioscience) configured on the PacBio Sequel II sequencer automatically converts the fluorescence signal into base sequence information, and at the same time performs in-single-molecule sequence correction and subsequent filtering to obtain HIFI reads with a single-molecule accuracy above Q20, that is, a BAM file is formed.

[0104] The second step is data splitting: Use the PacBio official built-in software lima to split the HIFI reads to distinguish different samples. The base number, average read length, and sequence quality of the HIFI reads of the 13013519th sample in this example are shown in Table 7.

[0105] Table 7 Base number, average read length, and sequence quality of the HIFI reads of the 13013519th sample in Example 1

[0106] Step 3 Data comparison: Use the official built-in software pbMM2 of PacBio to compare with the 38th version of the genome, and cooperate with software such as samtools and bedtools to extract the gene sequences on the target chromosome to generate a conventional fasta file. The sample No. 13013519 in this Example 1 is of the KIR AA1 gene combination type. The read lengths and the number of reads of the 9 KIR genes it carries are shown in Table 8.

[0107] Table 8 Statistics of the off-machine data of sample No. 13013519 in Example 1

[0108] Step 4 Data filtering: Using the same method as above, align the above fasta sequence data to the reference sequences of each specific KIR gene, and filter out the sequences of non-target genes. The comparison effect is shown in Figure 3 . It can be seen from the visualization Figure 3 that: The target sequences of all 9 KIR genes (2DL1, 2DL3, 2DL4, 2DS4, 3DL1, 3DL2, 3DL3, 2DP1, and 3DP1) carried by the sample in Example 1 can be clearly aligned to the reference genome.

[0109] Step 5 Extraction of coding region sequence (CDS): Use a third-party software such as sam2tsv to extract the CDS sequences of KIR genes.

[0110] Step 6 Genotyping: Align the extracted CDS sequences with the CDS sequences of all KIR alleles in the IPD-KIR database. If the detected CDS sequences are 100% consistent with a certain genotyping result in the database and reach a certain depth, they will be classified into that genotype.

[0111] The KIR genotype of sample No. 13013519 in this Example 1 can be determined as: 3DL3*01001,01002 - 2DL3*00101 - 2DP1*00201 - 2DL1*00302 - 3DP1*00302 - 2DL4*00102 - 3DL1*01502 - 2DS4*00101 - 3DL2*00201.

[0112] References:

[21] Deng Z, Zhen J, Harrison GF, et al. Adaptive Admixture of HLA Class I Allotypes Enhanced Genetically Determined Strength of Natural Killer Cells in East Asians. Mol Biol Evol. 2021, 38(6):2582-2596.

[22] Deng Z, Zhen J, Zhu B, Zhang G, Yu Q, Wang D, Xu Y, He L, Lu L. Allelic diversity of KIR3DL1 / 3DS1 in a southern Chinese population. Hum Immunol. 2015, 76(9):663-666.

[23] Zhen J, He L, Xu Y, et al. Allelic polymorphism of KIR2DL2 / 2DL3 in a southern Chinese population. Tissue Antigens. 2015, 86(5):362-367. Example 2 In this example, a sample with the KIR AB8 genotype (carrying 13 KIR genes including 2DL1, 2DL3, 2DL4, 2DL5, 2DS1, 2DS3, 2DS4, 3DL1, 3DL2, 3DL3, 3DS1, 2DP1, and 3DP1) was randomly selected after low-resolution detection of the presence or absence of KIR genes by Q-PCR method. The third-generation long-read sequencing genotyping was performed on all KIR genes carried by this sample using the present invention to verify the implementation effect of the present invention.

[0113] The experimental method of this Example 2 (including PCR amplification, electrophoresis and quality inspection of amplification products, library construction, preparation of the library for sequencing, sequencing on the machine, bioinformatics analysis, and KIR gene genotyping) is consistent with the experimental method of Example 1.

[0114] For the electrophoresis diagram of the PCR products in this Example 2, see Figure 4 . Figure 4Among them, M: DL15000 Marker; A1: Amplicon1, the first pair of PCR primers specifically amplify the full-length sequences of 14 KIR genes except 2DL4 and 3DP1 by multiplex PCR; A2: Amplicon 2, the second pair of PCR primers specifically amplify the full-length sequence of the 2DL4 gene; A3: The third pair of PCR primers specifically amplify the full-length sequence of the 3DP1 gene and a partial sequence of the first KIR gene upstream of its flank.

[0115] In bioinformatics analysis, the base number, average read length, and sequence quality of the HIFI reads of sample No. 13013484 in Example 2 are shown in Table 9.

[0116] Table 9 Base number, average read length, and sequence quality of HIFI reads of sample No. 13013484 in Example 2

[0117] In bioinformatics analysis, the sample No. 13013484 in Example 2 is of the KIR AB8 gene combination type. The read lengths and read numbers of the 13 KIR genes it carries are shown in Table 10.

[0118] Table 10 Statistics of the off-machine data of sample No. 13013484 in Example 2

[0119] The fasta sequence data of Example 2 were aligned to the reference sequences of each specific KIR gene, and the sequences of non-target genes were filtered out. The alignment results are shown in Figure 5 . It can be seen from the visualization Figure 5 that: The target sequences of all 13 KIR genes (2DL1, 2DL3, 2DL4, 2DL5, 2DS1, 2DS3, 2DS4, 3DL1, 3DL2, 3DL3, 3DS1, 2DP1, 3DP1) carried by the sample in Example 2 can be clearly aligned to the reference genome.

[0120] According to the alignment results of the extracted CDS sequences with the CDS sequences of all KIR alleles in the IPD-KIR database, the KIR genotype of sample No. 13013484 in Example 2 can be determined as: 3DL3*00802,01002-2DL3*00101-2DP1*00201-2DL1*00302,06901-3DP1*00302-2DL4*00103,00501-3DL1*02001-3DS1*085-2DL5A*00501-2DS3*00201-2DS1*00201-2DS4*00101-3DL2*00201,00701.

[0121] Example 3 In this example, samples from 20 healthy unrelated individuals were used for KIR gene third-generation long-read sequencing genotyping using the present invention. These samples had already been genotyped for KIR genes after whole-genome sequencing and extraction of KIR-related sequences (hereinafter referred to as "KIR gene genotyping by whole-genome sequencing method"). The results of KIR gene amplicon sequencing genotyping of the present invention were blindly compared with the results of "KIR gene genotyping by whole-genome sequencing method" to verify the implementation effect of the present invention.

[0122] The experimental method of this Example 3 (including PCR amplification, electrophoresis and quality inspection of amplification products, library construction, preparation of the library for sequencing, sequencing on the machine, bioinformatics analysis and KIR gene genotyping) was consistent with the experimental method of Example 1. The difference was that the test samples in this Example 3 were samples from 20 healthy unrelated individuals.

[0123] The KIR genotypes of the 20 samples in this Example 3 are shown in detail in Table 11. The results of KIR gene amplicon sequencing genotyping using the present invention were consistent with the results of "KIR gene genotyping by whole-genome sequencing method".

[0124] Table 11 Results of third-generation long-read sequencing genotyping of 16 KIR genes in the population (n = 20 cases)

[0125]

Claims

1. A PCR amplification primer set for 16 full-length sequences of KIR genes, characterized in that: Comprising 3 pairs of specific PCR primers, wherein the 3 pairs of specific PCR primers include a first pair of PCR primers, a second pair of PCR primers and a third pair of PCR primers; The sequences of the first pair of PCR primers are shown in SEQ ID NOs. 1-2, the sequences of the second pair of PCR primers are shown in SEQ ID NOs. 3-4, and the sequences of the third pair of PCR primers are shown in SEQ ID NOs. 5-6; The first pair of PCR primers is used for specific multiplex PCR amplification of the full-length sequences of 2DL1, 2DL2, 2DL3, 2DL5, 2DS1, 2DS2, 2DS3, 2DS4, 2DS5, 3DL1, 3DL2, 3DL3, 3DS1 and 2DP1 genes, the second pair of PCR primers is used for specific amplification of the full-length sequence of the 2DL4 gene, and the third pair of PCR primers is used for specific amplification of the full-length sequence of the 3DP1 gene and the partial sequence of a KIR gene on its flank upstream that is closest to the 3DP1 gene.

2. The PCR amplification primer set for the full-length sequences of 16 KIR genes according to claim 1, characterized in that: The third pair of PCR primers is used to specifically amplify the full-length sequence of the 3DP1 gene and a partial sequence of a KIR gene that is closest to the 3DP1 gene upstream of the flank: The KIR gene whose flanking upstream is closest to the 3DP1 gene includes any one of 2DL1, 2DL2, 2DL3, 2DL5, 2DS2, 2DS3, 2DS5, 3DL3 and 2DP1 genes; The partial sequence of the KIR gene closest to the 3DP1 gene includes the sequence from the fifth intron to the 3'-untranslated region of the KIR gene closest to the 3DP1 gene.

3. The PCR amplification primer set for the full-length sequences of 16 KIR genes according to claim 1, characterized in that: The PCR amplification primer set is used for amplifying the full-length sequences of all 16 human KIR genes and library construction by long-distance PCR method in the third-generation long-read sequencing typing of KIR genes; The lengths of the amplified products of the first pair of PCR primers, the second pair of PCR primers and the third pair of PCR primers are all between 9.2 and 16.7 kb, covering all exons and introns of each KIR gene.

4. A PCR amplification kit for the full-length sequences of 16 KIR genes, characterized in that: It comprises a PCR amplification primer set and a PCR amplification detection reagent for the full-length sequences of 16 KIR genes as described in any one of claims 1 to 3.

5. A method for genotyping 16 KIR genes using third-generation long-read sequencing, characterized in that: The following steps are involved: 1) The first pair of PCR primers, the second pair of PCR primers and the third pair of PCR primers are respectively mixed with the genomic DNA of the sample to be tested, the long fragment rapid amplification DNA polymerase mixture and ddH2O to prepare a PCR amplification reaction system to obtain a first PCR amplification reaction system, a second PCR amplification reaction system and a third PCR amplification reaction system, wherein the sequence of the first pair of PCR primers is shown in SEQ ID NO.1~2, the sequence of the second pair of PCR primers is shown in SEQ ID NO.3~4, and the sequence of the third pair of PCR primers is shown in SEQ ID NO.5~6; 2) The first PCR amplification reaction system, the second PCR amplification reaction system and the third PCR amplification reaction system are respectively amplified under corresponding PCR conditions to obtain a first amplification product, a second amplification product and a third amplification product; the three amplification products are mixed into one tube at a mass ratio of 30:2:3 to obtain a mixed test PCR amplification product; 3) The mixed test PCR amplification products are subjected to magnetic bead purification, concentration and fragment size detection, library construction, on-machine library preparation, on-machine sequencing and bioinformatics analysis to determine the KIR genotype of the sample to be tested.

6. The method for genotyping 16 KIR genes by third-generation long-read sequencing as claimed in claim 5, characterized in that: The PCR amplification reaction systems are: The first PCR amplification reaction system is 25.0 μL, including the following components: 2.0 μL of 10 μM first pair of PCR primers, 2.0 μL of 50-100 ng / μL genomic DNA, 12.5 μL of long fragment rapid amplification DNA polymerase mixture, and 8.5 μL of ddH2O; The second PCR amplification reaction system is 25.0 μL, including the following components: 2.0 μL of 10 μM second pair of PCR primers, 2.0 μL of 50-100 ng / μL genomic DNA, 12.5 μL of long fragment rapid amplification DNA polymerase mixture, and 8.5 μL of ddH2O; The third PCR amplification reaction system is 25.0 μL, including the following components: 6.0 μL of 1.0 μM third pair of PCR primers, 2.0 μL of 50-100 ng / μL genomic DNA, 12.5 μL of long fragment rapid amplification DNA polymerase mixture and 4.5 μL of ddH2O.

7. The method for genotyping 16 KIR genes by third-generation long-read sequencing as claimed in claim 5, characterized in that: The conditions for the PCR amplification are: The reaction conditions of the first PCR amplification reaction system are: 98°C for 2 min; 98°C for 10 Sec; 59°C for 10 Sec; 72°C for 4 min, 35 cycles; 72°C for 5 min; The reaction conditions of the second PCR amplification reaction system are: 98°C for 2 min; 98°C for 10 Sec; 56°C for 10 Sec; 72°C for 4 min, 35 cycles; 72°C for 5 min; The reaction conditions of the third PCR amplification reaction system are: 98°C for 2 min; 98°C for 10 Sec; 65°C for 10 Sec; 72°C for 5 min, 35 cycles; 72°C for 5 min.

8. The method for genotyping 16 KIR genes by third-generation long-read sequencing as claimed in claim 5, characterized in that: The lengths of the amplified products of the first pair of PCR primers, the second pair of PCR primers and the third pair of PCR primers are all between 9.2 and 16.7 kb, covering all exons and introns of each KIR gene.

Citation Information

Patent Citations

  • Synchronous sequencing-based typing method for 14 genes of functional killer cell immunoglobulin-like receptors KIRs

    CN107058544A

  • Method for simultaneous sequence-based typing of 14 functional killer cell immunoglobulin-like receptor (KIR) genes

    US10266877B2