Method and equipment for identifying HLA gene genetic variation based on NGS data and storage medium
By constructing a standard coordinate system for reference sequence of individual HLA allele for NGS data comparison, the difficulty in detecting SNP caused by HLA genotype differences was solved, and the accurate detection of HLA polymorphic sites and the accurate identification of disease susceptibility sites were achieved.
Patent Information
- Application Number
- CN202510524834.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-24
AI Technical Summary
When the existing NGS data are very different from the HLA genotype and the reference genome, it is difficult to accurately compare, resulting in difficulty in detecting SNPs in high polymorphic regions, affecting the reliability and accuracy of disease association analysis.
By obtaining the HLA allelic sequences with relatively close HLA genotypes as references, a standard coordinate system is constructed for comparison, and a genotype file in standard VCF format is generated to ensure that the SNP loci are comparable among individuals.
The accurate detection of HLA polymorphic sites is achieved, the accuracy of identification of disease-susceptible SNP sites is improved, and more accurate HLA region allelic frequency and genotype frequency is provided, filling the technical gap in genetic variation detection of NGS data in highly variable regions.
Smart Images

Figure CN120452533A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of HLA gene genetic analysis, and specifically to a method, device, and storage medium for identifying HLA gene genetic variations based on NGS data. Background Art
[0002] The human leukocyte antigen (HLA) system is a group of genes that play an important role in the immune system. The molecules encoded by these genes form the major compatibility complex (MHC), of which MHC I and II are professional antigen peptide presenting molecules on the cell surface. They can present antigen peptides to the T cell receptors (TCR) on the surface of T cells and are considered to be the most immunogenic antigens in the process of allogeneic transplant rejection. Matching the HLA genotypes of donors and recipients is crucial to ensuring the safety and prognosis of organ transplantation. With the vigorous development of organ transplantation over the past two decades, HLA genotype databases and typing methods have made great progress. The IPD-IMGT / HLA database is the most authoritative human HLA database, containing records of all HLA alleles officially named by the WHO HLA Nomenclature Committee. As of August 2024, the IPD-IMGT / HLA database contains 27,717 HLA and related alleles. HLA genes are difficult to study due to their complex structure, polygenic and highly polymorphic characteristics, and extremely complex evolutionary patterns. Despite decades of research, the characterization and interpretation of the genetic and genomic variability of HLA genes remains challenging, and HLA mutations are an important way to understand many diseases.
[0003] Currently, SNP microarray data, relying on HLA imputation, are widely used in HLA-associated phenotype (HLA-SNP) studies. This approach facilitates HLA association studies in larger populations at a lower cost and has yielded numerous important findings. However, many SNP genotypes are inferred based on linkage disequilibrium between HLA alleles and SNP markers, making the reliability of the inferred genotypes difficult to assess. Next-generation sequencing (NGS), with its advantages of high throughput, high accuracy, and rich information content, can address this shortcoming. Numerous HLA typing methods have been developed that rely on NGS data, many of which, such as HLA-HD and Optitype, have gained widespread recognition and adoption due to their reliable and accurate typing results. When using case-control samples for HLA genetic variation and disease association studies, it is crucial to ensure that the variants being called are based on the same genomic reference system. Standard NGS mutation calling workflows typically use the reference genome published by the Genome Reference Consortium as the baseline sequence for alignment and variant calling. Due to the obstruction of HLA gene polymorphism, if the genotype of the tested sample differs significantly from the sequence of the HLA genotype integrated in the reference genome, it will be difficult to correctly align the sequencing reads to the human reference genome, thus making the strategy of sequence alignment and mutation detection on the same reference genome in conventional analysis ineffective, and unable to ensure the authenticity and reliability of the disease association analysis results.
[0004] A widely used approach for assessing the association between HLA SNPs and disease risk based on next-generation sequencing data is to identify genetic variants using the GATK best practice pipeline, followed by GWAS studies using HLA imputation. However, a limitation of conventional NGS analysis workflows, such as the GATK best practice pipeline, is that if the HLA genotype of the sample being tested differs significantly from the HLA genotype sequence integrated into the reference genome, reads will be difficult to accurately map to the human reference genome, hindering accurate mutation detection in highly polymorphic regions and disease association analysis. Genes such as DRB3 and DRB4 are not present in all humans and are not integrated on chromosome 6 in the human reference genome GRCh38. Incorporating the ALT sequence from chromosome 6 into the reference genome to identify genetic variants in these regions can easily result in multiple alignments, which can be filtered out by some downstream analysis software, significantly impacting the analysis results. The HLA region has a complex linkage disequilibrium structure. HLA imputation methods use multi-ethnic reference panels to infer the genotype probability of undetected loci, but the reliability of the inferred genotypes is difficult to determine. Furthermore, for less-studied populations and HLA gene regions, genotypes are even more difficult to infer due to a lack of prior information.
[0005] The information in the background technology is only intended to illustrate the general background of the invention and should not be regarded as an admission or any form of suggestion that this information constitutes the prior art known to a person skilled in the art. Summary of the Invention
[0006] To address at least some of the technical issues in the prior art, the present invention accurately detects HLA SNP genotypes by obtaining HLA allele sequences with similar HLA genotypes as HLA gene reference sequences. These sequences are then aligned to the HLA gene sequences and positional coordinates integrated into the human reference genome, such as GRCh38.p14. This ensures comparable detection of detected SNPs, enabling more accurate identification of target HLA genetic SNPs. Furthermore, the analysis encompasses multiple HLA functional genes documented in known databases such as IPD-IMGT / HLA, thus addressing the inability of standard NGS analysis protocols to detect SNPs in highly polymorphic regions. Specifically, the present invention encompasses the following.
[0007] In a first aspect of the present invention, a method for identifying genetic variations in HLA genes based on NGS data is provided, comprising the following steps:
[0008] (1) Providing NGS data of the sample, comparing the HLA allele sequence information that is closest to the HLA genotype of the sample in the HLA reference sequence database on a gene-by-gene basis, and constructing the HLA gene reference sequence of the sample, wherein the HLA reference sequence database contains data on multiple alleles of multiple HLA and related genes;
[0009] (2) constructing a standard coordinate alignment matrix for the sample HLA gene reference sequence based on genes, and generating a sample file based on a standard coordinate system, wherein the standard coordinates are based on the HLA reference genome, a standard coordinate system constructed based on the human reference genome;
[0010] (3) aligning the NGS data of the sample to the sample HLA gene reference sequence to obtain sample HLA gene alignment data, wherein the sample HLA gene reference sequence is the HLA gene reference sequence constructed in step (1);
[0011] (4) calculating the HLA SNP genotype likelihood based on the sample HLA gene comparison data to obtain a genotype VCF file in the standard VCFv4.2 format, wherein the sample HLA gene comparison data is the sample HLA gene comparison data generated in step (3);
[0012] (5) performing standard coordinate conversion on the genotype VCF file based on the standard coordinate alignment matrix to obtain a standard coordinate conversion genotype VCF file in the standard VCFv4.2 format, wherein the standard coordinate alignment matrix is the standard coordinate alignment matrix of the sample HLA gene reference sequence generated in step (2), and the genotype VCF file is the genotype VCF file generated in step (4);
[0013] (6) Repeat steps (1), (2), (3), (4), and (5) to obtain standard coordinate genotype VCF files for multiple genes of the sample, merge the multiple HLA gene files, and generate a sample genotype VCF file that conforms to the VCFv4.2 format; and
[0014] (7) Identify HLA gene genetic variations based on sample genotype VCF files.
[0015] In certain embodiments, according to the method for identifying HLA gene genetic variations based on NGS data according to the present invention, the multiple HLA and related genes include HLA-A, HLA-B, HLA-C, HLA-DMA, HLA-DMB, HLA-DOA, HLA-DOB, HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-E, HLA-F, HLA-G, HFE, MICA, MICB, TAP1, and TAP2.
[0016] In certain embodiments, according to the method for identifying HLA gene genetic variations based on NGS data according to the present invention, the HLA reference sequence database in step (1) is established based on the HLA allele sequence information included in the HLA thematic database.
[0017] In certain embodiments, according to the method for identifying HLA gene genetic variations based on NGS data of the present invention, the HLA allele sequence corresponding to the HLA gene coding chain located on the negative strand is reverse complemented and then included in the HLA reference sequence database, and the HLA allele sequence corresponding to the HLA gene coding chain located on the positive strand is directly included in the HLA reference sequence database;
[0018] In certain embodiments, according to the method for identifying HLA gene genetic variations based on NGS data according to the present invention, the HLA gene reference sequence of the sample in step (1) is composed of HLA allele sequences that are close to the HLA genotype of the sample, wherein the number of allele sequences is 0, 1, or 2.
[0019] In certain embodiments, according to the method for identifying HLA gene genetic variations based on NGS data of the present invention, in step (2), the standard coordinate system HLA reference genome is established by intercepting the HLA gene sequence and position information integrated in the human reference genome, and the standard coordinate system HLA reference genome and the sample HLA gene reference sequence are subjected to a multi-sequence alignment to obtain a standard coordinate alignment matrix of the sample HLA gene reference sequence.
[0020] In a second aspect, the present invention provides a device for identifying HLA gene genetic variations based on NGS data, comprising:
[0021] a data acquisition unit configured to acquire NGS data of a sample;
[0022] a data processing unit configured to compare, on a gene-by-gene basis, an HLA reference sequence database to obtain HLA allele sequence information that is closest to the sample HLA genotype, generate an HLA gene reference sequence, construct a standard coordinate alignment matrix for the sample HLA gene reference sequence, align the sample NGS data to the HLA gene reference sequence, calculate the HLA SNP genotype probability, generate a genotype VCF file, generate a standard coordinate genotype VCF file for the HLA gene using the standard coordinate alignment matrix, repeat the above steps to obtain standard coordinate genotype VCF files for multiple HLA genes, merge the multiple HLA genes, generate a sample genotype VCF file in the VCFv4.2 format, and perform identification of HLA gene genetic variations;
[0023] The HLA reference sequence database contains data on multiple alleles of multiple HLA and related genes, and the standard coordinates are based on the human reference genome.
[0024] In a third aspect of the present invention, a device for identifying genetic variations of HLA genes based on NGS data is provided, wherein the device comprises: a memory, a processor, and an execution program stored in the memory and capable of running on the processor, wherein the execution program is configured to implement the steps of the method described in the first aspect.
[0025] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein an execution program is stored on the computer-readable storage medium, and when the execution program is executed by a processor, the steps of the method described in any one of the first aspects of the present invention can be implemented.
[0026] Technical effects of the present invention:
[0027] HLA is one of the most polymorphic human gene systems and is closely associated with immunopathological diseases. However, due to the high throughput, high accuracy, and rich information content of HLA gene polymorphisms, NGS data, despite their advantages of high throughput, high accuracy, and high information content, have not been able to fully reveal the associations of HLA polymorphic loci with complex diseases. Due to technical gaps, the frequencies of many highly polymorphic HLA region alleles in highly recognized population genomic databases, such as the NCBIALFA project and the 1000Genome Project, are missing or inaccurate in the population. The method and system of the present invention can detect 26 polymorphic genes included in the IPD-IMG / HLA database. By obtaining HLA allele sequences with similar HLA genotypes as reference gene sequences, the genotypes of HLA SNP loci are accurately detected. The detected SNP loci are then aligned to the human reference genome coordinate system to make them comparable between individuals. This allows for more accurate identification of disease-susceptible HLA genetic SNP loci with high research and diagnostic value, and accordingly, provides more accurate HLA region allele and genotype frequencies.
[0028] This invention fills the technical gap in accurately detecting genetic variations in the highly variable regions of the HLA gene using NGS data. It can detect 26 HLA genes included in the IPD-IMGT / HLA database. Its variation detection results can be further used to identify disease-susceptible HLA genetic variation sites with high research value or to obtain allele frequencies in specific populations. It has important significance and value for the development of HLA and immunogenetics. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 An exemplary overall flow chart of the method of the present invention.
[0030] Figure 2 Example screenshot of HLA allele information extracted from the HLAdat file in the IPD / IMGT-HLA database.
[0031] Figure 3 Flowchart of an exemplary method of constructing an HLA reference sequence database according to the present invention.
[0032] Figure 4 The present invention exemplifies a flow chart for constructing a standard coordinate alignment matrix.
[0033] Figure 5 Schematic diagram of an exemplary MAFFT multiple sequence alignment of the present invention.
[0034] Figure 6 Schematic diagram of an exemplary standard coordinate alignment matrix of the present invention.
[0035] Figure 7The present invention exemplifies a flow chart for establishing a sample HLA gene reference sequence (0 / 1 / 2).
[0036] Figure 8 The present invention exemplifies a flow chart of obtaining an HLA allele(s) support list.
[0037] Figure 9 Flowchart of exemplary genotype VCF standard coordinates of the present invention.
[0038] Figure 10 Flowchart of exemplary heterozygous VCF record merging according to the present invention.
[0039] Figure 11 Screenshot of the coverage of exon 2 in a sample with homozygous DRB1*15:01 gene under the GATK best practice process (realign.bam).
[0040] Figure 12 Screenshot of the coverage of exon 2 in a sample with homozygous DRB1 genotype DRB1*07:01 under the GATK best practice process (realign.bam).
[0041] Figure 13 Screenshot of the coverage of exon 2 in a sample with the DRB1 genotype of DRB1*07:01.
[0042] Figure 14 Comparison of the correlation between 380 coding region SNPs CMPD major and minor AF and 147 cases of AF in the present invention. DETAILED DESCRIPTION
[0043] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as limiting the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.
[0044] It should be understood that the terms described in the present invention are only for describing particular embodiments and are not intended to limit the present invention. In addition, for numerical ranges in the present invention, it should be understood that the upper and lower limits of the ranges and each intermediate value therebetween are specifically disclosed. Each smaller range between any stated value or intermediate value within a stated range and any other stated value or intermediate value within the stated range is also included in the present invention. The upper and lower limits of these smaller ranges may be independently included or excluded within the scope.
[0045] Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the invention belongs. Although only preferred methods and materials are described herein, any methods and materials similar or equivalent to those described herein may also be used in the practice or testing of the present invention. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods and / or materials related to the documents. In the event of any conflict with any incorporated document, the content of this specification shall prevail.
[0046] Example
[0047] The method of this embodiment can be divided into a preparation phase and an analysis phase. The preparation phase mainly includes constructing an HLA reference sequence database and generating a standard coordinate alignment matrix. The analysis phase includes establishing a sample HLA gene reference sequence, calculating HLA SNP genotype likelihoods, converting genotypes to standard coordinates using VCF, and identifying HLA genetic variations.
[0048] Due to the advantages of comprehensive genetic coverage, low cost, high speed, interpretability, and extensive richness of WES data, the analysis phase was applied to gDNA WES data from 147 Chinese blood donors. The methods and systems of the present invention can also be applied to gDNA NGS data from other populations or case-control samples, where the NGS detection range needs to include the target HLA gene region.
[0049] 1. Preparation stage:
[0050] 1. Download the HLA dat format file from the IPD-IMGT / HLA database and extract the gene, name_brief, name_full, length, sequence and gene_type information. This example covers 26 functional genes in the database, including classical and non-classical HLA I / II genes and 5 non-HLA genes (see Tables 1 and Figure 2 ).
[0051] Table 1 List of 26 HLA genes
[0052]
[0053]
[0054] 2. The HLA gene sequences and coordinates integrated into the human reference genome (GRCh38.p14 primary assembly chromosome 6) were selected as the standard coordinate system HLA reference genome. For the HLA-DRB3 and HLA-DRB4 genes not integrated into chromosome 6, the sequences and coordinates integrated into chr6_GL000255v2_alt and chr6_GL000256v2_alt, respectively, were used as standard coordinates to generate a custom HLA_gene_hg38.txt file (see Table 1).
[0055] 3. Construction of HLA reference sequence database based on HLA gene alleles in IPD-IMGT / HLA database ( Figure 3 ), including the following processing steps: 1) retaining the longest HLA allele sequence information for HLA alleles with identical protein products; 2) taking the reverse complement of the HLA allele sequence located on the negative strand of the HLA gene coding chain and adding it to the HLA reference sequence database; otherwise, directly including it; 3) establishing an index for the HLA reference sequence; 4) generating the HLA allele UTR / exon / intron Bed file.
[0056] 4. Gene sequences from the HLA reference genome and HLA reference sequence database in the standard coordinate system are used as units to generate fasta files, and multiple sequence alignment is performed to construct the HLA gene standard coordinate alignment matrix ( Figure 4 ).
[0057] Multiple sequence alignment software was used to generate high-precision multiple sequence alignment files on a gene-by-gene basis, and the correctness of the alignment results was manually verified. In this example, MAFFT-auto XX.fasta>XX.mafft was used, where XX represents the 26 polymorphic functional genes implemented in this example.
[0058] Construct a standard coordinate alignment matrix based on the MAFFT file ( Figure 5 、 Figure 6 ).
[0059] 2. Analysis phase:
[0060] Quality control of the raw sequencing data was performed using the fastq tool in this example, using the command fastp --in1fq1 --in2fq2 --out2 fq2_trim -j json_trim -h html_trim -3 -W 4-M 15-q 15-l 36-w ncpu . Alternatively, trimmomatic or other software with similar functionality for operations such as adapter removal and low-quality removal can be used.
[0061] Obtain the HLA allele sequence that is closer to the sample HLA genotype to establish the sample HLA gene reference sequence ( Figure 7 、 Figure 8 ): Extract the reads aligned to the HLA allele sequence in the HLA reference sequence database, and extract the alignment records with an alignment score (AS) greater than 130. Use bedtools to count the sum of coverage and depth of the bed region in units of exons 2 and 3 and all exons respectively. For highly variable HLA genes belonging to exons 2 and 3, such as classic HLA I and class II genes, sort from high to bottom based on the sum of coverage of exons 2 and 3, the sum of depth, the sum of coverage of all exons, and the sum of depth. Otherwise, sort from high to low based on the sum of coverage and the sum of depth of all exons. After removing alleles without read support, an ordered list of allele support is obtained. The sample HLA gene reference sequence established based on the HLA allele sequence can be divided into three cases according to the HLA allele support list:
[0062] (1) If the support coverage of all alleles is 0, that is, no reads support the target gene genotype, the analysis process for this gene is terminated. This situation may occur in the analysis of HLA genes that are not present in all people, such as DRB3 / 4 / 5 genes;
[0063] (2) If only one allele has read support, this allele is selected to establish the sample HLA gene reference sequence. In this case, the sample genotype may be homozygous or heterozygous with extremely high sequence similarity;
[0064] (3) Multiple alleles are supported by reads. In this case, the allele with the highest support (denoted as X allele) is first taken as the reference sequence, and the unaligned reads and reads with AS <= 130 are extracted. The extracted reads are then re-aligned to other supported alleles to obtain the allele support list again. If no reads support other alleles, the X allele is selected to establish the sample HLA gene reference sequence. Otherwise, the allele with the highest support among the other alleles (denoted as Y allele) is taken, and the X and Y alleles are selected to establish the sample HLA gene reference sequence. This situation may be that the sample genotype is heterozygous with some sequence differences.
[0065] SNP genotype identification: HLA reads were aligned to the sample HLA gene reference sequence, and alignment records with AS>130 were extracted. Minor alignments were removed and duplicates were removed. After that, the HLA SNP genotype probability was calculated using bcftools mpileup-A.
[0066] Genotype VCF standard coordinates ( Figure 9 、 Figure 10 ): Based on the standard coordinate alignment matrix generated in the preparation stage, the VCF records with IPD-IMGT / HLA alleles (sample HLA gene reference sequence) as the reference sequence are converted to the standard coordinate system HLA reference genome, and the SNP genotype file in the VCF4.2 standard format is generated.
[0067] SNP calling: Use the bcftools call command to identify SNPs.
[0068] The method and system of the present invention were applied to 147 Chinese gDNA WES datasets. In Comparative Example 1, a direct comparison was made with the GATK best-practice analysis pipeline. In Comparative Example 2, the consistency of missense mutation population frequencies calculated with HLA genotype frequencies from the CMDP database was compared. In Comparative Example 3, the consistency of allele population frequencies was compared with databases such as ALAF and 1000Genome. These three comparative examples demonstrate the accuracy of the method and system of the present invention in identifying SNPs in HLA polymorphic regions.
[0069] Comparative Example 1
[0070] In this comparative example, both the GATK best practice analysis process and the method of the present invention were used to perform analysis on gDNA WES data. IGV (Integrative Genomics Viewer) is a high-performance visualization tool that can be used to view and compare sequencing data. IGV can intuitively show that when using the GATK best practice analysis process, samples with high HLA genotype similarity to the reference genome can correctly detect mutations there, but samples with large differences will have local regions with no read support at all ( Figure 11 、 Figure 12 ).
[0071] Taking the DRB1 gene as an example, the DRB1 haplotype integrated on chromosome 6 of the human reference genome is DRB1*15:01. If the HLA genotypes of the samples are highly similar to DRB1*15:01, the read coverage is normal ( Figure 11 ), but for genotypes with large differences, there will be a situation where there is no read support in the local region of the DRB1 gene ( Figure 12 The present invention first finds the allele closest to the sample genotype sequence as the sample HLA gene reference sequence for comparison and SNP genotype calculation, and then converts the genotype VCF standard coordinates. Therefore, SNPs in highly polymorphic regions can be accurately detected ( Figure 13 ).
[0072] Comparative Example 2
[0073] The accuracy of SNP population frequency directly reflects the accuracy of SNP detection. The China Marrow Database (CMDP) performed high-resolution typing of the five HLA genes A, C, B, DRB1, and DQB1 of serum samples using the SBT method, which can reflect changes in amino acid levels, and published the allele and haplotype frequencies of these genes in 169,995 Chinese volunteers. In order to verify whether the population frequency calculated by the present invention is consistent with the actual situation, the present invention intercepted the coding region sequence of the HLA genotype in the CMPD database and performed positional correspondence through multiple sequence alignment. The HLA gene coding region allele frequency was calculated based on the Chinese HLA genotype frequency published by the CMPD database. By applying the method of the present invention to 147 Chinese samples, 380 SNP sites that met MAF (this allele frequency)>0 and were missense mutations were detected in the HLA-A, C, B, DRB1, and DQB1 coding regions. Based on 380 SNP loci, the allele frequencies of the major allele and minor allele of CMDP were compared with the allele frequencies of 147 Chinese people calculated based on the SNP detection results of the present method. The Pearson correlation coefficients were 0.949 and 0.941, respectively. Figure 14 The absolute difference between the major allele frequency (AF) of CMDP and the allele frequency (AF) of 147 Chinese individuals was less than 0.15 for 99.2% (377 / 380) of the SNPs, and less than 0.1 for 91.6% (348 / 380).
[0074] Comparative Example 3
[0075] The NCBI ALAF database calculated SNP allele frequencies for over two million samples in dbGap, with East Asian populations (EAS) represented by China (7,612) and Japan (4,138). We screened 253 SNPs in the HLA-A, C, B, DRB1, and DQB1 coding regions for missense alleles and AF values > 0.1 for these alleles in 147 Chinese samples analyzed using the methods described in this paper. For these 251 SNPs, polymorphisms were confirmed using the CMDP database (Table 2). However, 20.3% (51 / 251) of these SNPs were completely identical to the reference bases in the ALAF East Asian (EAS) region, with no polymorphism detected (Table 2). Of these SNPs, 84.3% (43 / 51) were located in the highly polymorphic region of Exon 2, where the genotypes at which the alleles were not detected differed significantly from those in GRCh38.p14, resulting in incorrect read alignment.
[0076] Table 2 51 SNPs in the HLA-A / C / B / DRB1 / DQB1 coding region that were not polymorphic in the ALAF East Asian population
[0077]
[0078]
[0079] We further queried the allele frequencies for the CMDP minor alleles at 20 SNPs in the DRB1 Exon2 region using the latest published population frequency databases from NCBI. The 1000Genomes, 1000Genomes (30X), and gnomAD (Genomes) databases revealed a significant number of sites in the DRB1 Exon2 hypervariable region that had no polymorphism detected in the population (Table 3).
[0080] Table 3
[0081]
[0082]
[0083] Although the present invention has been described with reference to exemplary embodiments, it should be understood that the invention is not limited to the disclosed exemplary embodiments. Various modifications and variations may be made to the exemplary embodiments of the present specification without departing from the scope or spirit of the present invention. The scope of the claims is to be given the broadest interpretation so as to encompass all such modifications and equivalent structures and functions.
Claims
1. A method for identifying HLA gene genetic variations based on NGS data, comprising the following steps: (1) Providing NGS data of the sample, comparing the HLA allele sequence information that is closest to the HLA genotype of the sample in the HLA reference sequence database on a gene-by-gene basis, and constructing the HLA gene reference sequence of the sample, wherein the HLA reference sequence database contains data on multiple alleles of multiple HLA and related genes; (2) constructing a standard coordinate alignment matrix for the sample HLA gene reference sequence based on genes, and generating a sample file based on a standard coordinate system, wherein the standard coordinates are based on the HLA reference genome, a standard coordinate system constructed based on the human reference genome; (3) aligning the NGS data of the sample to the sample HLA gene reference sequence to obtain sample HLA gene alignment data, wherein the sample HLA gene reference sequence is the HLA gene reference sequence constructed in step (1); (4) calculating the HLA SNP genotype likelihood based on the sample HLA gene comparison data to obtain a genotype VCF file in the standard VCFv4.2 format, wherein the sample HLA gene comparison data is the sample HLA gene comparison data generated in step (3); (5) performing standard coordinate conversion on the genotype VCF file based on the standard coordinate alignment matrix to obtain a standard coordinate conversion genotype VCF file in the standard VCFv4.2 format, wherein the standard coordinate alignment matrix is the standard coordinate alignment matrix of the sample HLA gene reference sequence generated in step (2), and the genotype VCF file is the genotype VCF file generated in step (4); (6) Repeat steps (1), (2), (3), (4), and (5) to obtain standard coordinate genotype VCF files for multiple genes of the sample, merge the multiple HLA gene files, and generate a sample genotype VCF file that conforms to the VCFv4.2 format; and (7) Identify HLA gene genetic variations based on sample genotype VCF files.
2. The method for identifying HLA gene genetic variation based on NGS data according to claim 1, characterized in that: The multiple HLA and related genes include HLA-A, HLA-B, HLA-C, HLA-DMA, HLA-DMB, HLA-DOA, HLA-DOB, HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQA2, HL A-DQB1, HLA-DQB2, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-E, HLA-F, HLA-G, HFE, MICA, MICB, TAP1, TAP2.
3. The method for identifying HLA gene genetic variation based on NGS data according to claim 1, characterized in that: The HLA reference sequence database in step (1) is established based on the HLA allele sequence information included in the HLA database.
4. The method for identifying HLA gene genetic variation based on NGS data according to claim 3, characterized in that: The HLA allele sequence corresponding to the HLA gene coding chain located on the negative strand is reverse complemented and then included in the HLA reference sequence database, and the HLA allele sequence corresponding to the HLA gene coding chain located on the positive strand is directly included in the HLA reference sequence database.
5. The method for identifying HLA gene genetic variation based on NGS data according to claim 1, characterized in that: In step (1), the sample HLA gene reference sequence consists of HLA allele sequences that are close to the sample HLA genotype, wherein the number of allele sequences is 0, 1, or 2.
6. The method for identifying HLA gene genetic variation based on NGS data according to claim 1, characterized in that: In step (2), a standard coordinate system HLA reference genome is established by intercepting the HLA gene sequence and position information integrated in the human reference genome, and a multi-sequence alignment is performed on the standard coordinate system HLA reference genome and the sample HLA gene reference sequence to obtain a standard coordinate alignment matrix for the sample HLA gene reference sequence.
7. A device for identifying HLA gene genetic variations based on NGS data, characterized in that: include: a data acquisition unit configured to acquire NGS data of a sample; a data processing unit configured to compare, on a gene-by-gene basis, an HLA reference sequence database to obtain HLA allele sequence information that is closest to the sample HLA genotype, generate an HLA gene reference sequence, construct a standard coordinate alignment matrix for the sample HLA gene reference sequence, align the sample NGS data to the HLA gene reference sequence, calculate the HLA SNP genotype probability, generate a genotype VCF file, generate a standard coordinate genotype VCF file for the HLA gene using the standard coordinate alignment matrix, repeat the above steps to obtain standard coordinate genotype VCF files for multiple HLA genes, merge the multiple HLA genes, generate a sample genotype VCF file in the VCFv4.2 format, and perform identification of HLA gene genetic variations; The HLA reference sequence database contains data on multiple alleles of multiple HLA and related genes, and the standard coordinates are based on the human reference genome.
8. A device for identifying HLA gene genetic variation based on NGS data, characterized in that: The device includes: a memory, a processor, and an execution program stored in the memory and capable of running on the processor, wherein the execution program is configured to implement the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an execution program, which, when executed by a processor, can implement the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
HLA genotype-SNP linkage database, its constructing method, and HLA typing method
CN103221551A
Method for identifying genome selection and utilization interval of wheat breeding based on density of polymorphic sites in resequencing data
CN111798922A
Method for Imputing Classical Alleles of HLA-DRB1 Using Single Nucleotide Polymorphisms in East Asian Population
KR1020160029947A
Improved human leukocyte antigen (HLA) genotyping
WO2024006705A1