Method, device and storage medium for identifying genetic variations of hla genes based on ngs data

By constructing an HLA gene reference sequence database and aligning it to the standard coordinate system of the human reference genome, the accuracy problem of NGS data in HLA gene genetic variation detection was solved, achieving accurate detection of highly polymorphic regions and improving the accuracy of disease association analysis.

CN120452533BActive Publication Date: 2025-11-04BEIJING HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510524834.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-11-04
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

Existing NGS data is difficult to accurately compare and detect HLA gene genetic variations, especially in highly polymorphic regions, leading to insufficient reliability and accuracy of disease association analysis results. In particular, in areas with high HLA polymorphism, conventional procedures are prone to missing genetic variations and causing errors in analysis results.

Method used

By obtaining HLA allele sequences with similar individual HLA genotypes as reference sequences, an HLA gene reference sequence database is constructed and aligned to the standard coordinate system of the human reference genome. This allows for the accurate detection and standardization of SNP genotypes, ensuring that test results are comparable under the same reference system.

Benefits of technology

It enables precise detection of HLA polymorphic sites, improves the accuracy and reliability of disease association analysis, can identify HLA genetic SNP sites with high research value, fills the detection gap of NGS data in highly variable regions, and improves the accuracy of HLA regional allele frequencies and genotype frequencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452533B_ABST
    Figure CN120452533B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for identifying genetic variation of HLA genes based on NGS data and a storage medium. By obtaining a sequence close to an individual HLA genotype sequence as a sample HLA gene reference sequence, accurate detection of an HLA SNP genotype is performed, and then alignment is performed in a coordinate system, so that the detected SNP sites have comparability between individuals, and a disease-susceptible HLA genetic SNP site with high research and diagnosis value can be more accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of HLA genetic analysis, and in particular to a method, device and storage medium for identifying HLA genetic variations based on NGS data. BACKGROUND

[0002] The human leukocyte antigen (HLA) system is a group of genes that play an important role in the immune system. These genes encode molecules that form the major histocompatibility complex (MHC), of which MHC I and II are professional antigen peptide presentation molecules on the cell surface, which can present antigen peptides to T cell receptors (TCR) on the surface of T cells, and are considered to be the most immunogenic antigens in the process of allograft rejection. Matching of HLA genotypes between donors and recipients is crucial for ensuring the safety and prognosis of organ transplantation. With the rapid development of organ transplantation over the past two decades, HLA genotype databases and typing methods have made great progress. The IPD-IMGT / HLA database is the most authoritative human HLA database, containing records of all HLA alleles officially named by the WHO HLA Nomenclature Committee. As of August 2024, the IPD-IMGT / HLA database contains 27717 HLA and related alleles. HLA genes are difficult to study due to their complex structure, multiple gene characteristics, high polymorphism, and unusually complex evolutionary patterns. Despite decades of research, the characterization and interpretation of HLA genetic and genomic variability remain challenging, and HLA mutations are an important way to understand many diseases.

[0003] Currently, SNP microarray chip data, which relies on HLA imputation, is widely used in HLA SNP research. This method helps to study the association of HLA genetic variation with larger population at lower cost, and has produced a lot of important research results, but many SNP genotype is obtained by HLA allele and SNP marker linkage disequilibrium, and the reliability of the inferred genotype is difficult to judge. The second-generation sequencing technology (NGS) has the advantages of large throughput, high accuracy and rich information, which can make up for this deficiency. Many HLA typing methods relying on NGS data have been developed, and many of them, such as HLA-HD, Optitype, etc., have been widely recognized and used because of their reliable and accurate typing results. When using case-control grouping samples to study the association between HLA genetic variation and disease, it is necessary to ensure that the detected variations are under the same genomic reference system. The standard NGS mutation detection process usually uses the reference genome released by the Genome Reference Consortium as the benchmark sequence for alignment and mutation detection. Hindered by HLA gene polymorphism, if the genotype of the detected sample is significantly different from the sequence of the integrated HLA genotype of the reference genome, the sequencing reads will be difficult to correctly align to the human reference genome, thus the strategy of sequence alignment and mutation detection in the conventional analysis under the same reference genome fails, and the authenticity and reliability of the results of disease association analysis cannot be ensured.

[0004] The current widely used method for assessing the association between HLA SNP and disease risk based on second-generation sequencing data is to identify genetic variations through the GATK best practice pipeline, and then use HLA imputation for GWAS research. However, the conventional NGS analysis pipeline represented by the GATK best practice pipeline has the following shortcomings: if the HLA genotype of the detected sample is significantly different from the integrated HLA genotype sequence of the reference genome, the reads will be difficult to correctly align to the human reference genome, and the accurate detection of mutations in high polymorphism regions and the association analysis of diseases cannot be supported. Genes such as DRB3 and DRB4 do not exist in all humans and are not integrated on chromosome 6 of the human reference genome GRCh38. If the ALT sequence of chromosome 6 is included in the reference genome in order to identify genetic variations in these regions, multi-site alignment will occur, which may be filtered out by some software in downstream analysis, thereby greatly affecting the analysis results. The HLA region has a complex linkage disequilibrium structure. The HLA imputation method infers the genotype of the undetected site by introducing multi-ethnic reference panels, but the reliability of the inferred genotype is difficult to judge. In addition, for the less studied regional population and HLA gene region, the genotype is more difficult to obtain by inference due to the lack of prior information.

[0005] The information in the background art is merely intended to explain the general background of the application and should not be considered as admitting or in any form implying that these information constitutes the prior art known to those skilled in the art. SUMMARY

[0006] To solve at least part of the technical problems in the prior art, the present application obtains HLA allele sequences with similar individual HLA genotypes as HLA gene reference sequences for accurate detection of HLA SNP genotypes, and aligns to the integrated HLA gene sequence and position coordinates of the human reference genome such as GRCh38.p14, so that the detected SNP sites have comparability, the target HLA genetic SNP sites can be more accurately identified, and the analysis range can cover multiple HLA functional genes recorded in known databases such as IPD-IMGT / HLA, avoiding the problem that the standard NGS analysis process cannot support SNP detection in high polymorphism regions. Specifically, the present application includes the following contents.

[0007] In a first aspect of the present application, a method for identifying HLA genetic variations based on NGS data is provided, which comprises the following steps:

[0008] (1) providing NGS data of a sample, aligning the NGS data to an HLA reference sequence database to obtain HLA allele sequence information close to the HLA genotype of the sample in units of genes, and constructing a sample HLA gene reference sequence, wherein the HLA reference sequence database comprises data of multiple alleles of multiple HLA and related genes;

[0009] (2) constructing a standard coordinate alignment matrix of the sample HLA gene reference sequence in units of genes, and generating a sample file based on a standard coordinate system, wherein the standard coordinates are referenced to an HLA reference genome constructed based on a standard coordinate system of a human reference genome;

[0010] (3) aligning the NGS data of the sample to the sample HLA gene reference sequence to obtain sample HLA gene alignment data, wherein the sample HLA gene reference sequence is the HLA gene reference sequence constructed in step (1);

[0011] (4) calculating HLA SNP genotype likelihood based on the sample HLA gene alignment data to obtain a genotype VCF file in standard VCFv4.2 format, wherein the sample HLA gene alignment data is the sample HLA gene alignment data generated in step (3);

[0012] (5) standardizing the genotype VCF file based on the standard coordinate alignment matrix to obtain a standard coordinate standardized genotype VCF file in standard VCFv4.2 format, wherein the standard coordinate alignment matrix is the standard coordinate alignment matrix of the sample HLA gene reference sequence generated in step (2), and the genotype VCF file is the genotype VCF file generated in step (4);

[0013] (6) repeating steps (1), (2), (3), (4) and (5) to obtain standard coordinate standardized genotype VCF files of multiple genes of the sample, merging the multiple HLA gene files to generate a sample genotype VCF file conforming to the VCFv4.2 format; and

[0014] (7) identifying HLA genetic variations based on the sample genotype VCF file.

[0015] In some embodiments, the method for identifying genetic variations of HLA genes based on NGS data according to the present application, wherein the plurality of HLA and related genes comprises HLA-A, HLA-B, HLA-C, HLA-DMA, HLA-DMB, HLA-DOA, HLA-DOB, HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-E, HLA-F, HLA-G, HFE, MICA, MICB, TAP1, TAP2.

[0016] In some embodiments, the method for identifying genetic variations of HLA genes based on NGS data according to the present application, wherein the HLA reference sequence database in step (1) is established based on the HLA allele sequence information recorded in the HLA thematic database.

[0017] In some embodiments, the method for identifying genetic variations of HLA genes based on NGS data according to the present application, wherein the HLA allele sequence corresponding to the coding strand of HLA gene located on the negative strand is incorporated into the HLA reference sequence database after reverse complementation, and the HLA allele sequence corresponding to the coding strand of HLA gene located on the positive strand is directly incorporated into the HLA reference sequence database.

[0018] In some embodiments, the method for identifying genetic variations of HLA genes based on NGS data according to the present application, wherein the sample HLA gene reference sequence in step (1) is composed of HLA allele sequences close to the sample HLA genotype, wherein the number of allele sequences is 0, 1 or 2.

[0019] In some embodiments, the method for identifying genetic variations of HLA genes based on NGS data according to the present application, wherein the standard coordinate system HLA reference genome in step (2) is established by intercepting the HLA gene sequence and position information integrated on the human reference genome, and the standard coordinate alignment matrix of the sample HLA gene reference sequence is obtained by performing multiple sequence alignment on the standard coordinate system HLA reference genome and the sample HLA gene reference sequence.

[0020] In a second aspect, the present application provides a device for identifying genetic variations of HLA genes based on NGS data, comprising:

[0021] a data acquisition unit configured to acquire NGS data of a sample;

[0022] The data processing unit is configured to align, in units of bases, HLA allele sequence information close to the sample HLA genotype in an HLA reference sequence database to generate an HLA gene reference sequence, construct a standard coordinate alignment matrix of the sample HLA gene reference sequence, align the sample NGS data to the HLA gene reference sequence, calculate the HLA SNP genotype possibility, generate a genotype VCF file, generate a standard coordinate genotype VCF file of the HLA gene by means of the standard coordinate alignment matrix, then repeat the above steps to obtain a plurality of standard coordinate genotype VCF files of the HLA gene, and combine a plurality of HLA genes to generate a sample genotype VCF file conforming to the VCFv4.2 format, and identify HLA genetic variation.

[0023] The HLA reference sequence database comprises data of a plurality of alleles of a plurality of HLA and related genes, and the standard coordinates are referenced to a human reference genome.

[0024] In a third aspect of the present application, an apparatus for identifying HLA genetic variation based on NGS data is provided, wherein the apparatus comprises a memory, a processor, and an execution program stored on the memory and capable of running on the processor, and the execution program is configured to implement the steps of the method of the first aspect.

[0025] In a fourth aspect of the present application, a computer readable storage medium is provided, wherein the computer readable storage medium stores an execution program, and the execution program, when executed by a processor, can implement the steps of the method of any one of the first aspect of the present application.

[0026] Technical effects of the present application:

[0027] HLA belongs to the most polymorphic gene system in human, which is closely related to the diseases with immunopathological components. The NGS data with the advantages of large throughput, high accuracy and rich information cannot well reveal the correlation with complex diseases from the level of HLA polymorphic sites due to the obstruction of HLA gene polymorphism. Due to the technical blank, the population genome database with high recognition, such as NCBI ALFA project, 1000 Genome Project, and the alleles of HLA region with high polymorphism have a missing or inaccurate frequency in the population. The method and system of the present application can detect 26 polymorphic genes in the IPD-IMG / HLA database, accurately detect the HLA SNP site genotype by obtaining the HLA allele sequence close to the individual HLA genotype as the reference gene sequence, and align in the human reference genome coordinate system to make the detected SNP site comparable between individuals, which can more accurately identify the disease-susceptible HLA genetic SNP site with high research and diagnostic value, and accordingly give more accurate HLA region allele frequency and genotype frequency.

[0028] The present application fills the technical blank of accurate detection of genetic variation in the highly variable region of HLA gene using NGS data, can detect 26 HLA genes in the IPD-IMGT / HLA database, and the variation detection result can be further used to identify disease-susceptible HLA genetic variation sites with high research value or obtain the allele frequency of a specific population, which has important significance and value for the development of HLA and immunogenetics. BRIEF DESCRIPTION OF DRAWINGS

[0029] Figure 1 An exemplary overall flowchart of the method of the present application.

[0030] Figure 2 An exemplary screenshot of HLA allele information extracted from the HLAdat file of the IPD / IMGT-HLA database.

[0031] Figure 3 An exemplary flowchart of constructing an HLA reference sequence database according to the present application.

[0032] Figure 4 An exemplary flowchart of constructing a standard coordinate alignment matrix according to the present application.

[0033] Figure 5 An exemplary MAFFT multiple sequence alignment diagram according to the present application.

[0034] Figure 6 An exemplary standard coordinate alignment matrix diagram according to the present application.

[0035] Figure 7Exemplary build sample HLA gene reference sequence (0 / 1 / 2) flowchart.

[0036] Figure 8 Exemplary get HLA allele(s) support list flowchart.

[0037] Figure 9 Exemplary genotype VCF standardize flowchart.

[0038] Figure 10 Exemplary heterozygous VCF record merge flowchart.

[0039] Figure 11 DRB1 gene is homozygous DRB1*15:01 sample 2 exon 2 coverage screenshot (realign.bam) under GATK best practice flow.

[0040] Figure 12 DRB1 genotype is homozygous DRB1*07:01 sample 2 exon 2 coverage screenshot (realign.bam) under GATK best practice flow.

[0041] Figure 13 DRB1 genotype is DRB1*07:01 sample 2 exon 2 coverage screenshot.

[0042] Figure 14 380 coding region SNP CMPD major and minor AF comparison to 147 AF associated with the present application. DETAILED DESCRIPTION

[0043] Various exemplary embodiments of the present application will now be described in detail, without intent to limit the application, but to provide a more full understanding of the application. The detailed description uses numerical examples to illustrate the application. These examples should not be construed as a limitation on the scope or spirit of the application, but merely an illustration of certain embodiments of the application.

[0044] It should be understood that the terms used herein are for the purpose of describing particular embodiments and are not intended to limit the application. Additionally, for numerical ranges that are disclosed herein, it is intended that every

[0045] Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as those commonly understood by one of ordinary skill in the art to which this application belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present application, the preferred methods and materials are described. All publications mentioned in this specification are herein incorporated by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. In case of conflict, the content of the present specification will control.

[0046] Embodiments

[0047] The method of the present embodiment can be divided into a preparation phase and an analysis phase. The preparation phase mainly includes constructing an HLA reference sequence database and generating a standard coordinate alignment matrix. The analysis phase includes establishing a sample HLA gene reference sequence, calculating HLA SNP genotype likelihood, standardizing genotype VCF coordinates, and identifying HLA genetic variations.

[0048] Due to the advantages of comprehensive gene coverage, low cost, high speed, interpretability, and extensive richness of WES data, the example application in the analysis phase is on 147 cases of Chinese blood donor gDNA WES data. The method and system of the present embodiment can also be applied to gDNA NGS data of other populations or case-control samples, wherein the NGS detection range needs to include the target HLA gene region.

[0049] I. Preparation phase:

[0050] 1. Download the HLA dat format file from the IPD-IMGT / HLA database, and extract the gene, name_brief, name_full, length, sequence, and gene_type information. The present embodiment includes 26 functional genes in the database, including classic and non-classic HLA I / II genes and 5 non-HLA genes (see Table 1 and Figure 2 ).

[0051] Table 1 List of 26 HLA genes

[0052]

[0053]

[0054] 2. Select the integrated HLA gene sequence and coordinates of the human reference genome (GRCh38.p14 primary assembly chromosome 6) as the standard coordinate system HLA reference genome, and use the integrated sequence and coordinates of chr6_GL000255v2_alt and chr6_GL000256v2_alt for HLA-DRB3 and HLA-DRB4 genes not integrated in chromosome 6 as the standard coordinates to generate the custom HLA_gene_hg38.txt file (see Table 1).

[0055] 3. Construct the HLA reference sequence database based on the alleles of the HLA gene in the IPD-IMGT / HLA database (see Table 2), including the following processing steps: 1) retain the longest HLA allele sequence information for HLA alleles with identical protein products; 2) include the reverse complement of HLA allele sequences with coding chains located on the negative strand in the HLA reference sequence database, or directly include them; 3) index the HLA reference sequence; 4) generate HLA allele UTR / exon / intron Bed files. Figure 3

[0056] 4. Generate fasta files for gene sequences from the standard coordinate system HLA reference genome and HLA reference sequence database in units of genes, and construct HLA gene standard coordinate alignment matrices (see Table 3) after multiple sequence alignment. Figure 4

[0057] Use multiple sequence alignment software to generate high-precision multiple sequence alignment files in units of genes, and manually check and calibrate the correctness of the sequence alignment results. In this embodiment, MAFFT-auto XX.fasta>XX.mafft is used, where XX represents the 26 polymorphic functional genes implemented in this embodiment.

[0058] Construct standard coordinate alignment matrices (see Table 4) based on MAFFT files. Figure 5 、 Figure 6

[0059] II. Analysis stage:

[0060] Quality control the original sequencing data, in this embodiment, the fastq tool is used, the command is fastp--in1 fq1--in2 fq2--out2 fq2_trim-j json_trim-h html_trim-3-W 4-M 15-q 15-l 36-w ncpu. Trimmomatic or other software with similar functions can also be used to perform operations such as adapter removal and low-quality removal.

[0061] ​​​Obtaining HLA allele sequences close to the sample HLA genotype to establish a sample HLA gene reference sequence Figure 7 , Figure 8 ): Extract the alignment records with an alignment score (AS) greater than 130 from the reads aligned to the HLA reference sequence database HLA allele sequence. Use bedtools to count the coverage and depth of the bed region in units of exons 2 and 3 and all exons, respectively. For high-variable HLA genes belonging to exons 2 and 3, such as classical HLA I and II genes, sort the sum of coverage, depth, sum of all exon coverage, and depth from high to low, and remove the alleles without read support to obtain an ordered allele support list. The sample HLA gene reference sequence established based on the HLA allele sequence can be divided into three cases according to the HLA allele support list:

[0062] (1) All allele support coverage is 0, i.e. no read supports the target gene genotype, then the analysis program of this gene is terminated. This situation will occur in the analysis of HLA genes that do not exist in all humans, such as DRB3 / 4 / 5 genes;

[0063] (2) Only one allele has read support, then select this allele to establish the sample HLA gene reference sequence, this situation may be the sample genotype is homozygous, or the sequence similarity is very high;

[0064] (3) Multiple alleles have read support, for this case, first take the highest support allele (denoted as X allele) as the reference sequence, align and extract reads that are not aligned or AS <= 130, then align the extracted reads to other support alleles again, and obtain the allele support list again. If no read supports other alleles, select the X allele to establish the sample HLA gene reference sequence, otherwise take the highest support allele (denoted as Y allele) among the other alleles, and select the X and Y alleles to establish the sample HLA gene reference sequence. This situation may be the sample genotype is a hybrid with some sequence difference.

[0065] SNP genotype identification: Align the HLA reads to the sample HLA gene reference sequence, extract the alignment records with AS>130, and remove the secondary alignment, then use bcftools mpileup-A to calculate the HLA SNP genotype possibility.

[0066] Genotype VCF standard coordinateFigure 9 、 Figure 10 ):Based on the standard coordinate alignment matrix generated in the preparation stage, the VCF record with IPD-IMGT / HLA allele (sample HLA gene reference sequence) as the reference sequence is converted into the standard coordinate system HLA reference genome, and a SNP genotype file in VCF4.2 standard format is generated.

[0067] SNP calling: identify SNPs using the bcftools call command.

[0068] The method and system of the application are applied to 147 cases of Chinese gDNA WES data, directly compared with the GATK best practice analysis process in Comparative Example 1, consistent with the missense mutation population frequency calculated by the CMDP database HLA genotype frequency in Comparative Example 2, and consistent with the allele population occurrence frequency of the ALAF, 1000Genome database, etc. in Comparative Example 3. The three comparative examples illustrate the accuracy of the method and system of the application in SNP identification in the HLA polymorphic region.

[0069] Comparative Example 1

[0070] In this comparative example, the GATK best practice analysis process and the method of the application are simultaneously applied to gDNA WES data for analysis. IGV (Integrative Genomics Viewer) is a high-performance visualization tool that can be used to view and compare sequencing data. Through IGV, it can be found that when using the GATK best practice analysis process, samples with high similarity to the reference genome integrated HLA genotype can correctly detect mutations at this site, but samples with large differences will have a completely unsupported local region of reads. Figure 11 、 Figure 12 )。

[0071] Taking the DRB1 gene as an example, the DRB1 haplotype integrated on chromosome 6 of the human reference genome is DRB1*15:01, and if the sample HLA genotype is similar to DRB1*15:01, the read coverage is normal ( Figure 11 ), but for genotypes with large differences, the DRB1 gene will have a completely unsupported local region of reads ( Figure 12 ). The application first finds an allele close to the sample genotype sequence as the sample HLA gene reference sequence for alignment and SNP genotype calculation, and then standardizes the genotype VCF, so that the SNP in the high polymorphic region can be accurately detected ( Figure 13 )。

[0072] Comparative Example 2

[0073] The accuracy of SNP population frequency directly reflects the accuracy of SNP detection. The Chinese Marrow Donor Program (CMDP) performed high-resolution typing of five HLA genes (A, C, B, DRB1, and DQB1) of serum samples by SBT method, which can reflect the amino acid level changes, and published the allele and haplotype frequencies of these genes in 169995 Chinese volunteers. In order to test whether the population frequency calculated by the present application is consistent with the true situation, the coding region sequence of HLA genotype in the CMPD database was intercepted, and the position correspondence was carried out by multiple sequence alignment, and the allele frequency of the coding region of HLA gene was calculated by the HLA genotype frequency of Chinese people published by the CMPD database. By applying the method of the present application to 147 Chinese samples, 380 SNP sites meeting MAF (this allele frequency) > 0 and being missense mutations were detected in the coding regions of HLA-A, C, B, DRB1, and DQB1. Based on the 380 SNP sites, the consistency of the allele frequency of the major allele (major allele) and the minor allele (minor allele) of the CMDP with the allele frequency of the 147 Chinese people calculated based on the SNP detection results of the method of the present application was compared, and the Pearson correlation coefficient was 0.949 and 0.941, respectively. Figure 14 ) 99.2% (377 / 380) of SNP sites, the absolute difference between the major allele frequency (major allele frequency) of the CMDP and the allele frequency (allele frequency, AF) of the 147 Chinese people is less than 0.15; 91.6% (348 / 380) of the absolute difference is less than 0.1.

[0074] Comparative Example 3

[0075] The NCBI ALAF database calculates the SNP allele frequency of more than two million samples in dbGap, among which the East Asian population (EAS) is represented by China (7612) and Japan (4138). We screened 253 minor alleles as missense mutations in the coding regions of HLA-A, C, B, DRB1, and DQB1, and the SNP sites with allele AF > 0.1 in the 147 Chinese samples analyzed based on the method of the present application. For these 251 sites, polymorphism can be confirmed by the CMDP database (Table 2), but 20.3% (51 / 251) of the SNPs are completely consistent with the reference base in the ALAF East Asian (EAS) and no polymorphism is detected (Table 2). 84.3% (43 / 51) of these sites are located in the Exon2 high polymorphism region, and the genotype of the undetected allele has a large difference with the local difference of GRCh38.p14, which leads to the failure of correct alignment of the reads.

[0076] Table 2 HLA-A / C / B / DRB1 / DQB1 coding region 51 SNP sites where no population polymorphism was detected in ALAF East Asian population

[0077]

[0078]

[0079] Further, we queried the latest published population frequency database on NCBI for the allele frequency of East Asian population for the CMDP minor alleles at the 20 SNP sites in the DRB1 Exon2 region. The 1000 Genomes, 1000 Genomes (30X), gnomAD (Genomes) databases showed a large number of sites in the DRB1 Exon2 hypervariable region where no population polymorphism was detected (Table 3).

[0080] Table 3

[0081]

[0082]

[0083] While the application has been described with reference to the example embodiments, it will be understood by those skilled in the art that various changes in form and details can be made therein without departing from the spirit and scope of the application. The scope of the claims should not be limited by the preferred embodiments set forth in the examples, but should be given the broadest interpretation consistent with the principles, and the scope of the application set forth in the claims.

Claims

1. A method for identifying HLA gene genetic variations based on NGS data, comprising the following steps: (1) Provide NGS data of the sample, compare the HLA allele sequence information that is close to the HLA genotype of the sample in the HLA reference sequence database on a gene-by-gene basis, and construct the HLA gene reference sequence of the sample, wherein the HLA reference sequence database contains data of multiple alleles of multiple HLA and related genes; (2) Construct a standard coordinate alignment matrix of the sample HLA gene reference sequence on a gene-by-gene basis, and generate a sample file based on the standard coordinate system, wherein the standard coordinates are referenced by the HLA reference genome based on the standard coordinate system constructed based on the human reference genome; (3) Align the NGS data of the sample to the HLA gene reference sequence of the sample to obtain the HLA gene alignment data of the sample, wherein the HLA gene reference sequence of the sample is the HLA gene reference sequence constructed in step (1). (4) Calculate the probability of HLA SNP genotype based on the sample HLA gene alignment data to obtain a genotype VCF file in standard VCFv4.2 format, wherein the sample HLA gene alignment data is the sample HLA gene alignment data generated in step (3); (5) Standardize the genotype VCF file based on the standard coordinate alignment matrix to obtain a standard coordinate genotype VCF file in standard VCFv4.2 format, wherein the standard coordinate alignment matrix is ​​the standard coordinate alignment matrix of the sample HLA gene reference sequence generated in step (2), and the genotype VCF file is the genotype VCF file generated in step (4); (6) Repeat steps (1), (2), (3), (4), and (5) to obtain the standardized coordinate genotype VCF file for multiple genes in the sample, merge multiple HLA gene files, and generate a sample genotype VCF file conforming to the VCFv4.2 format; and (7) Identify HLA gene genetic variations based on sample genotype VCF files.

2. The method for identifying HLA gene genetic variations based on NGS data according to claim 1, characterized in that, The multiple HLA and related genes include HLA-A, HLA-B, HLA-C, HLA-DMA, HLA-DMB, HLA-DOA, HLA-DOB, HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQA2, HL A-DQB1, HLA-DQB2, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-E, HLA-F, HLA-G, HFE, MICA, MICB, TAP1, TAP2.

3. The method for identifying HLA gene genetic variations based on NGS data according to claim 1, characterized in that, The HLA reference sequence database in step (1) is established based on the HLA allele sequence information collected in the HLA topic database.

4. The method for identifying HLA gene genetic variations based on NGS data according to claim 3, characterized in that, The HLA allele sequences corresponding to the HLA gene coding strand on the negative strand are included in the HLA reference sequence database after reverse complementation, and the HLA allele sequences corresponding to the HLA gene coding strand on the positive strand are directly included in the HLA reference sequence database.

5. The method for identifying HLA gene genetic variations based on NGS data according to claim 1, characterized in that, Step (1) The sample HLA gene reference sequence consists of HLA allele sequences that are close to the sample HLA genotype, wherein the number of allele sequences is 0, 1 or 2.

6. The method for identifying HLA gene genetic variations based on NGS data according to claim 1, characterized in that, Step (2) The standard coordinate system HLA reference genome is established by extracting HLA gene sequences and location information integrated on the human reference genome. The standard coordinate system HLA reference genome and the sample HLA gene reference sequence are subjected to multiple sequence alignment to obtain the standard coordinate alignment matrix of the sample HLA gene reference sequence.

7. A device for identifying HLA gene genetic variations based on NGS data, the device being used to implement the method according to any one of claims 1-6, characterized in that, include: The data acquisition unit is configured to acquire NGS data of the sample; The data processing unit is configured to align HLA allele sequence information that is close to the sample HLA genotype by comparing it with the HLA reference sequence database on a gene-by-gene basis, generate HLA gene reference sequences, construct a standard coordinate alignment matrix of the sample HLA gene reference sequences, align sample NGS data to the HLA gene reference sequences, calculate the probability of HLA SNP genotypes, generate genotype VCF files, generate standardized coordinate genotype VCF files of HLA genes using the standard coordinate alignment matrix, repeat the above steps to obtain standardized coordinate genotype VCF files of multiple HLA genes, merge multiple HLA gene files, generate a sample genotype VCF file conforming to the VCFv4.2 format, and identify HLA gene genetic variations. The HLA reference sequence database contains data on multiple alleles of multiple HLA and related genes, and the standard coordinates are based on the human reference genome.

8. A device for identifying HLA gene genetic variations based on NGS data, characterized in that, The device includes: a memory, a processor, and an executable program stored in the memory and capable of running on the processor, the executable program being configured to implement the steps of the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an executable program that, when executed by a processor, can perform the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • HLA genotype-SNP linkage database, its constructing method, and HLA typing method

    CN103221551A

  • Method for identifying genome selection and utilization interval of wheat breeding based on density of polymorphic sites in resequencing data

    CN111798922A