Method for hla typing and electronic device therefor

By combining third-generation sequencing technology with conformity sequence correction, low-quality data is filtered out, solving the problem of low accuracy in existing HLA typing methods. This achieves high-resolution and high-accuracy HLA typing, which is suitable for clinical diagnosis and forensic identification.

CN120015120BActive Publication Date: 2025-12-05ANNOROAD GENE TECHNOLOGY (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510499415.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-12-05
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

Existing HLA typing methods suffer from low accuracy, especially when dealing with low-frequency or novel HLA alleles, where the reliability of the typing results is greatly reduced.

Method used

Using third-generation sequencing technology, the gene data of the sample were compared with the HLA gene database to filter out genotyping results with a coverage of less than 30% or an average sequencing depth of less than 30×, and the genotyping results were corrected using the consistency sequence to obtain the final HLA genotyping result.

Benefits of technology

It achieves high-resolution and high-accuracy HLA typing, providing more precise genotype information without relying on assembly and clustering, and is suitable for scenarios with high data quality requirements such as clinical diagnosis and forensic identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015120B_ABST
    Figure CN120015120B_ABST
Patent Text Reader

Abstract

The application provides a method for HLA typing and an electronic device thereof. The method comprises the following steps: S1) comparing the sequencing data of each sample gene with an HLA gene database to obtain a first candidate HLA gene typing set; S2) filtering the typing results with a coverage less than 30% or an average sequencing depth less than 30x in the first candidate HLA gene typing set to obtain a second candidate HLA gene typing set; S3) filtering the typing results with differences in the coding regions of the sample genes from the corresponding consensus sequences in the second candidate HLA gene typing set to obtain a final HLA typing result, wherein the sample HLA gene is obtained by sequencing with a third-generation sequencing platform.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of HLA typing, and in particular, to a method for HLA typing and an electronic device thereof. BACKGROUND

[0002] The Human Leukocyte Antigen (HLA) system, as part of the Major Histocompatibility Complex (MHC), is a region of the genome closely related to immune function. It is located on the short arm of chromosome 6 and contains multiple closely linked loci such as HLA-A, HLA-B, HLA-C, etc. These loci encode highly polymorphic proteins that have a crucial impact on immune responses and disease susceptibility between individuals.

[0003] Traditional HLA typing methods, especially those based on serology and cytology, have played an important role in early research, but due to their inherent limitations such as low resolution, complex operation and time-consuming, they have gradually become difficult to meet the needs of current precision medicine. In recent years, with the development of high-throughput sequencing technology, especially the application of second-generation sequencing technology, HLA typing based on sequencing data has become possible. However, this method still has significant defects, i.e., the resolution is limited by the length of the read, and in dealing with complex HLA gene regions, the accuracy is often affected by the splicing errors caused by short reads.

[0004] In view of the above limitations, the third-generation sequencing technology is considered an ideal choice to overcome these problems due to its ability to generate long reads and high-accuracy data. In particular, the HiFi sequencing technology of PacBio allows direct sequencing of the entire HLA gene region due to its unique long read characteristics, which theoretically improves the resolution and accuracy of HLA typing. Although PacBio provides hifi-hla software to process HiFi data and perform HLA typing, the performance of this method to some extent depends on the efficiency and accuracy of the clustering algorithm. When the sequencing depth is insufficient or the clustering effect is poor, the reliability of the typing results will be greatly reduced, especially when dealing with low-frequency or novel HLA alleles, this effect is particularly pronounced. SUMMARY

[0005] The main purpose of the present application is to provide a method for HLA typing and an electronic device thereof to solve the problem of low accuracy of typing results in the prior art.

[0006] In order to achieve the above object, according to a first aspect of the present application, there is provided a method for HLA typing, comprising: S1) aligning sequencing data of each sample gene with an HLA gene database to obtain a first set of candidate HLA gene typing; S2) filtering, in the first set of candidate HLA gene typing, typing results with coverage less than 30% or average sequencing depth less than 30x to obtain a second set of candidate HLA gene typing; S3) filtering, from the sample gene, typing results that are different from the coding region of the corresponding consensus sequence in the coding region of the sample gene according to the consensus sequence of the corresponding sample gene in each typing result in the second set of candidate HLA gene typing to obtain a final HLA typing result; wherein the sample gene is obtained by sequencing with a third-generation sequencing platform.

[0007] Further, the HLA gene database comprises IMGT / HLA data.

[0008] Further, in S2), the filtering comprises: i) multiplying the coverage corresponding to each site of each sample gene in the typing result in each first set of candidate HLA gene typing with the sequencing depth and dividing the sum by the sum of all sequencing depths to obtain the average sequencing depth; filtering the typing result with an average sequencing depth less than 30x to obtain a first filtered set; ii) sorting the typing results of each locus in the sample gene according to the size of the average sequencing depth of the sites of the sample gene in the first filtered set, retaining the typing results of the top two loci with the largest average sequencing depth, and the typing result of the locus with the largest average sequencing depth is recorded as the first typing result and the typing result of the locus with the second largest average sequencing depth is recorded as the second typing result; dividing the average sequencing depth of the first typing result by the average sequencing depth of the second typing result to obtain an average sequencing depth ratio; when the average sequencing depth ratio is greater than 2, the first typing result is retained; when the average sequencing depth ratio is less than or equal to 2, the first typing result and the second typing result are retained.

[0009] Further, in S3), the method for obtaining the consensus sequence comprises: counting, in each typing result in the second set of candidate HLA typing, the base with the most coverage by the gene fragment in the sequencing data of the sample gene at each site in each sample gene as the correct base; combining each correct base to obtain the consensus sequence.

[0010] In order to achieve the above object, according to a second aspect of the present application, there is provided an electronic device for HLA typing, comprising: an alignment unit, a filtering unit and a correction unit; wherein the alignment unit is configured to align sequencing data of each sample gene with an HLA gene database to obtain a first candidate HLA gene typing set; the filtering unit is configured to filter out typing results with coverage less than 30% or average sequencing depth less than 30x in the first candidate HLA gene typing set to obtain a second candidate HLA gene typing set; and the correction unit is configured to correct typing results in the second candidate HLA gene typing set to obtain a final HLA typing result; the correction comprises: filtering out typing results in the coding region of the sample gene that are different from the coding region of the corresponding consensus sequence according to the consensus sequence of the corresponding sample gene in each typing result in the second candidate HLA gene typing set; wherein the sample gene is sequenced by a third-generation sequencing platform.

[0011] Further, the HLA gene database comprises an IMGT / HLA database.

[0012] Further, the filtering unit comprises a first filtering unit and a second filtering unit; the first filtering unit is configured to filter out typing results with average sequencing depth less than 30x in the first candidate HLA gene typing set to obtain a first filtered set; the calculation method of the average sequencing depth comprises: multiplying the coverage of each sample gene in each typing result in the first candidate HLA gene typing set with the sequencing depth of the corresponding site, and dividing the sum of the products by the sum of all sequencing depths to obtain the average sequencing depth.

[0013] Further, the second filtering unit is configured to judge the typing results in the first filtered set; the judgment comprises: sorting the typing results of each locus in the sample gene according to the size of the average sequencing depth of the sites of the sample gene in the first filtered set, retaining the typing results of the first two loci with the largest average sequencing depth, and recording the typing result of the locus with the largest average sequencing depth as a first typing result and the typing result of the locus with the second largest average sequencing depth as a second typing result; dividing the average sequencing depth of the first typing result by the average sequencing depth of the second typing result to obtain an average sequencing depth ratio; when the average sequencing depth ratio is greater than 2, retaining the first typing result; when the average sequencing depth ratio is less than or equal to 2, retaining the first typing result and the second typing result.

[0014] Further, the correction unit comprises a consistency sequence obtaining unit; the consistency sequence obtaining unit is configured to obtain a consistency sequence of each sample gene in each typing result in each second candidate HLA genotype set; the consistency sequence obtaining method comprises: counting, in each typing result in the second candidate HLA genotype set, a base covered most by a gene fragment in the sequencing data of the sample gene at each site in each sample gene, and recording the base as a correct base; and combining each correct base to obtain a consistency sequence.

[0015] In order to achieve the above-mentioned purpose, according to a third aspect of the present application, a computer readable storage medium is provided, the storage medium comprising a stored program, wherein when the program is running, the HLA typing method described above is controlled.

[0016] In order to achieve the above-mentioned purpose, according to a fourth aspect of the present application, a processor is provided, the processor being configured to run a program, wherein when the program is running, the HLA typing method described above is executed.

[0017] By using the technical solution of the present application, the sequencing data of each sample gene is compared with the HLA gene database, the first candidate HLA genotype set is obtained, the coverage and the average sequencing depth in the typing result in the set are calculated and judged, the typing result with a coverage less than 30% or an average sequencing depth less than 30x is filtered out, and the consistency sequence corresponding to the sample gene sequence in the typing result is used for correction, so that the HLA typing result with improved accuracy can be obtained. The HLA analysis method of the present application does not need to be clustered, the algorithm is simple, and is more suitable for popularization in the application of HLA typing. BRIEF DESCRIPTION OF DRAWINGS

[0018] The accompanying drawings, which form a part of the present application, are included to provide a further understanding of the application, and are incorporated herein for purposes of illustrating the illustrative embodiments of the present application and the explanations provided herein. In the drawings:

[0019] Figure 1 A flowchart of a method for HLA typing according to an embodiment of the present application is shown.

[0020] Figure 2 A schematic diagram of an electronic device for HLA typing according to an embodiment of the present application is shown.

[0021] Figure 3 A hardware structure block diagram of an electronic device for HLA typing according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0022] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other in the case of no conflict. The present application will be described in detail below in combination with the embodiments.

[0023] Terminology:

[0024] Consensus sequence: refers to the nucleotide sequence with the highest frequency at each position in the alignment results of a group of sequences.

[0025] Resolution of HLA typing: refers to the accuracy of HLA typing, the accuracy of identification of HLA alleles. According to the different typing accuracy, the resolution of HLA typing can reach four main levels, which are low resolution (2-bit resolution), intermediate resolution, high resolution (4, 6, 8-bit resolution) and allelic level. Among them, the low resolution, i.e. 2-bit resolution, is classified based on serology (the definition of serological classification is that HLA antigens are classified based on their ability to react with specific antibodies), or the allelic group to which the allele belongs, representing a 2-bit typing result (for example: A02, A03, A11, C03, etc.). The intermediate resolution generally includes a subset of alleles with the same first field, but excludes some alleles sharing the field, which is used for preliminary screening, and further confirms the HLA type of the candidate donor. The high resolution typing result is defined as a group of alleles, which encode the same protein sequence of the region of HLA protein molecule (called antigen binding site), and exclude the alleles that are not expressed as cell surface proteins. The allelic level typing provides complete HLA gene sequence information, including the nucleotide sequence of all exons and introns.

[0026] Coverage: the number of bases covered to a gene divided by the total length of the gene.

[0027] Average sequencing depth: the sum of the depths of the bases covered to a gene divided by the number of bases covered.

[0028] As mentioned in the background, when the sequencing depth is insufficient or the clustering effect is poor, the reliability of HLA typing results based on third-generation sequencing will be greatly reduced, especially when dealing with low-frequency or new HLA alleles. Based on this, the inventors tried to develop a new HLA typing scheme, and thus proposed a series of protection schemes in the present application.

[0029] In a first typical embodiment of the present application, a method for HLA typing is provided, which comprises S1) aligning the sequencing data of each sample gene with an HLA gene database to obtain a first set of candidate HLA gene typing; S2) filtering the typing results with coverage less than 30% or average sequencing depth less than 30x in the first set of candidate HLA gene typing to obtain a second set of candidate HLA gene typing; S3) filtering the typing results in the coding region of the sample gene that are different from the corresponding coding region of the consensus sequence in each typing result in the second set of candidate HLA gene typing to obtain the final HLA typing result; wherein the sample gene is sequenced by a third-generation sequencing platform, preferably HiFi sequencing of PacBio. The flowchart of the HLA typing method of the present application is shown in Figure 1

[0030] The current HLA typing methods mainly include HLA serological typing, cytological typing, and typing based on second-generation sequencing data. Such typing methods are tedious in experimental process, high in data analysis cost, and low in result accuracy and resolution. The HLA typing technology based on third-generation sequencing data can produce long read and high-accuracy data, which can solve the above problems to some extent. However, the method of HLA analysis based on third-generation sequencing depends on the efficiency and accuracy of the clustering algorithm to some extent. When the sequencing depth is insufficient or the clustering effect is poor, the reliability of the typing result will be greatly reduced, especially when dealing with low-frequency or new HLA alleles, which greatly affects the accuracy of the typing result. Therefore, it is particularly important to develop a new HLA typing method. The HLA typing method of the present application is based on the sequencing accuracy of the third-generation sequencing technology, and does not need to be assembled and clustered, which overcomes the above technical problems and can realize ultra-high resolution and high accuracy of HLA typing.

[0031] The purpose of splitting the HLA gene in the present application is to split the sample amount to meet the requirements of sample loading for on-machine detection under the condition of low sample amount, to more accurately align the HLA analysis that needs to be analyzed, to simplify the data processing method, and to improve the analysis efficiency.

[0032] ​The HLA typing method of the present application accurately controls the threshold of filtering inaccurate typing results (coverage less than 30%, average sequencing depth less than 30x) based on the calculation of the coverage and sequencing depth of each site, and effectively filters out low-quality data through the strategy of accurate calculation, screening and multiple alignment, thereby improving the accuracy and consistency of typing. Moreover, the present application further aligns the consistent sequences obtained with candidate genes, filters out candidate genes with variations in the coding region, thereby improving the accuracy of typing, further confirming the correctness of the typing results, and providing more accurate typing results.

[0033] The HLA typing method of the present application realizes high typing accuracy and high resolution without relying on assembly and clustering, and at least achieves 6-bit resolution, which is suitable for clinical diagnosis, forensic identification and other scenarios with extremely high requirements for data quality. When the HLA typing method of the present application is applied to HLA typing, it can provide more accurate genotype information for clinical transplantation matching, drug metabolism prediction and disease association research.

[0034] High-resolution typing can be further divided into: 4-bit resolution: for example, A*01:01, which can distinguish specific HLA proteins. 6-bit resolution: for example, A*01:01:01, which can distinguish specific HLA coding sequences. 8-bit resolution: for example, A*01:01:01:01, which can distinguish specific HLA gene sequences, including untranslated regions and introns.

[0035] The above-mentioned sequencing data of the sample gene is obtained by sequencing on a third-generation sequencing platform, preferably HiFi sequencing of PacBio, or other sequencing means capable of obtaining full-length HLA genes can achieve the effect of the present application.

[0036] Preferably, the above-mentioned alignment software includes but is not limited to minimap2, pbmm2 or winnowmap2. Preferably, the above-mentioned splitting software includes but is not limited to lima software.

[0037] In a preferred embodiment, the HLA gene database includes the IMGT / HLA database. The HLA database used for alignment is updated and supplemented by the above-mentioned alignment HLA database at least every three months from the new HLA typing of the IMGT / HLA official website.

[0038] In a preferred embodiment, in S2), the method for calculating the average sequencing depth comprises: i) multiplying the coverage corresponding to each locus of each sample gene in the typing result in each first candidate HLA genotyping set by the sequencing depth, and dividing the sum of the products by the sum of all sequencing depths to obtain the average sequencing depth; filtering the typing result with an average sequencing depth lower than 30x to obtain a first filtered set; sorting the typing result of each locus in the sample gene according to the size of the average sequencing depth of the locus of the sample gene in the first filtered set, retaining the typing result of the top two loci with the largest average sequencing depth, and recording the typing result of the locus with the largest average sequencing depth as the first typing result and the typing result of the locus with the second largest average sequencing depth as the second typing result; dividing the average sequencing depth of the first typing result by the average sequencing depth of the second typing result to obtain an average sequencing depth ratio; when the average sequencing depth ratio is greater than 2, retaining the first typing result; and when the average sequencing depth ratio is less than or equal to 2, retaining the first typing result and the second typing result. The sum of the coverage corresponding to each locus of each sample gene in the typing result in each first candidate HLA genotyping set multiplied by the sequencing depth is divided by the sum of all sequencing depths to obtain the average sequencing depth, and the formula of this step is: (sum(cover x depth) / sum(depth)), wherein "sum" represents summation, "cover" represents coverage, and "depth" represents sequencing depth.

[0039] When the average sequencing depth ratio is less than or equal to 2, it indicates that the average sequencing depth ratio of the two haplotypes is relatively uniform, which meets the expectation of the design of the typing experimental system, and therefore the two typing results can be retained; when the average sequencing depth ratio is greater than 2, it indicates that the average sequencing depth of the two haplotypes is significantly different, which does not meet the expectation of the design of the experimental system, and the haplotype with the smaller average sequencing depth (the second typing result) is false positive, and therefore the second typing result is filtered and the first typing result is retained. Those skilled in the art can adjust the judgment threshold of the average sequencing depth ratio according to the experimental system and the average sequencing depths of the first typing result and the second typing result.

[0040] In a preferred embodiment, in S3), the method for obtaining the consensus sequence comprises: counting, in each typing result of the second candidate HLA genotyping set, the base that is covered most by the gene fragments in the sequencing data of the sample gene at each locus in each sample gene, and recording the base as the correct base; and combining each correct base to obtain the consensus sequence.

[0041] The application filters out low-quality data through multiple alignment, filtering and correction using consistent sequences, improves the accuracy of the final typing result, has high resolution, simple and easy-to-operate steps, avoids the problems of complex steps and inaccurate results of HLA typing based on second-generation sequencing, avoids the disadvantages of HLA typing methods based on third-generation sequencing in terms of dependence on sequencing depth and clustering methods, and can provide more accurate genotype information for clinical transplantation matching, drug metabolism prediction and disease association research, and is more suitable for promotion.

[0042] In a second typical embodiment of the application, an electronic device for HLA typing is provided (schematic diagrams of units of the electronic device are shown in Figure 2 The electronic device includes an alignment unit 01, a filtering unit 02 and a correction unit 03. The alignment unit 01 is configured to align each sample gene with an HLA gene database to obtain a first candidate HLA gene typing set. The filtering unit 02 is configured to filter out typing results with coverage less than 30% or average sequencing depth less than 30x in the first candidate HLA gene typing set to obtain a second candidate HLA gene typing set. The correction unit 03 is configured to correct the typing results in the second candidate HLA gene typing set to obtain a final HLA typing result. The correction includes filtering out typing results that differ from the corresponding coding region of the consistent sequence in the coding region of the sample gene according to the consistent sequence of the corresponding sample gene in each typing result in the second candidate HLA gene typing set. The sample gene is obtained by sequencing using a third-generation sequencing platform.

[0043] In a preferred embodiment, the HLA gene database includes the IMGT / HLA database.

[0044] In a preferred embodiment, the filtering unit includes a first filtering unit and a second filtering unit. The first filtering unit is configured to filter out typing results with average sequencing depth less than 30x in the first candidate HLA gene typing set to obtain a first filtered set. The first filtering unit and the second filtering unit. The calculation method of the average sequencing depth includes multiplying the coverage of each sample gene in each typing result in the first candidate HLA gene typing set by the sequencing depth and dividing the sum of all sequencing depths to obtain the average sequencing depth.

[0045] In a preferred embodiment, the second filtering unit is configured to determine the typing results in the first filtered set; the determining comprises: sorting the typing results of each locus in the sample genes in the first filtered set according to the average sequencing depth of the loci in the sample genes, retaining the typing results of the top two loci with the largest average sequencing depth, the typing result of the locus with the largest average sequencing depth is recorded as the first typing result, and the typing result of the locus with the second largest average sequencing depth is recorded as the second typing result; dividing the average sequencing depth of the first typing result by the average sequencing depth of the second typing result to obtain an average sequencing depth ratio; when the average sequencing depth ratio is greater than 2, the first typing result is retained; and when the average sequencing depth ratio is less than or equal to 2, the first typing result and the second typing result are retained.

[0046] In a preferred embodiment, the correction unit comprises a consistency sequence acquisition unit; the consistency sequence acquisition unit is configured to obtain a consistency sequence of each sample gene in each typing result in each second candidate HLA gene typing set; the method for obtaining the consistency sequence comprises: counting, in each typing result in the second candidate HLA typing set, the base with the largest number of gene fragments in the sequencing data of the sample gene covering each locus in each sample gene, and recording the base as a correct base;

[0047] The correct bases are combined to obtain the consistency sequence. In a third typical embodiment of the present application, a computer readable storage medium is provided, which comprises a stored program, wherein when the program is running, the above-mentioned HLA typing method is controlled.

[0048] In a fourth typical embodiment of the present application, a processor is provided, which is configured to run a program, wherein when the program is running, the above-mentioned HLA typing method is executed.

[0049] It should be noted that, for the above-mentioned method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions involved are not necessarily necessary for the present application.

[0050] Those skilled in the art can clearly understand that the application can be implemented by means of software plus a detection device and the like hardware equipment through the description of the above embodiments. Based on such understanding, the part of data processing in the technical solutions of the application can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk and the like, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device and the like) to execute the method of each embodiment or some part of the embodiments of the application.

[0051] The application can be used in many general or special computing system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like.

[0052] The method provided by the application can be executed in a terminal, a computer terminal or a similar computing device. Taking the case of running on a terminal, Figure 3 is a hardware structure block diagram of a terminal for HLA typing according to an embodiment of the application. As shown in Figure 3 , the terminal can include one or more (only one is shown in Figure 3 ) processor A1 (the processor A1 can include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory B1 for storing data. Optionally, the terminal can further include a transmission device C1 for communication function and an input / output device D1. Those skilled in the art can understand that Figure 3 the structure shown is only schematic, which does not limit the structure of the terminal. For example, the terminal can further include more or less components than those shown in Figure 3 , or have a different configuration from that shown in Figure 3 .

[0053] Memory B1 can be used to store computer programs, such as software programs of application software and modules, such as computer programs corresponding to the methods of read splicing, clustering, consistency processing, etc. in the embodiments of the present application. Processor A1 performs various functional applications and data processing, i.e. implements the above-mentioned methods, by running the computer programs stored in memory B1. Memory B1 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, memory B1 can further include memory disposed remotely with respect to processor A1, which can be connected to the terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0054] Transmission device C1 is used to receive or send data via a network. Specific examples of the above-mentioned network can include a wireless network provided by a communication provider of the terminal. In one example, transmission device C1 includes a network adapter (Network Interface Controller, NIC) which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, transmission device C1 can be a radio frequency (Radio Frequency, RF) module used to communicate with the Internet in a wireless manner.

[0055] Obviously, those skilled in the art should understand that some modules or steps of the present application described above can be realized on a general computing device, which can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be realized by program codes executable by a computing device, so as to be stored in a storage device and executed by a computing device, or they can be respectively manufactured into individual integrated circuit modules, or multiple modules or steps among them can be manufactured into a single integrated circuit module to realize. Thus, the present application is not limited to any specific combination of hardware and software. The advantages of the present application will be further explained in detail below in combination with specific embodiments.

[0056] The advantages of the present application will be further explained in detail below in combination with specific embodiments.

[0057] Embodiment 1

[0058] The NA17288 sample standard (Coriell Institute, NA17288) was sequenced and typed, and the specific method was as follows:

[0059] 1. Using lima software, the main HLA gene library is split according to the barcode information of each standard sample, and the split bam file is converted to fastq data by the bam2fastq command of the smrtcmds tool, so as to obtain the fastq data of each sample (sample gene);

[0060] The above method for constructing the main HLA gene library comprises: 1. The full-length sequences of the 11 HLA genes are amplified by a long fragment multiplex PCR amplification system; 2. The sequencing adapters are connected at both ends of the PCR product to construct a HiFi library; 3. The library is bound to sequencing primers and enzyme complexes and then sequenced. The HiFi library is sequenced using a PacBio third-generation sequencing platform.

[0061] 2. The fastq data is aligned to the reference genome using the minimap2 software, and the alignment rate, capture efficiency, coverage depth and other information are counted, as shown in Tables 1 and 2.

[0062] Table 1

[0063] .

[0064] Table 2

[0065]

[0066] 3. The fastq data is aligned to the HLA database using the minimap2 software to obtain the aligned bam file, i.e. the first candidate HLA gene typing set, and the bam file index is constructed using the samtools software;

[0067] According to the bam file, the coverage reads of each base of each HLA gene on the alignment are generated, the coverage depth and base accuracy of each position are calculated, and thus the coverage, average sequencing depth and average alignment accuracy of each gene are calculated;

[0068] The candidate HLA genes with a coverage less than 30% or an average sequencing depth less than 30x are filtered out to obtain the candidate HLA gene set; the specific method and test are as follows:

[0069] (1) Filter out the typing results with a coverage less than 30% (according to the design scheme of the target gene PCR amplification primer, the minimum coverage of the target gene is greater than 30%, and less than 30% is a false positive gene).

[0070] (2) Calculate the weighted average depth of each 6-quantile typing result (indicating a specific HLA coding sequence) (i.e. calculate the average sequencing depth of each typing result), and the calculation method is:

[0071] The sum of coverage x sequencing depth corresponding to each 6-quantile typing result is divided by the sum of sequencing depth (sum(cover x depth) / sum(depth)).

[0072] Typing results with average sequencing depth less than 30x are filtered to obtain a first filtered set;

[0073] Typing results in the first filtered set are sorted according to the average sequencing depth of each locus, and the typing results of the top two loci with the largest average sequencing depth (two haplotypes for each gene) are retained.

[0074] The typing result with the largest average sequencing depth is recorded as the first typing, and the typing result with the second largest average sequencing depth is recorded as the second typing;

[0075] (3) The average sequencing depth of the first typing is divided by the average sequencing depth of the second typing to obtain an average sequencing depth ratio;

[0076] When the average sequencing depth ratio is greater than 2, the first typing is retained;

[0077] When the average sequencing depth ratio is less than or equal to 2, the first typing and the second typing are retained.

[0078] 4. The candidate genes (sample genes) in each candidate HLA gene set are aligned to generate a consensus sequence, and the candidate genes are aligned to filter out genes with variations in the coding region to obtain the final correct HLA gene typing.

[0079] The method for obtaining a consensus sequence comprises: counting the most covered base at each site in each sample gene in each typing result of the second candidate HLA typing set as the correct base, which is covered by the gene fragment in the corresponding HLA gene database; and combining each correct base to obtain a consensus sequence.

[0080] The typing results of this embodiment are described in Table 5.

[0081] Example 2

[0082] In this embodiment, sample standard NA17295 (Coriell Institute, NA17295) and sample standard NA17248 (Coriell Institute, NA17248) are typed using different coverage filtering thresholds, and the remaining steps are consistent with Example 1.

[0083] The typing results of the sample NA17295 are shown in Table 3. In Table 3, when the coverage filtering threshold is set to 20%, the Allele2 haplotype typing result of the HLA-DPA1 gene is incorrect.

[0084] The typing results of sample NA17248 are shown in Table 4. In Table 4, no Allele2 haplotype of HLA-DQB1 gene was identified when the coverage filter threshold was set to 40%.

[0085] Table 3

[0086] .

[0087] Table 4

[0088] .

[0089] In Table 3 and Table 4, "Allele 1" refers to the first allele at a gene locus, and "Allele 2" refers to the second allele at the gene locus. Except for HLA-DRB3, HLA-DRB4 and HLA-DRB5, the blank cells in the HLA typing indicate that there is only one allele at the gene, i.e. the gene is homozygous (the genotype of the two haplotypes is the same), and the other two are heterozygous.

[0090] HLA-DRB3, HLA-DRB4 and HLA-DRB5 are adjacent to HLA-DRB1, and all individuals carry HLA-DRB1, which is located in the core region of HLA-DR locus (i.e. short arm of chromosome 6 p21.3). However, each individual can carry 0, 1 or 2 of HLA-DRB3, HLA-DRB4 and HLA-DRB5 (i.e. at most one DRB3, DRB4 and DRB5 gene per haplotype), but not multiple genes of DRB3, DRB4 and DRB5 at the same time. Moreover, regardless of which locus, the typing of HLA-DRB1 has a strong correlation, and HLA-DRB1 and HLA-DRB3, HLA-DRB4 and HLA-DRB5 have a strong linkage disequilibrium. Therefore, there are differences in the typing of HLA-DRB3, HLA-DRB4 and HLA-DRB5 of different samples (the explanation and description of each case in the typing results of Table 5 are the same as here).

[0091] The above-mentioned reference typing is described in Characterization of 108 Genomic DNA Reference Materials for 11 Human Leukocyte Antigen Loci: A Get-RM Collaboratice Project. (Bettinotti M P, Deborah F, Duke J L, et al. The Journal of Molecular Diagnostics, 2018, DOI: 10.1016 / j.joldx.2018.05.009).

[0092] The numbers in Table 3 and Table 4 refer to the nomenclature of alleles according to the HLA gene, which is used to uniquely identify and distinguish HLA molecules. The nomenclature of HLA alleles is generally composed of four numerical regions, each region number has a specific meaning. The first region identifies the serological antigen type or allelic group of the allele; the second region identifies the allelic subtype within the uniform serological antigen type or allelic group accompanied by amino acid variation; the third region represents the allele without amino acid variation, i.e. the same mutation, which does not affect the amino acid sequence of the protein; the fourth region represents the base variation of the non-coding region. Taking HLA-A gene 01:01:01 as an example, the first "01" indicates that the first allele at this gene locus belongs to the 01 serotype, the second "01" indicates that it has the 01 protein class characteristic, and the third "01" indicates that the gene has the 01 exon synonymous mutation characteristic.

[0093] Comparative Example 1

[0094] This comparative example uses the Hifihla software of the prior art to sequence and type the NA17288 sample. The specific steps are as follows:

[0095] 1. Use the bam2fq program of the smartlink software to convert the split ccs bam data of each sample to fq format;

[0096] 2. Use samtools fqidx to generate the index of fq data;

[0097] 3. Use the cluster function of pbaa software to cluster the fq data;

[0098] 4. Use the hifihla call-consensus function to type the clustered results to obtain the typing results.

[0099] The results are shown in Table 5. In Table 5, the Allele 2 haplotype typing result of the HLA-DRB1 gene of Comparative Example 1 is wrong; three haplotypes of HLA-DRB3 / HLA-DRB4 in Comparative Example 1 appear, in which HLA-DRB4*01:03:01:02N is a wrong identified haplotype.

[0100] Table 5

[0101] .

[0102] It can be seen from the typing results that the method of the present application is consistent with the reference typing, and has higher accuracy compared with the Hifihla software (official software of PacBio).

[0103] From the above description, it can be seen that the above-mentioned embodiments of the present application achieve the following technical effects: by using the HLA typing method of the present application, the advantages of third-generation sequencing can be combined, and the disadvantages of over-reliance on sequencing depth and clustering effect in the prior art HLA analysis method based on third-generation sequencing data are avoided. Through the strategies of comparison, filtering and correction, the HLA typing of the present application has higher resolution and higher accuracy, the algorithm steps are simple, the calculation cost is lower, and it can be applied to scenes such as clinical diagnosis and forensic identification which have very high requirements for data quality. When the HLA typing method of the present application is applied to HLA typing, more accurate genotype information can be provided for clinical transplantation matching, drug metabolism prediction and disease association research.

[0104] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for HLA typing, characterized in that, The method includes: S1) The sequencing data of each sample gene is compared with the HLA gene database to obtain the first candidate HLA genotype set; S2) Filter out the genotyping results in the first candidate HLA genotyping set that have a coverage of less than 30% or an average sequencing depth of less than 30× to obtain the second candidate HLA genotyping set. S3) Based on the consistency sequence of the sample genes corresponding to each typing result in the second candidate HLA genotyping set. The final HLA typing result is obtained by filtering out the typing results in the coding region of the sample gene that differ from the coding region of the corresponding homologous sequence. The sample genes were obtained by sequencing using a third-generation sequencing platform; The gene sequences of the samples were obtained by PacBio HiFi sequencing. The filtering includes: i) The sum of the product of the coverage and sequencing depth of each sample gene locus in each of the genotyping results in each of the first candidate HLA genotyping sets, divided by the sum of all the sequencing depths, is used to obtain the average sequencing depth; The genotyping results with an average sequencing depth of less than 30× are filtered to obtain the first filter set; ii) Sort the genotyping results of each locus in the sample genes according to the average sequencing depth of the loci in the first filter set, and retain the genotyping results of the two loci with the highest average sequencing depth. The genotyping result of the locus with the largest average sequencing depth is recorded as the first genotyping, and the genotyping result of the locus with the second largest average sequencing depth is recorded as the second genotyping; Divide the average sequencing depth of the first genotype by the average sequencing depth of the second genotype to obtain the average sequencing depth ratio. When the average sequencing depth ratio is greater than 2, the first genotype is retained; When the average sequencing depth ratio is less than or equal to 2, the first genotype and the second genotype are retained.

2. The method according to claim 1, characterized in that, The HLA gene database includes IMGT / HLA data.

3. The method according to claim 1, characterized in that, In S3), the method for obtaining the consistency sequence includes: counting the bases that are covered most by gene fragments in the sequencing data of the sample gene at each site in each of the genotyping results of the second candidate HLA genotyping set, and recording them as the correct bases; Each of the correct bases is combined to obtain the consistent sequence.

4. An electronic device for HLA typing, characterized in that, The electronic device includes: a comparison unit, a filtering unit, and a correction unit; The comparison unit is used to compare the sequencing data of each sample gene with the HLA gene database to obtain the first candidate HLA genotype set. The filtering unit is used to filter the genotyping results in the first candidate HLA genotyping set that have a coverage of less than 30% or an average sequencing depth of less than 30×, to obtain the second candidate HLA genotyping set. The correction unit is used to correct the typing results in the second candidate HLA genotyping set to obtain the final HLA typing result; the correction includes: filtering out typing results in the coding region of the sample gene that differ from the coding region of the corresponding homologous sequence in each typing result in the second candidate HLA genotyping set; The sample genes were obtained by sequencing using a third-generation sequencing platform; the filtering unit includes a first filtering unit and a second filtering unit. The first filtering unit is used to filter the genotyping results in the first candidate HLA genotyping set whose average sequencing depth is less than 30×, to obtain the first filter set; The method for calculating the average sequencing depth includes: multiplying the coverage corresponding to the site of each of the sample genes in the genotyping results of each of the first candidate HLA gene genotyping sets by the sequencing depth, and then dividing the sum of the product of the coverage and sequencing depth by the sum of all the sequencing depths to obtain the average sequencing depth; The second filtering unit is used to determine the classification result in the first filtering set; The determination includes: sorting the genotyping results of each locus in the sample genes according to the average sequencing depth of the loci in the first filter set, and retaining the genotyping results of the two loci with the largest average sequencing depth values. The genotyping result of the locus with the largest average sequencing depth is recorded as the first genotyping, and the genotyping result of the locus with the second largest average sequencing depth is recorded as the second genotyping; The average sequencing depth ratio is obtained by dividing the average sequencing depth of the first genotype by the average sequencing depth of the second genotype. When the average sequencing depth ratio is greater than 2, the first genotype is retained; When the average sequencing depth ratio is less than or equal to 2, the first genotype and the second genotype are retained.

5. The electronic device according to claim 4, characterized in that, The HLA gene database includes the IMGT / HLA database.

6. The electronic device according to claim 4, characterized in that, The correction unit includes a consistency sequence acquisition unit; The consistency sequence acquisition unit is used to obtain the consistency sequence of each sample gene in each typing result of each second candidate HLA genotyping set; The method for obtaining the consistent sequence includes: In each genotyping result of the second candidate HLA genotyping set, the base that is most covered by the gene fragment in the sequencing data of the sample gene at each site in each sample gene is recorded as the correct base. Each of the correct bases is combined to obtain the consistent sequence.

7. A computer-readable storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, the method of HLA typing according to any one of claims 1-3 is controlled.

8. A processor, characterized in that, The processor is used to run a program, wherein the program executes the HLA typing method according to any one of claims 1-3.

Citation Information

Patent Citations

  • A method for HLA genotyping based on a third-generation sequencing platform

    CN108460246B

  • Method for typing HLA-DRB gene based on next-generation sequencing data and application

    CN117542412A