HLA typing method and electronic device thereof

Through the strategies of comparison, filtering and correction, the problem of low accuracy of HLA classification results is solved, and high accuracy and high resolution HLA classification is achieved, which is suitable for scenarios with high quality requirements such as clinical and forensic science.

CN120015120AActive Publication Date: 2025-05-16ANNOROAD GENE TECHNOLOGY (BEIJING) CO LTD

Patent Information

Application Number
CN202510499415.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-16
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The accuracy of HLA typing results in the prior art is low, especially when dealing with low frequency or novel HLA alleles, the reliability of the typing results will be greatly reduced.

Method used

By comparing the sequencing data of each sample gene with the HLA gene database, the first candidate HLA genotyping set was obtained, and the coverage and sequencing depth were filtered to obtain the second candidate HLA genotyping set. Then, corrections were performed according to the consensus sequence to obtain the final HLA typing results.

Benefits of technology

It improves the accuracy and consistency of HLA classification results, and can maintain high classification accuracy and high resolution without relying on assembly and clustering, and can achieve at least 6-bit resolution. It is suitable for scenarios such as clinical diagnosis and forensic identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015120A_ABST
    Figure CN120015120A_ABST
Patent Text Reader

Abstract

The invention provides an HLA typing method and an electronic device thereof. The method comprises the following steps: S1) comparing sequencing data of each sample gene with an HLA gene database to obtain a first candidate HLA gene typing set; s2, in the first candidate HLA genetic typing set, typing results with the coverage degree smaller than 30% or the average sequencing depth smaller than 30 * are filtered, and a second candidate HLA genetic typing set is obtained; s3) according to the consistency sequence of the corresponding sample gene in each typing result in the second candidate HLA gene typing set, the typing results which are different from the coding region of the corresponding consistency sequence in the coding region of the sample gene are filtered, a final HLA typing result is obtained, and the sample HLA gene is obtained through sequencing of a three-generation sequencing platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of HLA typing, and in particular to an HLA typing method and an electronic device thereof. Background Art

[0002] The Human Leukocyte Antigen (HLA) system, as part of the human Major Histocompatibility Complex (MHC), is a region of the genome that is closely related to immune function. It is located on the short arm of chromosome 6 and contains multiple closely linked loci, such as HLA-A, HLA-B, HLA-C, etc. These loci encode highly polymorphic proteins and have a crucial impact on the immune response and disease susceptibility between individuals.

[0003] Traditional HLA typing methods, especially those based on serology and cytology, have played an important role in early research, but due to their inherent limitations, such as low resolution, complex and time-consuming operations, they have gradually become unable to meet the needs of current precision medicine. In recent years, with the development of high-throughput sequencing technology, especially the application of second-generation sequencing technology, HLA typing based on sequencing data has become possible. However, this method still has significant defects, namely, the resolution is limited by the length of the read, and when dealing with complex HLA gene regions, the accuracy is often affected by splicing errors caused by short reads.

[0004] In response to the above limitations, the third-generation sequencing technology is considered to be an ideal choice to overcome these problems because of its ability to produce long read lengths and highly accurate data. In particular, PacBio's HiFi sequencing technology, due to its unique long read length characteristics, allows the entire HLA gene region to be sequenced directly, which can theoretically improve the resolution and accuracy of HLA typing. Although PacBio officially provides hifi-hla software to process HiFi data and perform HLA typing, the performance of this method depends to a certain extent on the efficiency and accuracy of the clustering algorithm. When the sequencing depth is insufficient or the clustering effect is poor, the reliability of the typing results will be greatly reduced, especially when dealing with low-frequency or novel HLA alleles. This effect is particularly evident. Summary of the invention

[0005] The main purpose of the present invention is to provide a method for HLA typing and an electronic device thereof to solve the problem of low accuracy of typing results in the prior art.

[0006] In order to achieve the above-mentioned object, according to a first aspect of the present invention, a method for HLA typing is provided, the method comprising: S1) comparing the sequencing data of each sample gene with the HLA gene database to obtain a first candidate HLA genotyping set; S2) filtering the typing results with a coverage of less than 30% or an average sequencing depth of less than 30× in the first candidate HLA genotyping set to obtain a second candidate HLA genotyping set; S3) based on the consensus sequence of the sample gene corresponding to each typing result in the second candidate HLA genotyping set, filtering the typing results in the coding region of the sample gene that are different from the coding region of the corresponding consensus sequence to obtain a final HLA typing result; wherein the sample gene is obtained by sequencing on a third-generation sequencing platform.

[0007] Furthermore, the HLA gene database includes IMGT / HLA data.

[0008] Further, in S2), filtering includes: i) dividing the sum of the coverage and sequencing depth corresponding to the locus of each sample gene in the typing results of each first candidate HLA genotyping set by the sum of all sequencing depths to obtain an average sequencing depth; filtering the typing results with an average sequencing depth lower than 30× to obtain a first filtered set; ii) sorting the typing results of each locus in the sample gene according to the average sequencing depth of the loci of the sample gene in the first filtered set, retaining the typing results of the first two loci with the largest average sequencing depth, the typing result of the locus with the largest average sequencing depth is recorded as the first typing, and the typing result of the locus with the second largest average sequencing depth is recorded as the second typing; dividing the average sequencing depth of the first typing by the average sequencing depth of the second typing to obtain an average sequencing depth ratio; when the average sequencing depth ratio is greater than 2, retaining the first typing; when the average sequencing depth ratio is less than or equal to 2, retaining the first typing and the second typing.

[0009] Furthermore, in S3), the method for obtaining the consensus sequence includes: counting the bases that are most covered by the gene fragments in the sequencing data of the sample gene at each site in each typing result of the second candidate HLA typing set, and recording them as correct bases; and combining each correct base to obtain a consensus sequence.

[0010] In order to achieve the above-mentioned object, according to a second aspect of the present invention, an electronic device for HLA typing is provided, the electronic device comprising: a comparison unit, a filtering unit and a correction unit; wherein the comparison unit is used to compare the sequencing data of each sample gene with the HLA gene database to obtain a first candidate HLA genotyping set; the filtering unit is used to filter the typing results with a coverage of less than 30% or an average sequencing depth of less than 30× in the first candidate HLA genotyping set to obtain a second candidate HLA genotyping set; the correction unit is used to correct the typing results in the second candidate HLA genotyping set to obtain a final HLA typing result; the correction comprises: according to the consensus sequence of the sample gene corresponding to each typing result in the second candidate HLA genotyping set, filtering the typing results in the coding region of the sample gene that are different from the coding region of the corresponding consensus sequence; wherein the sample gene is obtained by sequencing on a third-generation sequencing platform.

[0011] Furthermore, the HLA gene database includes the IMGT / HLA database.

[0012] Furthermore, the filtering unit includes a first filtering unit and a second filtering unit; the first filtering unit is used to filter the typing results with an average sequencing depth lower than 30× in the first candidate HLA genotyping set to obtain a first filtered set; the calculation method of the average sequencing depth includes: dividing the sum of the coverage corresponding to the site of each sample gene in the typing results of each first candidate HLA genotyping set multiplied by the sequencing depth by the sum of all sequencing depths to obtain the average sequencing depth.

[0013] Furthermore, the second filtering unit is used to judge the typing results in the first filtering set; the judgment includes: sorting the typing results of each locus in the sample gene according to the size of the average sequencing depth of the sites of the sample gene in the first filtering set, retaining the typing results of the first two loci with the largest average sequencing depth, the typing result of the locus with the largest average sequencing depth is recorded as the first typing, and the typing result of the locus with the second largest average sequencing depth is recorded as the second typing; dividing the average sequencing depth of the first typing by the average sequencing depth of the second typing to obtain the average sequencing depth ratio; when the average sequencing depth ratio is greater than 2, retaining the first typing; when the average sequencing depth ratio is less than or equal to 2, retaining the first typing and the second typing.

[0014] Furthermore, the correction unit includes a consensus sequence acquisition unit; the consensus sequence acquisition unit is used to obtain the consensus sequence of each sample gene of each typing result in each second candidate HLA genotyping set; the method for obtaining the consensus sequence includes: counting the bases that are most covered by the gene fragments in the sequencing data of the sample gene at each site in each typing result of the second candidate HLA typing set, and recording them as correct bases; combining each correct base to obtain a consensus sequence.

[0015] In order to achieve the above object, according to a third aspect of the present invention, a computer-readable storage medium is provided, the storage medium comprising a stored program, wherein when the program is executed, the above HLA typing method is controlled.

[0016] In order to achieve the above object, according to a fourth aspect of the present invention, a processor is provided, the processor being used to run a program, wherein the above HLA typing method is executed when the program is run.

[0017] By applying the technical solution of the present invention, the sequencing data of each sample gene is compared with the HLA gene database. After the first candidate HLA gene typing set is preliminarily obtained, the coverage and average sequencing depth in the typing results in the typing set are calculated and judged, and the typing results with a coverage of less than 30% or an average sequencing depth of less than 30× are filtered out. The consensus sequence corresponding to the sample gene sequence in the typing result is used for correction, and the HLA typing result with improved accuracy can be obtained. The HLA analysis method of the present application does not require clustering, the algorithm is simple, and it is more suitable for promotion to HLA typing applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings constituting a part of the present application are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0019] Figure 1 A schematic flow chart of an HLA typing method according to an embodiment of the present invention is shown.

[0020] Figure 2 A schematic diagram of an electronic device for HLA typing according to an embodiment of the present invention is shown.

[0021] Figure 3 A hardware structure block diagram of HLA typing according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0022] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below in conjunction with the embodiments.

[0023] Terminology explanation:

[0024] Consensus sequence: refers to the nucleotide sequence with the highest frequency of occurrence at each position in a set of sequence alignment results.

[0025] Resolution of HLA typing: refers to the accuracy of HLA typing, the accuracy of identifying HLA alleles. According to the different typing accuracy, HLA typing resolution can reach 4 main levels, namely low resolution (Low Resolution, 2-bit resolution), intermediate resolution (Intermediate Resolution), high resolution (High Resolution, 4, 6, 8-bit resolution) and allele level (Allelic Level). Among them, low resolution, that is, 2-bit resolution, is based on serology classification (serological classification is defined as HLA antigens are classified based on their ability to react with specific antibodies), or the allele group to which the allele belongs, representing a 2-bit typing result (for example: A02, A03, A11, C03, etc.). The general typing result of medium resolution includes a subset of alleles with the same first field, but excludes some alleles that share the field, which is used for further confirmation of the HLA type of the candidate donor after preliminary screening. High-resolution typing results are defined as a group of alleles that encode the same protein sequence in a region of the HLA protein molecule called the antigen binding site, excluding alleles that are not expressed as cell surface proteins. Allele-level typing provides complete HLA gene sequence information, including the nucleotide sequence of all exons and introns.

[0026] Coverage: The bases covered in a gene divided by the total length of the gene.

[0027] Average sequencing depth: the sum of the depths of base coverage of a gene divided by the number of bases covered.

[0028] As mentioned in the background technology, when the sequencing depth is insufficient or the clustering effect is poor, the reliability of the HLA typing results based on the third-generation sequencing will be greatly reduced, especially when dealing with low-frequency or new HLA alleles. Based on this, the inventors in this application tried to develop a new HLA typing scheme, and thus proposed a series of protection schemes of this application.

[0029] In the first typical embodiment of the present application, a method for HLA typing is provided, which includes S1) comparing the sequencing data of each sample gene with the HLA gene database to obtain a first candidate HLA genotyping set; S2) filtering the typing results with a coverage of less than 30% or an average sequencing depth of less than 30× in the first candidate HLA genotyping set to obtain a second candidate HLA genotyping set; S3) based on the consensus sequence of the sample gene corresponding to each typing result in the second candidate HLA genotyping set, filtering the typing results in the coding region of the sample gene that differ from the coding region of the corresponding consensus sequence to obtain the final HLA typing result; wherein the sample gene is obtained by sequencing on a third-generation sequencing platform, preferably PacBio's HiFi sequencing. The flow diagram of the HLA typing method of the present application is as follows: Figure 1 shown.

[0030] The current HLA typing methods mainly include HLA serotyping, cytological typing, and typing based on second-generation sequencing data. This type of typing method is cumbersome to operate during the experiment, the data analysis cost is high, and the result accuracy and resolution are low. The HLA typing technology based on the third-generation sequencing data has improved sequencing level and can produce data with long read length and high accuracy. Although it can solve the above problems to a certain extent, the method of HLA analysis based on the third-generation sequencing depends on the efficiency and accuracy of the clustering algorithm to a certain extent. When the sequencing depth is insufficient or the clustering effect is not good, the reliability of the typing results will be greatly reduced, especially when dealing with low-frequency or new HLA alleles, the accuracy of the typing results will be greatly affected. Therefore, it is particularly important to develop a new HLA typing method. The HLA typing method of the present application is based on the sequencing accuracy of the third-generation sequencing technology, without the need for assembly and clustering, overcomes the above technical problems, and can achieve ultra-high resolution and high accuracy HLA typing.

[0031] The purpose of splitting the HLA gene in this application is to be able to split the sample amount that meets the requirements for machine testing on the basis of a low sample amount, to be able to more accurately compare the HLA analysis that needs to be analyzed, to simplify the data processing method, and to improve the efficiency of the analysis.

[0032] The HLA typing method of the present application, based on the calculation of the coverage rate and sequencing depth of each site, accurately controls the threshold for filtering inaccurate typing results (coverage less than 30%, average sequencing depth less than 30×), and effectively filters out low-quality data through precise calculation, screening and multiple alignment strategies, improving the accuracy and consistency of typing. In addition, the present application further compares the obtained consensus sequence with the candidate gene, filters out the candidate gene with mutations in the coding region, so as to improve the accuracy of typing, further confirm the correctness of the typing results, and provide more accurate typing results.

[0033] The HLA typing method of the present application achieves high typing accuracy and high resolution without relying on assembly and clustering, and can reach at least 6-bit resolution, which is suitable for scenarios with extremely high data quality requirements such as clinical diagnosis and forensic identification. When the HLA typing method of the present application is applied to HLA typing, it can provide more accurate genotype information for clinical transplant matching, drug metabolism prediction and disease association research.

[0034] High-resolution typing can be further subdivided into: 4-bit resolution: such as A*01:01, which can distinguish specific HLA proteins. 6-bit resolution: such as A*01:01:01, which can distinguish specific HLA coding sequences. 8-bit resolution: such as A*01:01:01:01, which can distinguish specific HLA gene sequences, including untranslated regions and introns.

[0035] The sequencing data of the above-mentioned sample genes are obtained by sequencing on a third-generation sequencing platform, preferably PacBio's HiFi sequencing, or other sequencing methods that can obtain the full-length HLA gene can achieve the effect of the present application.

[0036] Preferably, the above comparison software includes but is not limited to minimap2, pbmm2 or winnowmap2. Preferably, the above splitting software includes but is not limited to lima software.

[0037] In a preferred embodiment, the HLA gene database includes the IMGT / HLA database. The HLA database used for comparison in the present application is updated and supplemented with new HLA typing updated on the IMGT / HLA official website at least every three months.

[0038] In a preferred embodiment, in S2), the method for calculating the average sequencing depth includes: i) dividing the sum of the coverage and sequencing depth corresponding to the locus of each sample gene in the typing results of each first candidate HLA genotyping set by the sum of all sequencing depths to obtain the average sequencing depth; filtering the typing results with an average sequencing depth lower than 30× to obtain the first filtered set; sorting the typing results of each locus in the sample gene according to the size of the average sequencing depth of the loci of the sample gene in the first filtered set, retaining the typing results of the first two loci with the largest average sequencing depth, the typing result of the locus with the largest average sequencing depth is recorded as the first typing, and the typing result of the locus with the second largest average sequencing depth is recorded as the second typing; dividing the average sequencing depth of the first typing by the average sequencing depth of the second typing to obtain the average sequencing depth ratio; when the average sequencing depth ratio is greater than 2, retaining the first typing; when the average sequencing depth ratio is less than or equal to 2, retaining the first typing and the second typing. The average sequencing depth is obtained by dividing the sum of the coverage multiplied by the sequencing depth for each sample gene site in the typing results of each first candidate HLA genotyping set by the sum of all sequencing depths. The formula for this step is: (sum(cover×depth) / sum(depth)), where "sum" means addition, "cover" means coverage, and "depth" means sequencing depth.

[0039] When the average sequencing depth ratio is less than or equal to 2, it means that the average sequencing depth ratio of the two haplotypes is relatively uniform, which meets the expectations of the design of the typing experimental system, so the two haplotypes can be retained; when the average sequencing depth ratio is greater than 2, it means that the average sequencing depths of the two haplotypes differ greatly, which does not meet the expectations of the design of the experimental system, and the haplotype with a smaller average sequencing depth (the second haplotype) is a false positive, so the second haplotype is filtered and the first haplotype is retained. A person skilled in the art can adjust the judgment threshold of the above average sequencing depth ratio according to the experimental system and the average sequencing depths of the first and second haplotypes.

[0040] In a preferred embodiment, in S3), the method for obtaining a consensus sequence includes: counting the bases that are most covered by gene fragments in the sequencing data of the sample gene at each site in each typing result of the second candidate HLA typing set, and recording them as correct bases; and combining each correct base to obtain a consensus sequence.

[0041] The present application screens out low-quality data through multiple alignments, filtering, and the use of consensus sequence correction strategies, thereby improving the accuracy of the final typing results. It has a high resolution and is simple and easy to operate. It avoids the problems of complex steps and inaccurate results of HLA typing based on second-generation sequencing, as well as the drawbacks of HLA typing methods based on third-generation sequencing that rely on sequencing depth and clustering methods. It can provide more accurate genotype information for clinical transplant matching, drug metabolism prediction, and disease association studies, and is more suitable for promotion.

[0042] In a second typical embodiment of the present application, an electronic device for HLA typing is provided (a schematic diagram of each unit of the electronic device is shown in FIG. Figure 2 As shown in FIG. 1 ), the electronic device includes: a comparison unit 01, a filtering unit 02 and a correction unit 03; the comparison unit 01 is used to compare each sample gene with the HLA gene database to obtain a first candidate HLA genotyping set; the filtering unit 02 is used to filter the typing results with a coverage of less than 30% or an average sequencing depth of less than 30× in the first candidate HLA genotyping set to obtain a second candidate HLA genotyping set; the correction unit 03 is used to correct the typing results in the second candidate HLA genotyping set to obtain the final HLA typing result, and the correction includes: according to the consensus sequence of the sample gene corresponding to each typing result in the second candidate HLA genotyping set, filtering the typing results in the coding region of the sample gene that are different from the coding region of the corresponding consensus sequence. Wherein, the sample gene is obtained by sequencing on a third-generation sequencing platform.

[0043] In a preferred embodiment, the HLA gene database includes the IMGT / HLA database.

[0044] In a preferred embodiment, the filtering unit includes a first filtering unit and a second filtering unit; wherein the first filtering unit is used to filter the typing results with an average sequencing depth lower than 30× in the first candidate HLA genotyping set to obtain a first filtered set; the first filtering unit and the second filtering unit; the calculation method of the average sequencing depth includes: dividing the sum of the coverage corresponding to the site of each sample gene in the typing results of each first candidate HLA genotyping set multiplied by the sequencing depth by the sum of all sequencing depths to obtain the average sequencing depth.

[0045] In a preferred embodiment, the second filtering unit is used to judge the typing results in the first filtering set; the judgment includes: sorting the typing results of each locus in the sample gene according to the average sequencing depth of the sites of the sample gene in the first filtering set, retaining the typing results of the first two loci with the largest average sequencing depth, the typing result of the locus with the largest average sequencing depth is recorded as the first typing, and the typing result of the locus with the second largest average sequencing depth is recorded as the second typing; dividing the average sequencing depth of the first typing by the average sequencing depth of the second typing to obtain the average sequencing depth ratio; when the average sequencing depth ratio is greater than 2, retaining the first typing; when the average sequencing depth ratio is less than or equal to 2, retaining the first typing and the second typing.

[0046] In a preferred embodiment, the correction unit includes a consensus sequence acquisition unit; the consensus sequence acquisition unit is used to obtain the consensus sequence of each sample gene of each typing result in each second candidate HLA genotyping set; the method for acquiring the consensus sequence includes: counting the bases that are most covered by the gene fragments in the sequencing data of the sample gene at each site in each typing result of the second candidate HLA typing set, and recording them as correct bases;

[0047] Each correct base is combined to obtain a consensus sequence. In a third typical embodiment of the present application, a computer-readable storage medium is provided, the storage medium includes a stored program, wherein when the program is executed, the above-mentioned HLA typing method is controlled.

[0048] In a fourth typical embodiment of the present application, a processor is provided, and the processor is used to run a program, wherein the above-mentioned HLA typing method is executed when the program is run.

[0049] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the present invention.

[0050] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus hardware devices such as detection devices. Based on such an understanding, the data processing part of the technical solution of the present application can be embodied in the form of a software product, and the computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment of the present application or some parts of the embodiments.

[0051] The present application can be used in many general or special computing system environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices.

[0052] The method provided in this application can be executed in a terminal, a computer terminal or a similar computing device. Taking running on a terminal as an example, Figure 3 FIG. 1 is a hardware structure diagram of an HLA typing terminal according to an embodiment of the present invention. Figure 3 As shown, the terminal may include one or more ( Figure 3 Only one is shown in the figure) processor A1 (processor A1 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory B1 for storing data. Optionally, the terminal may also include a transmission device C1 and an input and output device D1 for communication functions. It can be understood by those skilled in the art that Figure 3 The structure shown is for illustration only and does not limit the structure of the above terminal. Figure 3 More or fewer components as shown, or with Figure 3 Different configurations are shown.

[0053] The memory B1 can be used to store computer programs, for example, software programs and modules of application software, such as computer programs corresponding to methods such as read segment splicing, clustering, and consistency processing in the embodiments of the present invention. The processor A1 executes various functional applications and data processing by running the computer programs stored in the memory B1, that is, implementing the above method. The memory B1 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory B1 may further include a memory remotely arranged relative to the processor A1, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0054] The transmission device C1 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the terminal. In one example, the transmission device C1 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device C1 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0055] Obviously, those skilled in the art should understand that the above-mentioned partial modules or steps of the present application can be implemented in a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented with program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software. The beneficial effects of the present application will be further explained in detail below in conjunction with specific embodiments.

[0056] The beneficial effects of the present application will be further explained in detail below in conjunction with specific embodiments.

[0057] Example 1

[0058] The NA17288 sample standard (Coriell Institute, NA17288) was sequenced and typed as follows:

[0059] 1. Use lima software to split the library of the main HLA gene according to the barcode information of each standard sample, and convert the split bam file into fastq data through the bam2fastq command of the smrtcmds tool to obtain the fastq data (sample gene) of each sample;

[0060] The construction method of the above master HLA gene library includes: 1. Amplifying the full-length sequence of 11 HLA genes through a long-fragment multiplex PCR amplification system; 2. Connecting sequencing adapters at both ends of the PCR product to construct a HiFi library; 3. Binding the library to the sequencing primer and enzyme complex and sequencing on the machine. The HiFi library was sequenced using the PacBio third-generation sequencing platform.

[0061] 2. Use minimap2 software to align fastq data to the reference genome, and count the alignment rate, capture efficiency, coverage depth and other information. The results are shown in Tables 1 and 2.

[0062] Table 1

[0063] .

[0064] Table 2

[0065]

[0066] 3. Use minimap2 software to align fastq data to the HLA database to obtain the aligned bam file, i.e. the first candidate HLA genotyping set, and use samtools software to construct the bam file index;

[0067] Generate coverage reads of each base of each HLA gene in the alignment based on the bam file, calculate the coverage depth and base accuracy of each position, and thus calculate the coverage, average sequencing depth and average alignment accuracy of each gene;

[0068] Candidate HLA genes with coverage less than 30% or average sequencing depth less than 30× were filtered out to obtain a candidate HLA gene set; the specific method and test are as follows:

[0069] (1) Filter out typing results with coverage less than 30% (according to the design of the target gene PCR amplification primers, the minimum coverage of the target gene is greater than 30%, and genes with coverage less than 30% are false positive genes).

[0070] (2) Calculate the weighted average depth of each 6-quartile typing result (representing a specific HLA coding sequence) (i.e., calculate the average sequencing depth of each typing result) using the following calculation method:

[0071] The sum of coverage × sequencing depth corresponding to each 6-quartile typing result is divided by the sum of sequencing depths (sum(cover×depth) / sum(depth)).

[0072] The typing results with an average sequencing depth lower than 30× were filtered to obtain the first filter set;

[0073] The typing results of each locus in the typing results of the first filter set are sorted according to the average sequencing depth, and the typing results of the first two loci with the largest average sequencing depth (two haplotypes for each gene) are retained.

[0074] The typing result with the largest average sequencing depth was recorded as the first typing, and the typing result with the second largest average sequencing depth was recorded as the second typing;

[0075] (3) Dividing the average sequencing depth of the first genotype by the average sequencing depth of the second genotype to obtain an average sequencing depth ratio;

[0076] When the average sequencing depth ratio was greater than 2, the first genotype was retained;

[0077] When the average sequencing depth ratio is less than or equal to 2, the first and second genotypes are retained.

[0078] 4. Generate a consensus sequence for each candidate gene (sample gene) in the HLA gene set, compare it with the candidate gene, filter out genes with mutations in the coding region, and obtain the final correct HLA genotyping.

[0079] The method for obtaining a consensus sequence includes: counting the bases that are most covered by gene fragments in the corresponding HLA gene database at each site in each sample gene in each typing result of the second candidate HLA typing set, and recording them as correct bases; and combining each correct base to obtain a consensus sequence.

[0080] The typing results of this embodiment are shown in the records in Table 5.

[0081] Example 2

[0082] In this example, the sample standard NA17295 (Coriell Institute, NA17295) and the sample standard NA17248 (Coriell Institute, NA17248) were typed using different coverage filtering thresholds, and the remaining steps were consistent with Example 1.

[0083] The typing results of sample NA17295 are shown in Table 3. In Table 3, when the coverage filtering threshold is set to 20%, the Allele2 haplotype typing result of the HLA-DPA1 gene is wrong.

[0084] The typing results of sample NA17248 are shown in Table 4. In Table 4, when the coverage filtering threshold was set to 40%, the Allele2 haplotype of the HLA-DQB1 gene was not identified.

[0085] Table 3

[0086] .

[0087] Table 4

[0088] .

[0089] In Tables 3 and 4, "Allele1" refers to the first allele at the gene locus, and "Allele2" refers to the second allele at the gene locus. Except for HLA-DRB3, HLA-DRB4, and HLA-DRB5, the blank cells of the remaining HLA typing indicate that the gene has only one allele, that is, the gene is homozygous (the genotypes of the two haplotypes are the same), and the other two are heterozygous.

[0090] HLA-DRB3, HLA-DRB4 and HLA-DRB5 are located adjacent to HLA-DRB1. All individuals carry HLA-DRB1, which is located in the core region of the HLA-DR locus (i.e., p21.3 on the short arm of chromosome 6). For HLA-DRB3, HLA-DRB4 and HLA-DRB5, each individual may carry 0, 1 or 2 genes (i.e., each haplotype carries at most one gene of DRB3, DRB4 and DRB5), but will not carry multiple genes of DRB3, DRB4 and DRB5 at the same time. Moreover, no matter which locus, it is strongly correlated with the typing of HLA-DRB1, and HLA-DRB1 and HLA-DRB3, HLA-DRB4 and HLA-DRB5 have extremely strong linkage disequilibrium. Therefore, there are differences in the typing of HLA-DRB3, HLA-DRB4 and HLA-DRB5 for different samples (the explanation and description of each situation in the typing results of Table 5 are consistent here).

[0091] The above-mentioned reference typing is the typing described in Characterization of 108 Genomic DNA Reference Materials for 11 Human Leukocyte Antigen Loci: A Get-RM Collaboratice Project. (Bettinotti MP, Deborah F, Duke JL, et al. The Journal of Molecular Diagnostics, 2018, DOI: 10.1016 / j.joldx.2018.05.009).

[0092] The numbers in Tables 3 and 4 refer to the naming rules for alleles of the HLA gene, which are used to uniquely identify and distinguish HLA molecules. The naming of HLA alleles usually consists of four digital regions, and the numbers in each region have a specific meaning. The first region identifies the serological antigen type or allele group of the allele; the second region identifies the allele subtype with amino acid variation within the unified serological antigen type or allele group; the third region indicates alleles without amino acid variation, that is, the same mutation, which does not affect the amino acid sequence of the protein; the fourth region indicates base variation in the non-coding region. Taking the HLA-A gene, 01:01:01 as an example, the first "01" indicates that the first allele at the gene locus belongs to the 01 serotype typing, the second "01" indicates that it has the characteristics of the 01 protein category, and the third "01" indicates that the gene has the 01 exon synonymous mutation characteristics.

[0093] Comparative Example 1

[0094] This comparative example uses the existing Hifihla software to sequence and type the NA17288 sample. The specific steps are:

[0095] 1. Use the bam2fq program of SmartLink software to convert the CCS BAM data of each sample after splitting into FQ format;

[0096] 2. Use samtools fqidx to generate the index of fq data;

[0097] 3. Use the cluster function of the pbaa software to cluster the fq data;

[0098] 4. Use the hifihla call-consensus function to classify the clustering results and obtain the typing results.

[0099] The results are shown in Table 5. In Table 5, the Allele2 haplotype typing result of the HLA-DRB1 gene in Comparative Example 1 is wrong; three haplotypes appeared in HLA-DRB3 / HLA-DRB4 in Comparative Example 1, among which HLA-DRB4*01:03:01:02N was an erroneously identified haplotype.

[0100] Table 5

[0101] .

[0102] It can be seen from the typing results that the method of this application is consistent with the reference typing and has higher accuracy than the Hifihla software (PacBio official software).

[0103] From the above description, it can be seen that the above embodiments of the present invention achieve the following technical effects: the HLA typing method of the present application can combine the advantages of the third-generation sequencing, and avoid the drawbacks of the HLA analysis method based on the third-generation sequencing data in the prior art, which is too dependent on the sequencing depth and clustering effect. Through the strategies of comparison, filtering and correction, the HLA typing of the present application has a higher resolution and higher accuracy, the algorithm steps are simple, and the calculation cost is lower. It can be applied to clinical diagnosis, forensic identification and other scenes with extremely high data quality requirements. When the HLA typing method of the present application is applied to HLA typing, it can provide more accurate genotype information for clinical transplant matching, drug metabolism prediction and disease association research.

[0104] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for HLA typing, characterized in that: The method comprises: S1) comparing the sequencing data of each sample gene with the HLA gene database to obtain the first candidate HLA genotyping set; S2) filtering the typing results with coverage less than 30% or average sequencing depth less than 30× in the first candidate HLA genotyping set to obtain a second candidate HLA genotyping set; S3) according to the consensus sequence of the sample gene corresponding to each typing result in the second candidate HLA genotyping set, Filtering the typing results in the coding region of the sample gene that are different from the coding region of the corresponding consensus sequence to obtain the final HLA typing result; The sample genes were sequenced using a third-generation sequencing platform.

2. The method according to claim 1, characterized in that The HLA gene database includes IMGT / HLA data.

3. The method according to claim 1, characterized in that In S2), the filtering includes: i) dividing the sum of the coverage and sequencing depth corresponding to each locus of the sample gene in the typing results of each of the first candidate HLA genotyping sets by the sum of all the sequencing depths to obtain the average sequencing depth; Filter the typing results with an average sequencing depth lower than 30× to obtain a first filtered set; ii) sorting the typing results of each locus in the sample gene according to the average sequencing depth of the loci of the sample gene in the first filter set, and retaining the typing results of the first two loci with the largest average sequencing depth, The typing result of the gene locus with the largest average sequencing depth is recorded as the first typing, and the typing result of the gene locus with the second largest average sequencing depth is recorded as the second typing; Dividing the average sequencing depth of the first genotype by the average sequencing depth of the second genotype to obtain an average sequencing depth ratio; When the average sequencing depth ratio is greater than 2, retaining the first genotype; When the average sequencing depth ratio is less than or equal to 2, the first genotype and the second genotype are retained.

4. The method according to claim 1, characterized in that In S3), the method for obtaining the consensus sequence includes: counting the bases most covered by the gene fragments in the sequencing data of the sample gene at each site in each typing result of the second candidate HLA typing set, and recording them as correct bases; Each of the correct bases is combined to obtain the consensus sequence.

5. An electronic device for HLA typing, characterized in that: The electronic device comprises: a comparison unit, a filtering unit and a correction unit; The comparison unit is used to compare the sequencing data of each sample gene with the HLA gene database to obtain a first candidate HLA genotyping set; The filtering unit is used to filter the typing results with a coverage of less than 30% or an average sequencing depth of less than 30× in the first candidate HLA genotyping set to obtain a second candidate HLA genotyping set; The correction unit is used to correct the typing results in the second candidate HLA genotyping set to obtain a final HLA typing result; the correction comprises: according to the consensus sequence of the sample gene corresponding to each typing result in the second candidate HLA genotyping set, filtering the typing results in the coding region of the sample gene that are different from the coding region of the corresponding consensus sequence; The sample genes were sequenced using a third-generation sequencing platform.

6. The electronic device according to claim 5, characterized in that: The HLA gene database includes the IMGT / HLA database.

7. The electronic device according to claim 5, characterized in that: The filter unit includes a first filter unit and a second filter unit; The first filtering unit is used to filter the typing results with an average sequencing depth lower than 30× in the first candidate HLA genotyping set to obtain a first filtered set; The method for calculating the average sequencing depth includes: dividing the sum of the coverage and the sequencing depth corresponding to each site of the sample gene in the typing results of each of the first candidate HLA genotyping sets by the sum of all the sequencing depths to obtain the average sequencing depth.

8. The electronic device according to claim 7, characterized in that: The second filtering unit is used to determine the typing result in the first filtering set; The judgment comprises: sorting the typing results of each locus in the sample gene according to the average sequencing depth of the loci of the sample gene in the first filter set, retaining the typing results of the first two loci with the largest average sequencing depth, The typing result of the gene locus with the largest average sequencing depth is recorded as the first typing, and the typing result of the gene locus with the second largest average sequencing depth is recorded as the second typing; Dividing the average sequencing depth of the first genotype by the average sequencing depth of the second genotype to obtain the average sequencing depth ratio; When the average sequencing depth ratio is greater than 2, retaining the first genotype; When the average sequencing depth ratio is less than or equal to 2, the first genotype and the second genotype are retained.

9. The electronic device according to claim 5, characterized in that: The correction unit includes a consistency sequence acquisition unit; The consensus sequence acquisition unit is used to obtain the consensus sequence of each sample gene of each typing result in each of the second candidate HLA genotyping sets; The method for obtaining the consensus sequence includes: Counting the bases most covered by the gene fragments in the sequencing data of each sample gene at each site in each typing result of the second candidate HLA typing set, and recording them as correct bases; Each of the correct bases is combined to obtain the consensus sequence.

10. A computer-readable storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is executed, the HLA typing method according to any one of claims 1 to 5 is controlled.

11. A processor, characterized in that: The processor is used to run a program, wherein the program, when running, executes the HLA typing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Three-generation sequencing platform-based HLA genotyping method

    CN108460246A

  • A method for HLA genotyping based on a third-generation sequencing platform

    CN108460246B

  • Human leukocyte antigen typing method and device

    CN111798924A

  • HLA typing method based on next-generation sequencing data

    CN113409890A

  • Method for typing HLA-DRB gene based on next-generation sequencing data and application

    CN117542412A

Cited By

  • Preparation method of human HLA typing standard substance and HLA typing standard substance library

    CN120350095A