Method for carrying out automatic analysis and typing on HLA (human leukocyte antigen) based on third-generation sequencing
By optimizing algorithms and automated processes, combining multi-source data and dynamic threshold judgment, the accuracy and efficiency of HLA typing in nanopore sequencing are solved, and fast and accurate HLA typing analysis is achieved, which is suitable for clinical testing.
Patent Information
- Application Number
- CN202510332649.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-20
AI Technical Summary
It is difficult for the prior art to perform HLA classification quickly and accurately, especially in nanopore sequencing data, which have problems such as high comparison difficulty, high error rate, and inflexible threshold judgment, resulting in inaccurate classification results.
Two algorithms are used to initially judge the HLA type, combined with minimap2, samtools and Megalodon algorithms for data filtering and comparison, and used blastn software to compare with the IMGT database, and introduced automated processes and dynamic threshold judgments to correct sequencing errors and quality control to ensure the accuracy of the typing results.
It realizes fast and accurate HLA typing based on third-generation sequencing data, supports high-resolution comparison of nanopore long read sequencing, reduces error comparison, adapts to heterozygous base judgment in complex situations, improves the timeliness and throughput of clinical applications, and provides a comprehensive report on typing results.
Smart Images

Figure CN120299511A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and more particularly, to a method for automatically analyzing and typing HLA based on third-generation sequencing. Background Art
[0002] For HLA identification methods based on high-throughput sequencing data, the main technologies and implementation schemes are to perform experimental steps such as sample pretreatment, nucleic acid extraction, amplification, and sequencing library construction on the sample, use a high-throughput sequencer to measure the amplification products of the nucleic acid sequences in the sample, and after obtaining the nucleic acid sequence data output by the sequencer, use software such as sequence assembly, clustering, and alignment to compare and analyze these sequences with the HLA sequence database, and screen using certain alignment result filtering conditions to obtain reliable results, and then calculate the number of sequences belonging to each HLA typing in the nucleic acid sequence data of the sample and the proportion in all sequences, etc., and finally judge the presence of HLA typing in the detected sample.
[0003] Since there are many types of HLA sample typing and a large amount of data, the main technical difficulties in the rapid detection and analysis of clinical samples currently lie in the rapid assembly of input data, accurate alignment, haploid typing analysis, while ensuring accuracy and minimizing dependence on hardware performance. Therefore, in order to improve the throughput, resolution, timeliness, and accuracy of clinical HLA typing detection, it is necessary to establish a rapid alignment and analysis method and system for HLA typing nanopore sequencing data.
[0004] Due to its characteristics such as long read lengths, high throughput, and fast typing speed, HLA typing based on nanopore sequencing data is gradually replacing first-generation and second-generation sequencing as the mainstream sequencing method for HLA typing. Currently, R & D personnel have developed various typing methods based on second-generation sequencing technology. For example, Chinese Patent CN103221551A discloses an HLA genotype-SNP linkage database, its construction method, and an HLA typing method. It is a method that selects the HLA locus sequence as a reference sequence, finds SNP sites, obtains the linkage relationship between HLA and the reference sequence to construct a database, and finally performs HLA typing based on the SNP linkage relationships of different types. Chinese Patent Application CN109477143A discloses a method for typing human leukocyte antigens, which uses multiple sequence alignment (MSA) of known HLA allele reference sequences to determine the HLA allele reference sequence, and finally selects the reference sequence closest to the match of the individual to complete HLA typing. This method produces more accurate results for class I and class II HLA typing. Chinese Patent Application CN113409890A discloses an HLA typing method based on second-generation sequencing data. By constructing an HLA reference gene sequence database, generating a sequence coordinate mapping vector, performing quality control and filtering on the original data and then aligning, and then separating heterozygous and homozygous base positions through a clustering algorithm to obtain haploid sequences, and finally obtaining the typing result. International Patent WO2022 / 116456 provides a method and device for analyzing second-generation sequencing data of HLA chromosomal region heterozygous deletion to automatically detect whether HLA chromosomal region heterozygous deletion occurs in a sample. Currently, commonly used second-generation sequence alignment software includes bwa, bowtie 2, SOAP, BLAST, etc. Chinese Patent Application CN201810191663 discloses an HLA gene typing method based on a third-generation sequencing platform. Specifically, after the product obtained by PCR is qualified for detection, third-generation sequencing is performed to obtain the original data, which is aligned with the reference gene sequence for a long sequence alignment and then the sequencing errors are corrected. The haplotype sequence is obtained through phasing and then the typing judgment is carried out.
[0005] Although there are various HLA typing methods in the prior art, there are still the following problems: 1. Most of these current solutions do not support HLA typing of nanopore long read length sequencing. 2. High-resolution HLA typing is not supported in many solutions. 3. The HLA gene has high polymorphism, high sequence similarity, and many repetitive sequences, making the alignment difficult and resulting in a relatively large number of incorrect alignments, and it is difficult to ensure the accuracy of the typing result. 4. The method of using a set threshold to judge whether it is a heterozygous base is not flexible enough and is prone to errors in complex situations. In summary, how to achieve automatic analysis of HLA typing based on nanopore sequencing data remains an issue to be solved. Summary of the Invention
[0006] The object of the present invention is to solve the deficiencies in the prior art and provide a method for automatic analysis and typing of HLA based on third-generation sequencing, so as to solve the problem of how to quickly and accurately determine the HLA typing in a test sample based on third-generation sequencing data.
[0007] First, the HLA types of the first sequencing read data of different data types of the test sample are preliminarily judged by two different algorithms, then the second sequencing read data is selected, and then the second sequencing read data of the identified different HLA types is compared and verified with the corresponding HLA reference sequences and screened, so as to determine the HLA types existing in the test sample.
[0008] To achieve the above object, the present invention is realized through the following technical solutions: A method for automatic analysis and typing of HLA based on third-generation sequencing, the method comprising the following steps:
[0009] Step S1, splitting the HLA gene to be tested according to the sample sequence to obtain the original off-machine data;
[0010] Step S2, performing quality assessment on the original off-machine data obtained in step S1;
[0011] Step S3, preprocessing the data that has completed quality assessment in step S2, including removing reads with an average data quality lower than Q15, truncating the primer sequences designed for the HLA gene on both sides of the read, and retaining sequences with a read length between 1000bp and 30000bp;
[0012] Step S4, using the minimap2 algorithm to align the preprocessed data obtained in step S3 to the HLA-related reference genome, filtering out sequences that are not aligned to the HLA-related reference genome, further filtering out reads with an average alignment data quality lower than Q30, then outputting the data, and then calculating the coverage depth of the HLA gene in the target region according to the obtained data;
[0013] Step S5, performing sequence filtering on the data output in step S4, retaining sequences with a sequence length between HLA amplicon design region length - 1000bp and + 1000bp, and respectively performing downsampling on the filtered data according to a coverage depth of 2000 to obtain preliminarily filtered data;
[0014] Step S6, using samtools and the Megalodon algorithm, according to the HLA-related reference genome, performing mutation detection on the preliminarily filtered data obtained in step S5 to respectively obtain two mutation site detection results, then merging the two mutation site detection results, filtering out mutation sites with a base depth lower than 30, and then performing haplotype splitting according to the mutation detection results to obtain split data;
[0015] Step S7: Perform cluster analysis on the split data obtained in step S6. Set the required minimum read length, and cluster the reads longer than this length to generate different cluster sequences, where the length of each cluster sequence fragment is not less than 1000, and the number of sequences in each cluster sequence is not less than 30 reads, to obtain the data after cluster analysis;
[0016] Step S8: Perform gene assembly on the data after cluster analysis obtained in step S7;
[0017] Step S9: Correct the sequencing errors of the sequences obtained in step S8 to obtain the corrected consensus sequence;
[0018] Step S10: Extract the CDS region of the corrected sequence, and perform alignment using the blastn software and the reference IMGT database to obtain the genotyping result;
[0019] Step S11: If the total number of sequences in the genotyping result is less than 300, filter it out; otherwise, perform quality inspection on the genotyping result;
[0020] Step S12: Perform data quality control on the data after genotyping quality inspection obtained in step S11. Remove bases with an average data quality lower than Q15, primer dimer sequences, repetitive sequences, and adapter sequences, and then use the reference sequence of the IPD-IMGT / HLA database to align with the consensus sequence of the haploid genotyping obtained in step S10 to obtain the analysis result of the HLA allele genotype and complete the analysis.
[0021] Preferably, the specific steps of step S1 are as follows: Load the HLA gene to be tested onto the sequencing chip, start the ONT sequencing software Minknow to convert the electrical signals generated by nanopore sequencing into unclassified data, and then automatically split according to the sample barcode sequence to obtain the original off-machine data.
[0022] Preferably, the quality assessment in step S2 includes average read length, average data quality, sequence median length, sequence median quality, number of sequences, N50 read length, total number of bases, data quality distribution, and sequence length distribution.
[0023] Preferably, the HLA-related reference genomes in step S3 include HLA-A, HLA-B, HLA-C, HLA-DQA1, HLA-DQB1, HLA-DPA1, HLA-DPB1, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5.
[0024] Preferably, the haplotype splitting according to the HLA gene type in step S5 is specifically as follows:
[0025] Extract heterozygous mutations with an average data quality of the variant QUAL > 15 and a single nucleotide variant depth greater than 25% of the average depth.
[0026] Extract homozygous mutations with an average data quality of the variant QUAL > 15 and an SNV frequency greater than 90%.
[0027] If the number of heterozygous mutations exceeds 25% of the total number of mutations, extract the bases at the corresponding heterozygous mutation positions in the read sequences to form a new mutation combination sequence; if the number of heterozygous mutations is less than 25%, randomly select the bases at the homozygous mutation positions to form a new mutation combination sequence until its length exceeds 25% of the total number of mutations.
[0028] If there are no heterozygous SNVs, skip the analysis in step S7 and directly proceed to step S10.
[0029] Preferably, the specific steps of step S8 are as follows: Cluster the sequences in the first two clustering groups of each amplicon region in step S7, select sequences with a length ranging from amplicon - 1000bp to + 1000bp, downsample to a coverage of 500X, and use flye and canu software for genome assembly. Select the minimum value of the amplicon length - 1000bp or 10000bp as the minimum assembly overlap region for assembly.
[0030] Preferably, the quality inspection of the genotyping results in step S11 includes evaluation of amplification imbalance, evaluation of linkage disequilibrium, homozygous typing heterozygous analysis, serological conversion, differentiation of rare typing, identification of new typing, and visual alignment.
[0031] The present invention solves many problems in the prior art of HLA typing by nanopore sequencing through methods such as optimizing algorithms, integrating multi - source data, introducing automated processes, and dynamic threshold judgment, and achieves the following beneficial effects:
[0032] The method of the present invention can effectively process long read - length data generated based on third - generation sequencing technology, reduce the computational amount of accurate alignment, and can achieve accurate analysis and report generation of typical third - generation HLA data within 15 minutes, meeting the need for rapid HLA detection and analysis; at the same time, the method can also analyze short read - length data obtained by second - generation sequencing, having good data compatibility.
[0033] The present invention supports high-resolution HLA typing for nanopore long-read sequencing; not only improves alignment accuracy and reduces incorrect alignments, but also can flexibly handle the determination of heterozygous bases and adapt to complex situations; enhances the timeliness and throughput of clinical applications, meets the needs of large-scale screening; and provides a comprehensive typing result report to assist clinical decision-making. Brief Description of the Drawings
[0034] Figure 1 It is a flowchart of a method for automatically analyzing HLA typing;
[0035] Figure 2 It is an output example of quality assessment in the method for automatically analyzing HLA typing;
[0036] Figure 3 It is for minimap2 to perform positioning by constructing an index to obtain a chain;
[0037] Figure 4 It is a flowchart of the minimap2 algorithm implementation;
[0038] Figure 5 It is a flowchart of data quality control in the method for automatically analyzing HLA typing Detailed Embodiments
[0039] The present invention will be further described below in conjunction with embodiments, but it shall not be used as a basis for limiting the present invention.
[0040] A method for automatically analyzing and typing HLA based on third-generation sequencing is as follows:
[0041] Step S1, data acquisition and splitting
[0042] Load the HLA gene to be tested onto the sequencing chip, start the ONT sequencing software Minknow, select the dna_r10.4.1_e8.2_400bps_hac@v4.3.0 model, convert the electrical signals generated by nanopore sequencing into unclassified data fastq, and then automatically split according to the sample barcode sequence to obtain the original off-machine data.
[0043] Step S2, quality assessment
[0044] The obtained raw off-machine data is subjected to quality assessment using the open-source software nanostat, including: average read length, average data quality, median sequence length, median sequence quality, number of sequences, N50 read length, total number of bases, data quality distribution, sequence length distribution, etc.; if the assessment results show that the quality of the data results is lower than expected (Q30 > 80%, the minimum sequencing depth requirement is 20X, the average sequencing depth is 1000X, the sequence contig length or average reads length is 3000 - 8000bp), especially if the data quality distribution is uneven or the sequence length distribution is abnormal, then the library needs to be rebuilt or sequencing needs to be redone.
[0045] Step S3, data preprocessing
[0046] The data that has completed quality assessment is preprocessed, including removing reads with an average data quality lower than Q15, trimming the primer sequences corresponding to the HLA gene designed on both sides of the read, and retaining sequences with a read length between 1000bp and 30000bp.
[0047] Step S4, data alignment
[0048] The preprocessed data is aligned to the HLA-related reference genome using the minimap2 algorithm, setting the map-ont parameter, that is, using the ordinary minimizer as the seed. The specific method is to establish a minimizer index table according to the reference genome: store the minimizer index table of the genomic sequence in a hash table (minimizer refers to the seed with the minimum hash value within a sequence, which is to establish an index for the reference genome); for each sequence to be aligned (that is, the reads of nanopore sequencing), find all the minimizers of the sequence to be aligned, establish an anchor table according to the position of x in minimizer[x - w + 1, x]; connect the anchors according to the optimal chain score to perform positioning on the reference genome; then use dynamic programming to amplify the non-seed regions. As Figure 3As shown, a minimizer index table is established with reference to the genome. Each read has a sorted anchor table through alignment with the minimizer. A minimizer may find many matching segments in the reads. Therefore, according to the scoring, the qualified anchors are connected to form a chain, thereby locating on the reference genome. Sequences that are not aligned to the HLA-related reference genome are filtered out, and reads with an average numerical quality lower than Q30 are further filtered out. Based on the obtained data, calculate the coverage depth of the HLA gene target region, including: gene name, start point, end point, number of sequences, percentage of sequences passing the screening (screening principle: the sequence length reaches 80% of the target region length), number of covered bases, average coverage, average depth, average base quality, and average alignment quality, etc.
[0049] Table 1: Output list of the coverage depth of the target region
[0050]
[0051] Step S5, collect sampling sequences
[0052] Filter the data obtained in step S4 for sequences, retain sequences with a length ranging from HLA amplicon design region length - 1000bp to +1000bp, and perform downsampling on the filtered data according to a coverage depth of 2000 respectively through seqkit to obtain a preliminarily filtered bam file.
[0053] Step S6, variant detection
[0054] Use the HLA-related reference genome, and use samtools and the Megalodon algorithm to detect mutations in the preliminarily filtered bam file obtained in step S5. Combine the mutation site detection results of the two algorithms, and filter out mutation sites with a base depth lower than 30. Split the haplotypes of the preliminarily filtered bam file obtained in step S5 according to the mutation site detection results. The processing method for each HLA gene is as follows:
[0055] #1. Extract heterozygous mutations with an average data quality of QUAL > 15 and a single nucleotide variant depth greater than 25% of the average depth.
[0056] #2. Extract homozygous mutations with QUAL > 15 and an SNV frequency greater than 90%.
[0057] #3. If the number of heterozygous mutations exceeds 25% of the total number of mutations, extract the bases corresponding to the heterozygous mutation positions in the read sequences to form a new mutant combination sequence. If it is less than 25%, randomly select the bases at the homozygous mutation positions to form a new mutant combination sequence until its length exceeds 25% of the total number of mutations.
[0058] #4. If there are no heterozygous SNVs, do not perform the analysis of S7, and directly generate a bam file for re-alignment with the CDS region.
[0059] Step S7, clustering analysis
[0060] Perform clustering analysis on the mutant combination sequences obtained in the previous step from the data in #1, #2, and #3. Set the required minimum read length, and cluster the reads longer than this length to generate different clusters, where the length of each cluster sequence segment is not less than 1000, and the number of sequences in each cluster is not less than 30 reads. This process can also consider using Longshot for phasing (genotyping). Longshot can perform phasing on SNVs (single nucleotide variations), insertions, deletions, MNP (multiple adjacent SNVs), and "complex" variations.
[0061] Step S8, genome assembly
[0062] For the sequences in the first two clustering groups of each amplicon region in step S7, select sequences with lengths from amplicon - 1000bp to +1000bp, downsample to a coverage of 500X, and use the flye and canu software for genome assembly. Select the minimum value of the amplicon length - 1000bp or 10000bp as the minimum assembly overlap region for assembly.
[0063] Step S9, sequencing error correction
[0064] Due to the relatively high error rate of nanopore sequencing, these errors usually manifest as insertions, deletions, and substitution errors, which makes the sequences directly obtained by nanopore sequencing may not be completely accurate. Therefore, in order to improve the accuracy of sequencing data, additional error correction measures must be taken. It is necessary to correct the error sequences and improve the accuracy of the final assembly through reference genome alignment, overlapping information between reads, and machine learning models. Use the Nanopolish software to perform sequencing error correction on the sequences obtained in the previous step to obtain the corrected consensus sequence.
[0065] Step S10, genotyping result comparison
[0066] Extract the CDS region of the corrected sequence and align it using the Blastn software against the reference IMGT database. Check whether the HLA type matches the four-part format, such as DQB1*05:02:01:01, intercept the first three parts and return. Filter the results of non-complete alignment, preferentially extract the alignment result with the fewest number of mismatches and gaps and add a Tag, merge the results with the same three-domain results, calculate how many results are merged, extract the one with the most merged results, and store it in the cluster alignment result of a single locus, so as to obtain the HLA typing result of each gene haplotype.
[0067] Among them, the blastn processing logic:
[0068] #1. Filter the blastn initial file for complete alignment
[0069] #2. Extract data according to the priority of (0, 0)(1, 0)(0, 1). If there is no result of (0, 0), the script will automatically extract the results of (1, 0)(0, 1).
[0070] # Extract the one with the largest number of merged results as the optimal solution of the current cluster and submit it to the next level. All merged results will be retained in Blast / clsuter_no / sampleid_Cluster_X_blastresult_Filter.xls;
[0071] # If there are no results of (1, 0)(0, 1), the script will extract and submit the sub-optimal solution according to the ascending combination of mismatches and gaps. The results will still be retained in bam / barcodeX_Cluster_X_blastresult_Filter.xls
[0072] #3. The optimal solution of each cluster will be submitted to the amplicon level, and the data will be stored in amplicon / amplicon_blastresults.xls
[0073] #4. The results of each amplicon will be uniformly included in the HLA typing result of the barcode, arranged in each row in the order of cluster1 / 2 / 3 / 4 / 5.
[0074] # Combine with the clustering statistics results of cluster_stat.txt, and the data will be stored in Blast / sampleid_HLA_Result.xls
[0075] #5. Perform the final step of data merging on sampleid_HLA_Result.xls. Merge the results with exactly the same typing and tags, and add the proportions of the corresponding clusters to obtain the final result.
[0076] # The data is stored in Blast / sampleid_HLA_result.txt
[0077] Step S11, Quality inspection of typing results
[0078] If the total number of sequences in the typing result is less than 300, the result will be automatically filtered. Then, quality inspection of the typing results is carried out, specifically including: evaluation of amplification imbalance, evaluation of linkage disequilibrium, analysis of homozygous typing and heterozygosity, serological conversion, differentiation of rare typing, identification of new typing, visual alignment, etc. It is written in R language, including merging duplicate results, adding the sequence numbers and clusters, and taking the cluster with the largest number of sequences. The filtering index takes the top two clusters and the sequence number of the second-ranked cluster is not less than 1 / 5 of the sequence number of the first-ranked cluster, which is limited according to HLA biological theory. For example, DRB345 has no -1 -2, and at most 2 types of DRB345 can exist at the same time. Those with only 1 cluster are directly formed into diploid typing. The CWD frequency determines rare typing. It can comprehensively evaluate the accuracy and reliability of the typing results and ensure its applicability to clinical diagnosis and research.
[0079] Step S12, Data quality control
[0080] By removing bases with an average value quality lower than Q15, primer dimer sequences, repetitive sequences, removing adapter sequences (tag sequences, primer sequences, sequencing primer binding sequences), and comparing with the reference sequences of the current version of the IPD-IMGT / HLA database, the alignment of the sequencing sequences with the reference database is achieved, and data processing after alignment is carried out to obtain the analysis results of HLA allele genotypes and complete the secondary analysis of the data. The data output for each sample is 250M, and the processes of primary analysis and secondary analysis of the data.
[0081] Requirements for HLA typing data quality control, including Q30 > 80%, sequencing depth requirement: minimum sequencing depth of 20X, average sequencing depth of 1000X, sequence contig length or average read length of 3000 - 8000 bp. Among them, the sequencing uniformity is that the read fragments are evenly distributed in the detection area; the sequencing coverage is that the sequencing sequences of each gene locus cover 100% of the target area, especially the key exon regions are completely covered and reach the required sequencing depth. Allele balance refers to the ratio of two bases at the polymorphic position of alleles, distinguishing homozygous or heterozygous sites; when the ratio of the two bases is similar, it is judged as allele balance, and when it is lower than the acceptable threshold, which is set to 1 / 3, it is judged as allele imbalance. Currently, 50% is used as the cutoff to determine mismatches (subsequent experiments are needed to define parameters and the gray area).
[0082] The bioinformatics quality control standards for HLA gene detection include sample quality control, in-batch positive and negative control product and blank control sample (NTC) quality control. The specific process is shown in Figure 5 。
[0083] The above shows and describes the basic principles, main features and advantages of the present invention. However, the above are only specific embodiments of the present invention, and the technical features of the present invention are not limited thereto. Any other embodiments obtained by those skilled in the art without departing from the technical solution of the present invention should be covered within the patent scope of the present invention.
Claims
1. A method for automatic analysis and typing of HLA based on third-generation sequencing, characterized in that, The method includes the following steps: Step S1: Split the HLA gene to be tested according to the sample sequence to obtain the original off-machine data; Step S2: Perform quality assessment on the original off-machine data obtained in Step S1; Step S3: Preprocess the data that has completed quality assessment in Step S2, including removing reads with an average data quality lower than Q15, intercepting the primer sequences corresponding to the HLA gene on both sides of the read, and retaining sequences with a read length between 1000bp and 30000bp; Step S4: Use the minimap2 algorithm to align the preprocessed data obtained in Step S3 to the HLA-related reference genome, filter out sequences that are not aligned to the HLA-related reference genome, further filter out reads with an average alignment data quality lower than Q30, then output the data, and calculate the coverage depth of the HLA gene in the target region based on the obtained data; Step S5: Perform sequence filtering on the data output in Step S4, retain sequences with a length between HLA amplicon design region length - 1000bp and +1000bp, and perform downsampling on the filtered data at a coverage depth of 2000 respectively to obtain preliminarily filtered data; Step S6: Use the samtools and Megalodon algorithms to perform mutation detection on the preliminarily filtered data obtained in Step S5 according to the HLA-related reference genome, obtain two mutation site detection results respectively, then merge the two mutation site detection results, filter out mutation sites with a base depth lower than 30, and then perform haplotype splitting according to the mutation detection results to obtain split data; Step S7: Perform cluster analysis on the split data obtained in Step S6, set the required minimum read length, and cluster reads longer than this length to generate different cluster sequences, where the length of each cluster sequence fragment is not less than 1000, and the number of sequences in each cluster sequence is not less than 30 reads, to obtain data after cluster analysis; Step S8: Perform gene assembly on the data after cluster analysis obtained in Step S7; Step S9: Correct the sequencing errors of the sequences obtained in Step S8 to obtain the corrected consensus sequence; Step S10: Extract the CDS region of the corrected sequence and perform alignment using the blastn software and the reference IMGT database to obtain the typing result; Step S11: If the total number of sequences in the typing result is less than 300, filter them out, otherwise, perform quality inspection on the typing result; Step S12: Perform data quality control on the data after typing quality inspection obtained in Step S11, remove bases with an average data quality lower than Q15, primer dimer sequences, repetitive sequences, and adapter sequences, and then use the reference sequence of the IPD-IMGT / HLA database to align with the consensus sequence of the haploid typing obtained in Step S10 to obtain the analysis result of the HLA allele type and complete the analysis.
2. The method for automatic analysis and typing of HLA based on third-generation sequencing according to claim 1, wherein, Step S1 is specifically as follows: Load the HLA gene to be tested onto the sequencing chip, start the ONT sequencing software Minknow to convert the electrical signals generated by nanopore sequencing into unclassified data, and then automatically split according to the sample barcode sequence to obtain the original off-machine data.
3. The method for automatic analysis and typing of HLA based on third-generation sequencing according to claim 1, wherein The quality assessment described in step S2 includes average read length, average data quality, median sequence length, median sequence quality, number of sequences, N50 read length, total number of bases, data quality distribution, and sequence length distribution.
4. The method for automatic analysis and typing of HLA based on third-generation sequencing according to claim 1, wherein The HLA-related reference genomes described in step S3 include HLA-A, HLA-B, HLA-C, HLA-DQA1, HLA-DQB1, HLA-DPA1, HLA-DPB1, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5.
5. The method for automatic analysis and typing of HLA based on third-generation sequencing according to claim 1, wherein The haplotype splitting according to the HLA gene type described in step S5 is specifically as follows: Extract the heterozygous mutations with an average data quality QUAL > 15 for the variations and a single nucleotide variation depth greater than 25% of the average depth. Extract the homozygous mutations with an average data quality QUAL > 15 for the variations and an SNV frequency greater than 90%. If the number of heterozygous mutations exceeds 25% of the total number of mutations, extract the bases at the corresponding heterozygous mutation positions in the read sequences to form a new mutant combination sequence; if the number of heterozygous mutations is less than 25%, randomly select the bases at the homozygous mutation positions to form a new mutant combination sequence until its length exceeds 25% of the total number of mutations. If there are no heterozygous SNVs, skip the analysis in step S7 and directly proceed to step S10.
6. The method for automatic analysis and typing of HLA based on third-generation sequencing according to claim 1, wherein Step S8 is specifically as follows: For the sequences in the first two clustering groups of each amplicon region in step S7, select the sequences with a length ranging from amplicon - 1000 bp to + 1000 bp, downsample to a 500X coverage, and use the flye and canu software for genome assembly. Select the minimum value of the amplicon length - 1000 bp or 10000 bp as the minimum assembly overlap region for assembly.
7. The method for automatic analysis and typing of HLA based on third-generation sequencing according to claim 1, wherein The quality inspection of the typing results described in step S11 includes assessment of amplification imbalance, assessment of linkage imbalance, homozygous typing heterozygous analysis, serological conversion, discrimination of rare typings, identification of new typings, and visual alignment.
Citation Information
Patent Citations
HLA genotype-SNP linkage database, its constructing method, and HLA typing method
CN103221551A
A method for HLA genotyping based on a third-generation sequencing platform
CN108460246B
Methods of human leukocyte antigen typing
CN109477143A
Analysis method and analysis and processing device for loss of heterozygosity in human HLA chromosomal region
WO2022116456A1
Three-generation sequencing platform-based HLA genotyping method
CN108460246A
Cited By
Weighted syncmer-based extensive genome length and read length comparison method
CN121528309A