A comprehensive analysis method of HLA genes based on high-throughput sequencing data

By employing a multi-algorithm parallel typing and result integration approach, combining the iNeo-HLA-Qual and iNeo-HLA-Quant algorithms, the problem of insufficient qualitative and quantitative accuracy in HLA gene analysis of high-throughput sequencing data is solved, achieving higher typing accuracy and expression level calculation, and is suitable for comprehensive HLA gene analysis.

CN114678071BActive Publication Date: 2026-05-19HANGZHOU XINYUANLI BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU XINYUANLI BIOTECHNOLOGY CO LTD
Filing Date
2021-12-31
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing methods for HLA gene analysis of high-throughput sequencing data lack accuracy and completeness in both qualitative and quantitative aspects, especially in the lack of downstream analysis support that simultaneously provides qualitative and quantitative results.

Method used

A multi-algorithm parallel typing and result integration approach was adopted, combining the iNeo-HLA-Qual algorithm for qualitative analysis and the iNeo-HLA-Quant algorithm for quantitative analysis. Preprocessing and comprehensive reference genome (CRP) were used to improve the accuracy of data alignment, and the accuracy of typing results and expression level calculation were improved by weighting coefficients and read count corrections.

Benefits of technology

It improves the accuracy of HLA genotyping and expression level calculation, provides more comprehensive data support, is significantly correlated with HLA allele copy number, and is suitable for HLA gene analysis in clinical and research settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114678071B_ABST
    Figure CN114678071B_ABST
Patent Text Reader

Abstract

The application discloses a kind of HLA gene comprehensive analysis methods based on high-throughput sequencing data, comprising the following contents: step one, data preprocessing: DNA sequencing data is compared to reference genome, then according to the alignment result, the sequence needed is selected to form file, to be used as the input of downstream analysis;Step two, HLA typing scheme based on NGS sequencing alignment result iNeo-HLA-Qual, HLA typing scheme based on the performance of multiple algorithms in actual data set is weighted integrated;Step three, according to the read count calculation HLA typing expression amount scheme iNeo-HLA-Quant;The application carries out the performance of multiple algorithms in actual data set and is weighted integrated, the application not only simultaneously gives the downstream analysis of qualitative, quantitative two kinds of results, and various algorithms are integrated to realize synergistic effect in accuracy;The accuracy of analysis and the perfection of data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics, and in particular to a method for HLA gene synthesis analysis based on high-throughput sequencing data. Background Technology

[0002] Human leukocyte antigen (HLA) is a protein molecule in cells used to present antigenic peptides, and the genes encoding HLA are called HLA genes. In practical analysis, two main classes of classic HLA genes are typically studied: HLA-I and HLA-II genes. Both types of genes are located on chromosome 6 and are named according to their location. In practical applications, HLA-I genes primarily focus on three types: HLA-A, HLA-B, and HLA-C. Their encoded products are distributed on the surface of almost all nucleated cells (except nerve cells, mature erythrocytes, and trophoblast cells), and they mainly participate in the presentation of endogenous antigens. HLA-II genes typically focus on HLA-DR, HLA-DQ, and HLA-DP genes. These are mostly distributed on the surface of professional antigen-presenting cells such as T lymphocytes, macrophages, and dendritic cells, and participate in the presentation of exogenous antigens.

[0003] Because HLA functions by binding antigens and participating in antigen presentation, the high diversity of antigens determines the high diversity of HLA itself and HLA genes; that is, each type of HLA gene has a large number of genotypes. To facilitate research and application, the Nomenclature Committee for Factors of the HLA System under the World Health Organization has established detailed rules for the nomenclature of HLA genotypes. See the appendix for specific nomenclature rules. Figure 5 .

[0004] In biology, a cell possessing multiple sets of genetic material is called polyploid, and genes at identical locations on the two sets of genetic material are called alleles. Humans are diploid organisms, inheriting one set of genetic material from the mother and the other from the father. If a pair of alleles of a gene have identical genotypes, we call that gene homozygous; otherwise, it is heterozygous. For ease of description, we will refer to the two alleles of an HLA gene as allele 1 (allele1) and allele 2 (allele2) in the following text.

[0005] HLA genotyping is the process and method of determining the genotype of HLA alleles in a subject using technical means, while also understanding the homozygous / heterozygous status of each HLA gene. There are various methods for HLA genotyping, such as serological typing, sequence-specific oligonucleotide hybridization (SSO), capillary sequencing (Sanger sequencing), and next-generation sequencing (NGS). Different methods vary in experimental efficiency, sample requirements, and genotyping resolution. Thanks to the rapid development of NGS technology in recent years, which offers higher experimental efficiency, higher resolution, and lower sample requirements, NGS is being used more and more widely.

[0006] For high-throughput sequencing data, the genotyping algorithm largely determines the accuracy and resolution of the genotyping results. While different algorithms share similar technical principles—aligning the sequenced reads (short base sequences obtained after randomly breaking down nucleic acid sequences during sequencing) with a prepared reference genome to roughly identify reads belonging to HLA genes—they then assign these aligned reads to HLA genotypes based on their respective algorithms. Finally, the genotype of each HLA allele in the subject is determined based on the assignment results, thus identifying heterozygotes and completing the genotyping. However, different algorithms employ different strategies in key steps such as reference genome usage, assignment algorithm construction, and heterozygote determination. Therefore, the final genotyping results often differ between methods, and the accuracy of the same method can also vary on datasets with different characteristics.

[0007] Furthermore, in past clinical and research applications, the analysis of HLA genes has been mostly limited to HLA typing and qualitative analysis, with very few studies and applications targeting HLA gene expression levels. Therefore, there is currently a large gap in quantitative methods for HLA genes. This invention addresses these issues and provides an analytical method that can obtain more accurate and comprehensive data support for downstream analyses that require both qualitative and quantitative results. Summary of the Invention

[0008] To address the shortcomings of existing technologies, the present invention aims to provide a comprehensive HLA gene analysis method based on high-throughput sequencing data. This method determines the HLA gene type of a subject based on DNA sequencing data and calculates the expression level of the type based on RNA sequencing results, thereby improving the accuracy and completeness of data for downstream analysis that requires both qualitative and quantitative results.

[0009] To achieve the above objectives, the present invention adopts the following technical solution:

[0010] A comprehensive HLA gene analysis method based on high-throughput sequencing data includes the following:

[0011] Step 1, Data Preprocessing:

[0012] DNA sequencing data is aligned to a reference genome, and then the required partial sequences are selected based on the alignment results to form a file, which serves as input for downstream analysis.

[0013] Step 2: Parallel multi-algorithm fracturing and result integration:

[0014] 1. Summary of results from multi-algorithm integration:

[0015] a) Summarize the HLA typing results from multiple algorithms, standardize all typing results, and then perform deduplication.

[0016] b. Calculate the weight value W for all the summarized subtypes:

[0017] The weight values ​​are obtained by summing the weight coefficients w of each algorithm that supports the classification. These weight coefficients are based on the performance of the algorithm on a specific dataset. The metrics for measuring this performance are the classification accuracy of the algorithm or the correlation between the algorithm results and the gold standard results.

[0018] c, Identify allele 1:

[0019] After determining the weight values ​​of all fractals, arrange all fractals in descending order of weight value;

[0020] If there are ties for the top-ranked fractal, then all the ties are directly taken as the heterozygous result;

[0021] If there are no ties for the first place, the genotype ranked first is allele 1, and its corresponding weight value is denoted as W1.

[0022] d. Determine if allele 2 is homozygous:

[0023] After allele 1 is determined, if other genotypes are found, the second largest genotype is selected according to the size of W, with its corresponding weight value being W2. W2 is then compared with the pre-set weight threshold T. thr Compare;

[0024] If W2>T thr If the result is heterozygous, the corresponding genotypes are allele 1 and allele 2.

[0025] If W2≤T thr If the typing result is homozygous, the corresponding typing only has allele 1.

[0026] Once allele 1 is identified, if no other results are found, the genotyping result is homozygous, and the corresponding genotype contains only allele 1.

[0027] If a second-ranked fractal exists and satisfies the heterozygous threshold, then all the second-ranked fractals are retained.

[0028] e. The intersection of the typing results of normal tissue samples and tumor tissue samples is taken;

[0029] 2. Qualitative analysis of the iNeo-HLA-Qual algorithm:

[0030] a, Fractal read count:

[0031] Obtain the alignment results of the reference genome and DNA sequencing data, and count the number N of reads aligned to the corresponding HLA type based on the alignment results;

[0032] b. Select candidate subtypes:

[0033] Select the largest number of supported reads from all the fractals and count it as N. max Multiply this by a set coefficient C (a decimal between 0.5 and 1) to obtain the filtering threshold N. thr Only keep N>N thr The subtypes were used as candidates for further analysis;

[0034] c, allele combination set:

[0035] Candidate HLA types are sequentially designated as allele 1, and other types constitute allele 2 sets, thus obtaining the combination set of [allele 1 - allele 2 sets] corresponding to all HLA types;

[0036] d. Extract the set of allele 1-allele 2:

[0037] Extract an allele 1-allele 2 set from the combination set of [allele 1-allele 2 set] to obtain allele 1;

[0038] e, Fractal read recount:

[0039] After identifying allele 1, the read that simultaneously aligned with allele 1 among the remaining candidate genotypes is removed from the calculation results of the remaining candidate genotypes, resulting in the recounted read count N'. For a genotype HLA-A*02:01, its supporting read count is N'. A0201 ;

[0040] After all N' calculations are completed, the remaining subtypes are screened again in the manner described in step 2b of selecting candidate subtypes. The subtypes that pass the screening are used as the new allele 2 set for subsequent analysis.

[0041] f, allele 2 is determined and generated by allele combination:

[0042] Based on the size of N', the candidate genotypes obtained in step 2e are selected as alleles 2 and grouped together with alleles 1 determined in step 2d to form an allele combination. All combinations are stored in set C. allele middle;

[0043] g, repeatedly generates combinations:

[0044] Returning to step 2d, extract the allele 1-allele 2 set, extract a new allele 1-allele 2 set, and then repeat steps 2e-2f. The determined combination is stored in C. allele In the process, until all allele 1-allele 2 sets in the combination set of [allele 1-allele 2 set] have been analyzed;

[0045] h, calculate the combination score:

[0046] For each belonging to C allele Fractal combination C allele1-allele2 N allele1 With N' allele2 The sum is the fraction S of the combination, denoted as S0. allele1-allele2 ;

[0047] i. The typing result is determined:

[0048] Select the S with the highest score from all S's. allele1-allele2 As an alternative result, based on the determined heterozygosity threshold T heter To determine the final result;

[0049] If N' allele2 >T heter ×N allele1 If the result is heterozygous, the genotypes will be allele 1 and allele 2, respectively.

[0050] If N' allele2 ≤T heter ×N allele1 The genotyping result is homozygous, and the genotype is allele 1;

[0051] Step 3: Quantitative analysis of HLA typing expression levels based on read counts using the iNeo-HLA-Quant protocol.

[0052] The aforementioned method for HLA gene synthesis analysis based on high-throughput sequencing data involves a preprocessed synthesis reference genome (CRP) for HLA genotyping constructed from information in a publicly available database in step one. The CRP reference genome includes known HLA allele sequence information from the publicly available database.

[0053] In the aforementioned HLA gene synthesis analysis method based on high-throughput sequencing data, the weight value in step one is obtained by summing the weight coefficient w of each algorithm that supports the genotyping. The weight coefficient is obtained based on the performance of the algorithm in a specific dataset, and the indicator for measuring this performance is the genotyping accuracy of the algorithm or the correlation between the algorithm results and the gold standard results.

[0054] In the aforementioned HLA gene synthesis analysis method based on high-throughput sequencing data, all typing results in step two are standardized to the standard naming format specified by the WHO Nomenclature Committee For Factors of the HLA System.

[0055] The aforementioned method for HLA gene comprehensive analysis based on high-throughput sequencing data, specifically the method for taking the intersection of the typing results of normal tissue samples and tumor tissue samples in step 2.1e, is to retain only the typing results that are consistent between the two sets of data; the process of taking the intersection does not consider the distinction between homozygous and heterozygous, but only considers whether the same typing appears in both tumor tissue and normal control results.

[0056] The aforementioned method for comprehensive HLA gene analysis based on high-throughput sequencing data, specifically the iNeo-HLA-Quant quantitative analysis scheme for calculating HLA genotyping expression levels based on read counts in step three, includes:

[0057] c, Fractal read count:

[0058] Obtain RNA sequencing data, referencing genome alignment results and HLA typing results. Calculate the number N of reads aligned to the HLA typing based on the alignment results.

[0059] d, Fractal read count correction:

[0060] The modifications to N include two aspects:

[0061] First, the count value is corrected to the count value N' of the specific read. Specificity means that a read can only be matched to a specific genotype and cannot be matched to any other genotype.

[0062] Second, N' is standardized for allele length to eliminate read count bias caused by gene length.

[0063] In the aforementioned method for comprehensive HLA gene analysis based on high-throughput sequencing data, step 3b involves standardizing N' for allele length to eliminate read number bias caused by gene length. Specifically, N” = N' × 1000 / L, where L is the reference gene sequence length of a specific HLA allele in the reference genome, and N” is the absolute value of the expression level of a specific HLA allele.

[0064] The advantages of this invention are:

[0065] This invention uses information from publicly available databases to construct a preprocessing comprehensive reference genome (CRP) specifically for HLA genotyping, and performs sequencing data preprocessing based on this CRP. During its construction, known HLA allele sequence information from publicly available databases is added to the genome, thereby improving the overall alignment accuracy of HLA gene sequences.

[0066] This invention integrates the performance of multiple algorithms on real datasets using a weighted average. The iNeo_HLABoost qualitative module achieves higher accuracy than a single algorithm and shows greater stability in HLA-I and HLA-II genotyping results. This demonstrates that this invention helps improve the accuracy of HLA genotyping and achieves synergistic effects in terms of accuracy among various algorithms.

[0067] This approach can determine the HLA genotype of a subject based on DNA sequencing data, and calculate the expression level of the genotype based on RNA sequencing results, providing more accurate and comprehensive data support for downstream analyses that require both qualitative and quantitative results.

[0068] The expression level of HLA alleles was significantly correlated with the copy number of HLA alleles in the genotyping test results. Attached Figure Description

[0069] Figure 1 This is the overall process of the HLA integrated typing and expression level evaluation analysis method of the present invention;

[0070] Figure 2 This is a flowchart of the multi-algorithm classification result integration scheme, i.e., qualitative analysis, of the present invention;

[0071] Figure 3 This is a flowchart of the iNeo-HLA-Qual algorithm, i.e., qualitative analysis, of the present invention;

[0072] Figure 4 This is a flowchart of the iNeo-HLA-Quant (quantitative analysis) scheme for calculating HLA expression levels in this invention;

[0073] Figure 5This invention relates to the naming rules for HLA genotypes under the WHO Nomenclature Committee For Factors of the HLA System.

[0074] Figure 6 This is a correlation diagram between the results of the quantitative module in Experiment 3 of this invention and the results of HLA-LOH analysis.

[0075] Figure 7 This is a correlation diagram between the quantitative module results of this invention and the qPCR experimental analysis results. Detailed Implementation

[0076] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0077] A comprehensive HLA gene synthesis analysis method based on high-throughput sequencing data (iNeo-HLABoost), such as Figure 1 As shown, the steps include:

[0078] I. Data Preprocessing:

[0079] The main task of data preprocessing is to align DNA sequencing data to a reference genome, and then select the necessary partial sequences based on the alignment results to form files as input for downstream analysis.

[0080] As a preferred option, the reference genome used is a preprocessed comprehensive reference panel (CRP) for HLA genotyping, constructed from information in a publicly available database that has undergone special processing for HLA genotyping. During its construction, known HLA allele sequence information from publicly available databases is added to the genome, thereby improving the overall alignment accuracy of HLA gene sequences.

[0081] II. Parallel Fractalization and Result Integration Using Multiple Algorithms:

[0082] 2.1 Multi-algorithm classification result integration scheme

[0083] To reduce the bias of the HLA typing algorithm and improve the overall accuracy of the typing results, this paper designs the following weighted integration scheme for the typing results, as follows: Figure 2 As shown:

[0084] 2.1.1. Summary of results from multiple algorithms:

[0085] First, obtain the HLA typing results from different algorithms, then standardize all typing results to the standard naming format specified by the WHO Nomenclature Committee For Factors of the HLA System (e.g., HLA-A*06:02), and then perform deduplication.

[0086] Assuming the two algorithms give the classification results HLA-A*02:06, HLA-A*06:03 and HLA-A*04:08, HLA-A*06:03 respectively, then only the unique HLA-A*02:06, HLA-A*06:03 and HLA-A*04:08 will be retained in the summary results.

[0087] 2.1.2. Weight Calculation:

[0088] The weights W are calculated for all the subtypes summarized in Step 1. The weights are obtained by summing the weight coefficients w of each algorithm supporting that subtype, which are based on the algorithm's performance on a specific dataset. Metrics for measuring this performance include the algorithm's subtype accuracy or the correlation between the algorithm's results and the gold standard results. For example, if HLA-A*02:04 is supported by three algorithms A, B, and C in an analysis, then the weight W for that subtype would be... A0204 =w A +w B +w C .

[0089] 2.1.3. Identification of Allele 1:

[0090] After determining the weight values ​​for all genotypes, all genotypes are arranged in descending order of their weight values. If there are ties for the top-ranked genotype, all genotypes with that ties are directly considered heterozygous. If there are no ties for the top-ranked genotype, then the top-ranked genotype is allele 1, and its corresponding weight value is denoted as W1.

[0091] 2.1.4. Determination of allele 2 and homozygosity:

[0092] After allele 1 is determined, if other genotypes are found, the second largest genotype is selected according to the size of W, with its corresponding weight value being W2. W2 is then compared with the pre-set weight threshold T. thr Compare. If W2 > T thr If allele 1 is found to be heterozygous, the genotype is determined to be heterozygous, corresponding to allele 1 and allele 2. Otherwise, the genotype is determined to be homozygous, corresponding to allele 1 only. If no other results are found after determining allele 1, the genotype is determined to be homozygous, corresponding to allele 1 only. If a genotype with a tie for second place exists and meets the heterozygous threshold, both genotypes with the tie are retained.

[0093] For example, in a sample, the genotype with the largest weight value is HLA-A*02:02, while the weight values ​​W for HLA-A*02:01 and HLA-A*02:06 are... A0201 W A0206 Both are 1.23, tied for second place, higher than the set T. thr The final typing results are HLA-A*02:02, HLA-A*02:01, and HLA-A*02:06.

[0094] 2.1.5. Intersection of normal and tumor tissues:

[0095] When both normal tissue data and tumor data are available for the subject, the typing results are determined for each type of data according to steps 1 to 4. Then, the intersection of the two results is processed, retaining only the typing results that are consistent between the two sets of data. The intersection process does not consider the distinction between homozygous and heterozygous, but only whether the same typing appears in both tumor tissue and normal control results.

[0096] For example, if the tumor tissue typing result is homozygous HLA-A*02:06, while the normal tissue results are heterozygous HLA-A*02:06 and HLA-A*02:01, then the result after taking the intersection is HLA-A*02:06. Alternatively, if there are no common typing results after taking the intersection, the typing result is recorded as missing.

[0097] 2.2 iNeo-HLA-Qual Algorithm (Qualitative Analysis)

[0098] During the genotyping process, we may incorporate a self-developed genotyping algorithm based on NGS sequencing data, called iNeo-HLA-Qual. Therefore, a detailed description of this algorithm is provided, such as... Figure 3 As shown. The following calculation method is for determining the result within a single Locus. For different Locuses, the classification result is determined separately using the following method.

[0099] 2.2.1. Fractal read count:

[0100] Obtain the reference genome alignment results and count the number N of reads aligned to different HLA types based on the alignment results. For a type HLA-A*02:01, the number of supporting reads is counted as N. A0201 .

[0101] 2.2.2. Candidate subtyping selection:

[0102] Select the largest number of supported reads from all the fractals and count it as N. max Multiply this by a set coefficient C (a decimal between 0.5 and 1) to obtain the filtering threshold N.thr Only retain N>N thr The subtypes were used as candidates for further analysis.

[0103] 2.2.3 Allele combination set:

[0104] Candidate HLA types are sequentially designated as allele 1, and other types constitute allele 2 sets, thus obtaining the combination set of [allele 1 - allele 2 sets] corresponding to all HLA types.

[0105] 2.2.4. Extracting the set of allele 1-allele 2:

[0106] Extract a set of alleles 1 and alleles 2 from the set of allele 1-allele 2 combinations. Obtain allele 1.

[0107] 2.2.5. Fractal read recount:

[0108] After allele 1 is identified, the number of reads for the candidate genotypes must be recalculated. Specifically, the read that simultaneously aligns to the remaining candidate genotypes and is associated with allele 1 is removed from the calculation results of the remaining candidate genotypes, resulting in the recounted read count N'. For a genotype HLA-A*02:01, its supporting read count is N'. A0201 .

[0109] For example, if allele 1 is identified as HLA-A*02:02, its read count N A0202 The number of reads is 1024, and the remaining candidate genotypes are HLA-A*02:01 and HLA-A*02:03, with N reads. A0201 N A0203 The values ​​are 512 and 256 respectively. N A0201 128 reads can be simultaneously matched to HLA-A*02:02, while N A0203 If 64 reads can be simultaneously matched to HLA-A*02:02, then N' A0201 =512-128=384, N' A0203 =256-64=192. After all N' calculations are completed, the remaining subtypes are screened again in the manner described in step 2.2.2. The subtypes that pass the screening are used as the new allele 2 set for subsequent analysis.

[0110] 2.2.6. Determination of Allele 2 and Generation of Allele Combinations:

[0111] Candidate genotypes obtained in section 2.2.5 are selected as allele 2 in descending order of N' size, and combined with allele 1 determined in section 2.2.4 to form an allele combination. All combinations are stored in set C. allele middle.

[0112] 2.2.6. Repeatedly generating combinations:

[0113] Returning to step 2.2.4, extract the new set of alleles 1-2, then repeat steps 2.2.5-2.2.6, storing the determined combinations in C. allele In the process, until all allele 1-allele 2 sets in the combined set of allele 1-allele 2 are analyzed.

[0114] 2.2.7. Calculate the combination fraction:

[0115] For each belonging to C allele Fractal combination C allele1-allele2 N allele1 With N' allele2 The sum is the fraction S of the combination, denoted as S0. allele1-allele2 .

[0116] 2.2.8. Determination of typing results:

[0117] Select the S with the highest score from all S's. allele1-allele2 As an alternative result, based on the determined heterozygosity threshold T heter To determine the final result.

[0118] If N' allele2 >T heter ×N allele1 If the result is heterozygous, the genotypes will be allele 1 and allele 2, respectively.

[0119] Conversely, if the typing result is homozygous, the typing is allele 1.

[0120] III. HLA Expression Calculation Scheme iNeo-HLA-Quant (Quantitative Analysis):

[0121] HLA typing and expression levels are significant for neoantigen therapy selection because high typing expression levels indicate high expression levels of the corresponding HLA molecules on the cell surface. Antigens with good affinity for these molecules are more likely to be presented to T cells, thus potentially leading to better immunogenicity. Therefore, we designed a read-count-based scheme to evaluate the expression levels of different HLA typings, providing a basis for downstream neoantigen screening. Details of our designed expression quantification scheme, iNeo-HLA-Quant, are as follows: Figure 4 As shown:

[0122] 3.1. Fractal read count:

[0123] Obtain RNA sequencing data, referencing genome alignment results and HLA typing results. Calculate the number N of reads that are aligned to the HLA typing based on the alignment results.

[0124] 3.2. Fractal read count correction:

[0125] The modifications to N include two aspects:

[0126] First, the count value is corrected to the count value N' of the specific read. Specificity means that a read can only be matched to a specific genotype and cannot be matched to any other genotype.

[0127] Second, N' is standardized for allele length to eliminate read number bias caused by gene length. Specifically, N” = N' × 1000 / L, where L is the reference gene sequence length of a specific HLA allele in the reference genome, and N” is the absolute value of the expression level of a specific HLA allele.

[0128] The technical effects of the present invention are verified through experiments as follows:

[0129] Experiment 1: Application of the qualitative module of iNeo-HLABoost in clinical subjects

[0130] I. Experimental Materials

[0131] The blood samples were taken from nine healthy individuals, and the gold standard for typing was the results of first-generation HLA sequencing.

[0132] II. Experimental Procedure

[0133] Nine collected samples were subjected to HLA typing using PCR-SSO, and simultaneously, the iNeo-HLA-Boost algorithm described in this invention and other NGS-based HLA typing algorithms were used for further typing. The PCR-SSO HLA typing results were used as the gold standard for accuracy, and the accuracy of this invention's method was compared with that of other algorithms.

[0134] Nine collected samples were subjected to HLA genotyping using PCR-SSO. Simultaneously, the iNeo-HLA-Boost qualitative analysis module described in this invention, along with iNeo-Qual and other NGS-based HLA genotyping algorithms, were used for further genotyping. The PCR-SSO HLA genotyping results were used as the gold standard for accuracy, comparing the accuracy of this invention's method with other algorithms.

[0135] PCR-SSO is an HLA genotyping method based on specific DNA probe hybridization experiments. The method first requires DNA extraction from the sample to be genotyped. After obtaining the DNA, specific PCR primers are used to amplify exons 2 and 3 of the HLA-I gene and exon 2 of the HLA-II gene. After the amplification is completed, the sample is added to a 96-well hybridization plate, and pre-designed specific probes for different HLA genotypes are added for hybridization. Finally, the HLA genotype is determined based on the fluorescence generated by the hybridization reaction.

[0136] The specific process of iNeo-HLA-Qual described in this invention is described in Step 1 and Step 2 of the invention description.

[0137] Other HGS-based HLA typing schemes and their specific parameters used in this experiment are shown in Table 1 below:

[0138] Table 1

[0139]

[0140] After obtaining the genotyping results and the integrated results of the iNeo-HLA-Boost qualitative analysis module, the differences between the genotyping results of the iNeo-HLA-Boost qualitative analysis module, iNeo-Qual, HLA-PRG, xHLA, Phalt, and Optitype and the HLA genotyping results obtained by PCR-SSO were compared, and the genotyping accuracy was calculated.

[0141] An example of how accuracy is calculated is as follows:

[0142] Assume the genotyping results from the iNeo-HLABoost qualitative module and the genotyping results from the PCR-SSO module are shown in Table 2 below:

[0143] Table 2

[0144]

[0145]

[0146] It can be seen that the two methods have inconsistent genotyping results for one allele each in the HLA-A and HLA-B genes. Therefore, for this sample, the genotyping accuracy is (6-2) / 6 = 66.7%.

[0147] The experimental results for all nine samples are shown in Tables 3 and 4:

[0148] Table 3 Comparison of HLA-I genotyping results from different algorithms for nine samples.

[0149]

[0150] Table 4. Comparison of HLA-II genotyping results using different algorithms in nine samples.

[0151]

[0152] III. Results Analysis

[0153] The results in Tables 3 and 4 show that, using PCR-SSO results as the gold standard, the accuracy of the iNeo_HLABoost qualitative module is no lower than that of a single algorithm, and it also demonstrates relatively stable performance in HLA-I and HLA-II gene genotyping results. This proves that the present invention helps improve the accuracy of HLA genotyping, playing a complementary and integrated role.

[0154] Experiment 2: Application of the qualitative module of iNeo-HLABoost on the HapMap dataset

[0155] We obtained the East Asian population dataset publicly available from the HapMap project, and obtained the HLA typing results of this dataset obtained by PCR-SSO method from the research of X et al. We performed HLA typing analysis on the dataset according to the method described in Experiment 1, and compared the accuracy of the iNeo-HLABoost qualitative module with other algorithms.

[0156] The experimental results are shown in Tables 5 and 6:

[0157] Table 5

[0158]

[0159]

[0160] Table 6

[0161]

[0162] The results in Tables 5 and 6 show that using this scheme can make the overall integrated classification results exceed the classification results of any single scheme.

[0163] Experiment 3: Application of the iNeo-HLABoost quantification module in clinical subject samples:

[0164] I. Detection of Expression Levels in Typing Results

[0165] The RNA-Seq data of the samples were analyzed using the iNeo-HLABoost quantitative analysis module (iNeo-Quant) to analyze the data of 20 known HLA typing results. After standardization based on the comparison of HLA typing, the expression level results corresponding to each HLA typing were obtained.

[0166] The expression levels obtained were then compared with the HLA genotype-specific copy number results obtained by the LOHHLA (Loss Of Heterozygosity in Human Leukocyte Antigen) analysis method mentioned in the study by Anagnostou, V. et al., *Multimodal genomic features predict outcome of immune checkpoint blockade in non-small-cell lung cancer*, *Nature Cancer*, 1, 99–111 (2020). LOHHLA analyzes HLA gene deletion based on copy number variations. Theoretically, HLA gene deletion leads to a significant decrease in the mRNA level of that HLA genotype. Therefore, the correlation analysis between the LOHHLA results and the HLA expression levels of this invention can prove the correctness of the analytical results of this invention.

[0167] In addition, 10 samples with known HLA typing and RNA samples were selected. After designing qPCR probes targeting specific parts of their HLA typing sequences, qPCR experiments were performed on the RNA samples to obtain the expression levels of the corresponding HLA typing. Correlation analysis was then performed between these expression levels and the expression levels obtained from the iNeo-HLABoost quantitative analysis module (iNeo-Quant).

[0168] II. Experimental results are as follows Figure 6 Correlation plot between quantitative module results and HLA-LOH analysis results, and Figure 7 The correlation between the quantitative module results and the qPCR experimental analysis results is shown in the following figure:

[0169] Figure 6 The x-axis represents the HLA genotyping expression level obtained through the iNeo-HLABoost quantitative analysis module (iNeo-Quant), and the y-axis represents the HLA genotyping-specific copy number obtained through the LOHHLA analysis method.

[0170] Figure 7 The x-axis represents the HLA genotyping expression level obtained through the iNeo-HLABoost quantitative analysis module (iNeo-Quant), and the y-axis represents the HLA genotyping expression level obtained through qPCR experiments.

[0171] Results Analysis: Figure 6 It can be seen that the LOHHLA results are highly correlated with the HLA expression levels of this invention (P < 0.001), indicating that our calculation results are significantly correlated with the copy number of HLA alleles. Figure 7 It can be seen that the qPCR experimental results are highly correlated with the HLA expression level of the present invention, with a P value < 0.05, indicating that our calculation results are significantly correlated with the copy number of HLA alleles.

[0172] In summary, the three experiments demonstrate that the analytical method of this invention improves the accuracy and completeness of data for downstream analyses that require both qualitative and quantitative results.

[0173] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any way, and all technical solutions obtained by equivalent substitution or equivalent transformation fall within the protection scope of the present invention.

Claims

1. A method for HLA gene synthesis analysis based on high-throughput sequencing data, characterized in that, Includes the following: Step 1, Data Preprocessing: DNA sequencing data is aligned to a reference genome, and then the required partial sequences are selected based on the alignment results to form a file, which serves as input for downstream analysis. Step 2: Parallel multi-algorithm fracturing and result integration:

1. Summary of results from multi-algorithm integration: a) Summarize the HLA typing results from multiple algorithms, standardize all typing results, and then perform deduplication. b. Calculate the weight value W for all the summarized subtypes: The weight values ​​are obtained by summing the weight coefficients w of each algorithm that supports the classification. These weight coefficients are based on the performance of the algorithm on a specific dataset. The metrics for measuring this performance are the classification accuracy of the algorithm or the correlation between the algorithm results and the gold standard results. c, Identify allele 1: After determining the weight values ​​of all fractals, arrange all fractals in descending order of weight value; If there are ties for the top-ranked fractal, then all the ties are directly taken as the hybrid result; If there are no ties for the first place, the genotype ranked first is allele 1, and its corresponding weight value is denoted as W1. d. Determine if allele 2 is homozygous: After determining allele 1, if there are other genotype results, the second largest genotype is selected according to the size of W, and its corresponding weight value is W2. W2 is then compared with the set weight threshold Tthr. If W2>Tthr, the genotyping result is heterozygous, corresponding to allele 1 and allele 2; If W2≤Tthr, the genotyping result is homozygous, and the corresponding genotype has only allele 1; Once allele 1 is identified, if no other results are found, the genotyping result is homozygous, and the corresponding genotype contains only allele 1. If a second-ranked fractal exists and satisfies the heterozygous threshold, then all the second-ranked fractals are retained. e. The intersection of the typing results of normal tissue samples and tumor tissue samples is taken; 2. Qualitative analysis of the iNeo-HLA-Qual algorithm: a, Fractal read count: Obtain the alignment results of the reference genome and DNA sequencing data, and count the number N of reads aligned to the corresponding HLA type based on the alignment results; b. Select candidate subtypes: The largest number of supporting reads among all morphologies is selected as Nmax. This value is then multiplied by a set coefficient C (a decimal between 0.5 and 1) to obtain the filtering threshold Nthr for N. Only morphologies with N>Nthr are retained as candidates for further analysis. c, allele combination set: Candidate HLA types are sequentially designated as allele 1, and other types constitute allele 2 sets, thus obtaining the combination set of [allele 1 - allele 2 sets] corresponding to all HLA types; d. Extract the set of allele 1-allele 2: Extract an allele 1-allele 2 set from the combination set of [allele 1-allele 2 set] to obtain allele 1; e, Fractal read recount: After allele 1 is determined, the read that is matched with allele 1 and is also matched with the remaining candidate genotypes is removed from the calculation results of the remaining candidate genotypes, and the number of reads after recounting is obtained N'. For a genotype HLA-A*02:01, its supporting reads are counted as N'A0201. After all N' calculations are completed, the remaining subtypes are screened again in the manner described in step 2b of selecting candidate subtypes. The subtypes that pass the screening are used as the new allele 2 set for subsequent analysis. f, allele 2 is determined and generated by allele combination: According to the size of N', the candidate genotypes obtained in step 2e are taken as alleles 2 and placed together with alleles 1 determined in step 2d to form an allele combination. All combinations are stored in the set Callele. g, repeatedly generates combinations: Return to step 2d to extract the allele 1-allele 2 set, extract the new allele 1-allele 2 set, and then repeat steps 2e-2f. Store the determined combinations in Callele until all allele 1-allele 2 sets in the [allele 1-allele 2 set] combination set have been analyzed. h, calculate the combination score: For each combination of Callele belonging to Callele, Callele1-allele2, add Nallele1 and N'allele2 to get the fraction S of the combination, which is called Salule1-allele2; i. The typing result is determined: Select the largest score, Salule1-allele2, from all S as the candidate results, and determine the final result based on the determined heterozygosity threshold Theter. If N'allele2>Theter×Nallele1, then the final genotyping result is heterozygous, with allele 1 and allele 2 as the genotypes respectively; If N'allele2≤Theter×Nallele1, then the genotyping result is homozygous, and the genotype is allele 1; Step 3: Quantitative analysis of HLA typing expression levels based on read counts using iNeo-HLA-Quant. The weight values ​​in step one are obtained by summing the weight coefficients w of each algorithm that supports the classification. The weight coefficients are obtained based on the performance of the algorithm in a specific dataset, and the indicators for measuring this performance are the classification accuracy of the algorithm or the correlation between the algorithm results and the gold standard results. The specific method for taking the intersection of the typing results of normal tissue samples and tumor tissue samples in step 2.1e is to retain only the typing results that are consistent between the two sets of data; the process of taking the intersection does not consider the distinction between homozygous and heterozygous, but only considers whether the same typing appears in both tumor tissue and normal control results; The specific content of the iNeo-HLA-Quant quantitative analysis scheme for calculating HLA typing expression levels based on read counts in step three includes: a, Fractal read count: Obtain RNA sequencing data, referencing genome alignment results and HLA typing results, and count the number of reads N aligned to the HLA typing based on the alignment results; b, Fractal read count correction: The modifications to N include two aspects: First, the count value is corrected to the count value N' of the specific read. Specificity means that a read can only be matched to a specific genotype and cannot be matched to any other genotype. Second, N' is standardized for allele length to eliminate read count bias caused by gene length. In step 3b, the specific method for standardizing N' to eliminate read number bias caused by gene length is to set N” = N' × 1000 / L, where L is the reference gene sequence length of a specific HLA allele in the reference genome, and N” is the absolute value of the expression level of a specific HLA allele.

2. The method for HLA gene synthesis analysis based on high-throughput sequencing data according to claim 1, characterized in that, In step one, the reference genome is a preprocessed comprehensive reference genome for HLA typing constructed from information in a public database, referred to as the CRP reference genome. The CRP reference genome adds known HLA allele sequence information from the public database.

3. The method for HLA gene synthesis analysis based on high-throughput sequencing data according to claim 1, characterized in that, All typing results in step two are standardized to the standard naming format specified by the WHO Nomenclature Committee For Factors of the HLA System.