HLA genotyping method, device, equipment and medium

By constructing a polymorphism database and generating a single nucleotide diversity characteristic matrix, combining the HLA genotype data of the sample to be tested for error comparison of reconstruction and heterozygous combination, the shortcomings in efficiency and accuracy of the existing HLA genotyping methods are solved, and efficient and accurate HLA genotyping is achieved.

CN120148658AActive Publication Date: 2025-06-13SUZHOU LASSO BIOCHIP TECH CO LTD +1

Patent Information

Application Number
CN202510629351.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The existing HLA genotyping methods have insufficient efficiency and accuracy, especially due to the high polymorphism of HLA genes, the SNP differences between different types are small, making it difficult to efficiently and accurately typing through traditional methods.

Method used

By constructing a polymorphism database, a single nucleotide diversity feature matrix is ​​generated, the HLA genotype data of the sample to be tested is obtained for reconstruction, all possible heterozygous combinations are generated, and their single nucleotide diversity features are merged to form an expected set of features. Then, the feature vectors of the sample to be tested are compared bit by bit with the expected feature set, mismatch records are counted, and the heterozygous combination with the least number of mismatches is determined as the typing result.

Benefits of technology

This method can accurately determine the specific type of HLA gene in the sample, improve detection throughput, reduce costs, and also have a high typing accuracy. It is suitable for application scenarios such as clinical transplantation, immunology research and individualized treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148658A_ABST
    Figure CN120148658A_ABST
Patent Text Reader

Abstract

The invention relates to an HLA (human leukocyte antigen) genotyping method, device, equipment and medium. The method comprises the following steps: constructing a polymorphic database, and generating an SNP feature matrix of the polymorphic database; obtaining HLA genotype data of a to-be-detected sample, and reconstructing the HLA genotype data of the to-be-detected sample to obtain a feature vector of the to-be-detected sample; two alleles of the same gene locus are obtained from the polymorphism database, all heterozygous combinations of the two alleles are generated, mononucleotide diversity characteristics in all the heterozygous combinations are combined, and an expected characteristic set of all the heterozygous combinations is obtained; comparing the feature vector of the to-be-tested sample with the expected feature set bit by bit, counting mismatch records of data in the feature vector of the to-be-tested sample and the expected feature set, and recording the mismatch number of each heterozygous combination; and determining the heterozygous combination with the minimum mismatch number as an HLA typing result. By adopting the method, genotyping can be efficiently and accurately carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of bioengineering technology, and in particular, to an HLA gene typing method, device, equipment and medium. Background Art

[0002] The HLA (human leukocyte antigen) gene is of great significance in clinical detection due to its high polymorphism and close correlation with various diseases. Currently, there are various HLA typing methods, such as specific primer amplification technology (PCR-SSP), Sanger sequencing typing method (PCR-SBT), and high-throughput sequencing technology (NGS), etc.

[0003] As a high-throughput detection platform, a gene chip can detect SNP (single nucleotide polymorphism) information in the target region by using pre-designed probes on the chip, and can quickly obtain the complete sequence information of the exon region of the sample, thus providing a new idea for HLA typing. However, due to the extremely high polymorphism of the HLA gene and only minor SNP differences between different types, the current HLA gene typing methods still have deficiencies in terms of efficiency and accuracy. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide an efficient and accurate HLA gene typing method, and the method includes: Construct a polymorphism database, and generate a single nucleotide polymorphism feature matrix of the polymorphism database; obtain the HLA genotype data of the sample to be tested, reconstruct the HLA genotype data of the sample to be tested to obtain the feature vector of the sample to be tested; obtain two alleles at the same gene locus from the polymorphism database, generate all heterozygous combinations of the two alleles, merge the single nucleotide polymorphism features in each of the heterozygous combinations to obtain the expected feature set of each heterozygous combination; compare the feature vector of the sample to be tested with the expected feature set bit by bit, count the mismatch records of the data in the feature vector of the sample to be tested and the expected feature set, and record the number of mismatches of each heterozygous combination; determine the heterozygous combination with the least number of mismatches as the HLA typing result.

[0005] In one embodiment, constructing the polymorphism database and generating the single nucleotide diversity feature matrix of the polymorphism database includes: determining a reference HLA genotype and its corresponding reference sequence; obtaining single nucleotide diversity difference sites between the exon sequences of the remaining HLA genotypes in the polymorphism database except the reference HLA genotype and the reference sequence; and generating the single nucleotide diversity feature matrix of the polymorphism database based on the single nucleotide diversity difference sites of each HLA genotype.

[0006] In one embodiment, reconstructing the HLA genotype data of the test sample to obtain the feature vector of the test sample includes: sorting based on the exon region coordinates of the HLA genotype data of the test sample and splicing to obtain a complete sequence; and representing the heterozygous sites in the complete sequence with IUPAC degenerate codes to obtain the feature vector of the test sample.

[0007] In one embodiment, obtaining the HLA genotype data of the test sample includes: for the key HLA typing sites of the test sample, setting two different probes for detection, and if a single nucleotide diversity type conflict is detected, determining that the key HLA typing site is not detected and marking it.

[0008] In one embodiment, counting the mismatch records between the feature vector of the test sample and the data in the expected feature set includes: counting the mismatches caused by type conflicts, heterozygous non-coverage, and site non-detection; wherein, the type conflict includes the case where the base of the test sample does not match the base of the data in the expected feature set; the heterozygous non-coverage includes the case where the test sample is homozygous, but the expected feature set includes the base of one allele; and the site non-detection includes the case marked as not detected.

[0009] In one embodiment, determining the heterozygous combination with the fewest mismatches as the HLA typing result includes: in the case where multiple heterozygous combinations have the same number of mismatches, obtaining the frequencies of the heterozygous combinations with the same number of mismatches in the population based on an external database, and determining the heterozygous combination with the highest frequency as the HLA typing result.

[0010] In one embodiment, the method further includes: determining the confidence level of the current HLA typing result based on the mismatch rate and the frequency of the current HLA typing result in the population, and outputting the confidence level of the current HLA typing result.

[0011] The present application also provides an HLA genotyping device, which includes: a construction module for constructing a polymorphism database and generating a single nucleotide diversity feature matrix of the polymorphism database; a reconstruction module for obtaining HLA genotype data of a sample to be tested, reconstructing the HLA genotype data of the sample to be tested, and obtaining a feature vector of the sample to be tested; a generation module for obtaining two alleles at the same gene locus from the polymorphism database, generating all heterozygous combinations of the two alleles, and combining the single nucleotide diversity features in each of the heterozygous combinations to obtain an expected feature set for each of the heterozygous combinations; a comparison module for comparing the feature vector of the sample to be tested with the expected feature set bit by bit, counting the mismatch records of the data in the feature vector of the sample to be tested and the expected feature set, and recording the number of mismatches for each of the heterozygous combinations; and a determination module for determining the heterozygous combination with the least number of mismatches as the HLA typing result.

[0012] The present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method in the above implementation manner are implemented.

[0013] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method in the above implementation manner are implemented.

[0014] In the above HLA genotyping method, by comparing the detected single nucleotide diversity data with a pre-constructed polymorphism database and determining the heterozygous combination with the least number of mismatches as the typing result, the specific type of the HLA gene in the sample can be accurately determined. While improving the detection throughput and reducing the cost, this method also has a high typing accuracy and is applicable to application scenarios such as clinical transplantation, immunological research, and individualized treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained without creative efforts based on these drawings.

[0016] Figure 1 It is a flowchart of the HLA genotyping method in an embodiment; Figure 2 It is a system architecture diagram of the HLA genotyping method in an embodiment; Figure 3 It is a structural block diagram of the HLA genotyping device in an embodiment; Figure 4 It is the internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0017] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0018] The HLA (human leukocyte antigen) gene is of great significance in clinical detection due to its high polymorphism and close correlation with various diseases (such as autoimmune diseases, infectious diseases, tumors and organ transplantation rejection reactions). At present, there are various HLA typing methods, such as the specific primer amplification technique (PCR-SSP) that amplifies the target HLA gene with specific primers and determines the HLA genotype according to the electrophoresis results after amplification, the Sanger sequencing typing method (PCR-SBT) that directly sequences after PCR amplification and analyzes the base composition of the target sequence to determine the HLA genotype, and the next-generation sequencing technology (NGS) that deeply sequences the HLA gene using an NGS platform and determines the genotype by combining bioinformatics analysis, etc.

[0019] However, the traditional methods all have some problems that are difficult to solve more or less. For example, PCR-SSP (specific primer amplification technique) has low resolution and poor scalability, and it is difficult to detect new allele genotypes; although PCR-SBT (Sanger sequencing typing method) has relatively high resolution, its detection throughput is low, and due to problems such as low quality, heterozygous bases, and contamination signals in Sanger sequencing, HLA typing based on the detection results relies on manual interpretation; although NGS (next-generation sequencing technology) has high detection throughput, its detection cost is high, and HLA typing based on NGS detection data relies on sequence alignment, which takes a long time for analysis, and due to the high similarity between HLA types, there will be a situation where the alignment effect is not good.

[0020] As a high-throughput detection platform, a gene chip can detect SNP information in the target region using pre-designed probes on the chip, and can quickly obtain the complete sequence information of the exon region of the sample, thus providing a new idea for HLA typing. However, due to the extremely high polymorphism of the HLA gene and only tiny SNP differences between different types, there are still technical challenges in efficiently and accurately determining the HLA genotype of a sample from gene chip data.

[0021] To solve the above problems, the present invention proposes an HLA genotyping method, which can compare the gene chip detection data with a pre-constructed HLA genotype database to accurately determine the specific type of HLA gene in a sample.

[0022] Specifically, in one embodiment, as Figure 1 shown, an HLA genotyping method is provided. It can be understood that this method can be applied to a terminal, a server, or a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps: Step 110: Construct a polymorphism database and generate a single nucleotide diversity feature matrix of the polymorphism database; Step 120: Obtain the HLA genotype data of the sample to be tested, reconstruct the HLA genotype data of the sample to be tested, and obtain the feature vector of the sample to be tested; Step 130: Obtain two alleles at the same gene locus from the polymorphism database, generate all heterozygous combinations of the two alleles, and merge the single nucleotide diversity features in each heterozygous combination to obtain the expected feature set of each heterozygous combination; Step 140: Compare the feature vector of the sample to be tested with the expected feature set bit by bit, count the mismatch records of the data in the feature vector of the sample to be tested and the expected feature set, and record the number of mismatches of each heterozygous combination; Step 150: Determine the heterozygous combination with the least number of mismatches as the genotyping result.

[0023] For the convenience of expression, single nucleotide diversity is referred to as SNP (single nucleotide polymorphism) in the following text. The HLA gene has a high degree of polymorphism, and the differences between different alleles are usually concentrated in the exon regions encoding proteins.

[0024] Exemplarily, the polymorphism database of the HLA gene can select the IMGT / HLA database. The IMGT / HLA database is a special database maintained by the International Immunogenetics Information System (IMGT), which stores the detailed sequences and allele information of the human leukocyte antigen (HLA) gene.

[0025] Exemplarily, the HLA genotype data of the sample to be tested can be detected based on a solid-phase chip, or other means can be used, which are not specifically limited herein. Taking the solid-phase chip as an example, through microarray technology, the solid-phase chip can fix millions of oligonucleotide probes for detecting SNP sites on a glass slide, and according to the principle of base complementarity, capture and obtain the genetic information of the site to be tested through probe and fluorescence labeling design. In this application, the sample to be tested is detected to obtain the HLA genotype data of the sample to be tested; the HLA genotype data of the sample to be tested is reconstructed to obtain the SNP feature vector of the sample to be tested.

[0026] Select two alleles (allowing the same type) of the same gene locus from the database to generate all possible heterozygous combinations. For the two alleles within each combination, merge their SNP features and label them as the expected feature set (SNP feature set) of the combination. Compare the sample SNP feature vector with the expected feature set bit by bit and count various types of mismatches.

[0027] Arrange all the heterozygous combinations in ascending order according to the number of mismatches, and determine the heterozygous combination with the least number of mismatches as the typing result.

[0028] This embodiment has the following beneficial effects: By comparing the detection data with a pre-constructed polymorphism database and determining the heterozygous combination with the least number of mismatches as the typing result, the specific type of the HLA gene in the sample can be accurately determined. This method can improve the detection throughput and reduce the cost while having a high typing accuracy, and is applicable to application scenarios such as clinical transplantation, immunological research, and individualized treatment.

[0029] In one embodiment, constructing the polymorphism database and generating the single nucleotide polymorphism (SNP) diversity feature matrix of the polymorphism database includes: determining a reference HLA genotype and its corresponding reference sequence; obtaining the SNP diversity difference sites between the exon sequences of the remaining HLA genotypes in the polymorphism database except the reference HLA genotype and the reference sequence; generating the SNP diversity feature matrix of the polymorphism database based on the SNP diversity difference sites of each HLA genotype.

[0030] Specifically, the HLA gene has high polymorphism, and the differences between different alleles are usually concentrated in the exon region (especially the antigen-binding domain) encoding proteins. When constructing the HLA polymorphism database, extract the exon sequences of all alleles from the IMGT / HLA database, and use the first HLA genotype as the reference HLA genotype to obtain its corresponding reference sequence (Ref).

[0031] Record the SNP difference sites of each allele and the Ref to generate an SNP feature matrix, in the format of: {Allele: {Position 1: base, Position 2: base, ...}}.

[0032] The SNP feature matrix pre - calculation method provided in this embodiment takes the first allele as a reference, pre - calculates the SNP difference sites between all HLA alleles and the reference sequence, and forms a structured database. The computational load of real - time alignment is reduced through offline pre - calculation.

[0033] In one embodiment, reconstructing the HLA genotype data of the test sample to obtain the feature vector of the test sample includes: sorting based on the exon region coordinates of the HLA genotype data of the test sample and splicing to obtain a complete sequence; representing the heterozygous sites in the complete sequence with IUPAC degenerate codes to obtain the feature vector of the test sample.

[0034] In one embodiment, obtaining the HLA genotype data of the test sample includes: for the typing key sites of the test sample, setting two different probes for detection. If a single - nucleotide polymorphism type contradiction is detected, the typing key site is determined to be undetected and marked.

[0035] Specifically, input the SNP genotype data detected by the gene chip, sort it according to the exon region coordinates, and splice it into a complete sequence. Among them, the heterozygous sites are represented by IUPAC degenerate codes (such as Y = C / T) to generate the sample SNP feature vector. Among them, the IUPAC nomenclature is a method for systematically naming chemical substances stipulated by the International Union of Pure and Applied Chemistry (IUPAC).

[0036] For the typing key sites (i.e., the unique SNP sites that distinguish a certain single genotype from other genotypes), two different probes are set for detection. If a SNP type contradiction is detected, the result of the SNP site is determined to be undetected (represented by *).

[0037] The degenerate base coding mapping method in this embodiment converts the heterozygous sites (such as Y = C / T) detected by the chip into IUPAC codes and performs fuzzy matching with the SNP set of allele combinations in the database. It is compatible with the ambiguity of chip data and improves the accuracy of heterozygous typing.

[0038] In one embodiment, the statistics of the mismatch records between the feature vectors of the sample to be tested and the data in the expected feature set include: counting the mismatches caused by type conflicts, heterozygous non-coverage, and site non-detection; wherein, the type conflict includes the case where the base of the sample to be tested does not match the base of the data in the expected feature set; the heterozygous non-coverage includes the case where the sample to be tested is homozygous, but the expected feature set includes the base of one allele; the site non-detection includes the case marked as non-detected.

[0039] Specifically, two alleles of the same gene locus (allowing the same type) are selected from the polymorphism database to generate all possible heterozygous combinations (such as *01:01 / *02:01). For the two alleles within each combination, their SNP features are merged and marked as the expected feature set of the combination.

[0040] The SNP feature vector of the sample is compared bit by bit with the expected feature set, and the following types of mismatches are counted: Type conflict: The sample base does not match the expected base at all (such as the sample is C and the expected is A).

[0041] Heterozygous non-coverage: The sample is homozygous (such as C / C), but the combination only contains the base of one allele (such as C and T).

[0042] Site non-detection: The sample base is determined to be non-detected due to detection conflicts (* indicates).

[0043] Weight optimization: Assign a higher penalty weight to the mismatches in the key functional domains of exons (such as the antigen-binding region).

[0044] In this embodiment, through dynamic allele combination matching and hierarchical mismatch weight mechanism, according to the importance of HLA protein functional domains (such as antigen-binding grooves), different penalty weights are assigned to SNP mismatches in different regions (such as the mismatch weight of the key region = 3, non-key region = 1). Enhance clinical relevance and avoid misjudging the genotype due to non-critical mismatches.

[0045] In one embodiment, the determination of the heterozygous combination with the fewest mismatches as the typing result includes: arranging all the heterozygous combinations in ascending order according to the number of mismatches, and determining the heterozygous combination with the fewest mismatches as the typing result. In the case where the number of mismatches of multiple heterozygous combinations is the same, based on the external database, obtain the frequencies of the heterozygous combinations with the same number of mismatches in the population, and determine the heterozygous combination with the highest frequency as the typing result.

[0046] Specifically, all combinations are sorted in ascending order according to the total number of mismatches, and the combination with the fewest mismatches is preferentially selected. If multiple combinations have the same number of mismatches, the external database is further used to compare their frequencies in the population, and the one with the highest frequency is selected. Exemplarily, the external database here can be the CWD database (Common and Well-Documented alleles). The genotyping result and the confidence score are output.

[0047] In this embodiment, through the two-factor sorting strategy, the candidate combinations are preferentially sorted in ascending order according to the number of mismatches. If the number of mismatches is the same, they are sorted in descending order according to the population frequency of the allele combination. Combining statistical and biological significance improves the reliability of the results.

[0048] In one embodiment, the method further includes: determining the confidence of the current genotyping result based on the mismatch rate and the frequency of the current genotyping result in the population, and outputting the confidence of the current genotyping result.

[0049] Specifically, this embodiment also provides a confidence score model, which calculates the comprehensive confidence of the genotyping result based on the mismatch rate, the matching degree of the key region, and the population frequency. Quantifying the result credibility assists clinical decision-making.

[0050] In summary, this embodiment can efficiently implement HLA genotyping. By pre-computing the SNP feature matrix of HLA genotypes in the database, it performs genotyping calculations directly based on the matching of the SNP feature matrix without relying on sequence alignment, shortening the genotyping time to 1 / 5 of the traditional method. In addition, this embodiment also has the characteristics of high precision. By introducing the fuzzy matching of heterozygous sites and the hierarchical weight mechanism, and setting two probes for double detection at the key genotyping sites, the HLA genotyping accuracy can be increased to more than 99%.

[0051] To illustrate the HLA genotyping method of the present application in detail, the following is described with a most detailed embodiment: Taking the HLA genotyping method of the present application applied to the Figure 2 shown HLA genotyping system as an example, Figure 2The system architecture shown includes: a polymorphism database module 210, a chip data processing module 220, and a dynamic matching engine 230. Among them, the polymorphism database module is used to preprocess the full-genotype information of the IMGT database. Each gene takes the first allele genotype (*01:01:01:01) as a reference and records all SNP difference sites between different types and the reference. The chip data processing module is used to analyze the SNP genotype data output by the gene chip, reconstruct the complete sequence of the sample exon region, and mark the heterozygous sites as degenerate bases (such as R = A / G). The dynamic matching engine is used for mismatch analysis of the SNP characteristics based on the allele genotype combination and the sample SNPs to screen the optimal genotyping result.

[0052] Taking HLA-A gene typing as an example, first, data input is performed. The gene chip detects a total of 52 SNP sites in the HLA-A region of the sample to be tested, among which 3 are heterozygous (Y = C / T). Subsequently, sequence reconstruction is carried out to generate a sample SNP feature vector containing 52 sites, and the heterozygous sites are marked as Y. Then, the sample SNP feature vector of the sample to be tested is screened for mismatches with the alleles in the database. Taking the example that there are 1720 alleles from HLA-A*01:01:01 to HLA-A*74:13 in the database, the heterozygous combinations of all genotypes in this database are generated, and the sample SNP feature vector of the sample to be tested is screened for mismatches with each combination. The optimal combination is obtained as HLA-A*02:01:01 / *03:01:01, and the total number of mismatches is 2 (both in non-critical regions). Finally, the genotyping result is output: the genotyping result is HLA-A*02:01:01 / *03:01:01, and the confidence score is 98.7%.

[0053] It should be understood that although each step in the flowcharts involved in the above-described embodiments is shown in sequence according to the indication of the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same moment, but can be executed at different moments. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0054] Based on the same inventive concept, an embodiment of the present application further provides an HLA genotyping device for implementing the HLA genotyping method involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the HLA genotyping device provided below can refer to the limitations on the HLA genotyping method in the above text, and will not be repeated here.

[0055] In an exemplary embodiment, as Figure 3 shown, an HLA genotyping device is provided, and the device includes: A construction module 310, configured to construct a polymorphism database and generate a single nucleotide diversity feature matrix of the polymorphism database; A reconstruction module 320, configured to obtain HLA genotype data of a sample to be tested, reconstruct the HLA genotype data of the sample to be tested, and obtain a feature vector of the sample to be tested; A generation module 330, configured to obtain two alleles at the same gene locus from the polymorphism database, generate all heterozygous combinations of the two alleles, and merge the single nucleotide diversity features in each of the heterozygous combinations to obtain an expected feature set for each of the heterozygous combinations; A comparison module 340, configured to compare the feature vector of the sample to be tested with the expected feature set bit by bit, count the mismatch records of the data in the feature vector of the sample to be tested and the expected feature set, and record the number of mismatches for each of the heterozygous combinations; A determination module 350, configured to determine the heterozygous combination with the least number of mismatches as the genotyping result.

[0056] The construction module 310 is further configured to: Determine a reference HLA genotype and its corresponding reference sequence; obtain the single nucleotide diversity difference sites between the exon sequences of the remaining HLA genotypes in the polymorphism database except the reference HLA genotype and the reference sequence; generate a single nucleotide diversity feature matrix of the polymorphism database based on the single nucleotide diversity difference sites of each HLA genotype.

[0057] The reconstruction module 320 is further configured to: Sort based on the exon region coordinates of the HLA genotype data of the sample to be tested, splice to obtain a complete sequence; represent the heterozygous sites in the complete sequence with IUPAC degenerate codes to obtain the feature vector of the sample to be tested.

[0058] The reconstruction module 320 is further configured to: For the key typing sites of the sample to be tested, two different probes are set for detection. If a single nucleotide polymorphism type conflict is detected, the key typing site is determined to be undetected and marked.

[0059] The comparison module 340 is further configured to: Count the mismatches caused by type conflicts, heterozygous non-coverage, and undetected sites; wherein, the type conflict includes the case where the base of the sample to be tested does not match the base of the data in the expected feature set; the heterozygous non-coverage includes the case where the sample to be tested is homozygous, but the expected feature set includes the base of one allele; the undetected site includes the case marked as undetected.

[0060] The determination module 350 is further configured to: In the case where the number of mismatches of multiple heterozygous combinations is the same, obtain the frequency of the heterozygous combinations with the same number of mismatches in the population based on an external database, and determine the heterozygous combination with the highest frequency as the typing result.

[0061] The determination module 350 is further configured to: Determine the confidence level of the current typing result based on the mismatch rate and the frequency of the current typing result in the population, and output the confidence level of the current typing result.

[0062] Each module in the above HLA gene typing device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0063] In an exemplary embodiment, a computer device is provided, and the internal structure diagram of the computer device can be as Figure 4As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. The computer program, when executed by the processor, implements a method for calibrating the angle of a lidar.

[0064] Those skilled in the art can understand that Figure 4 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0065] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0066] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0067] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0068] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for HLA genotyping, characterized in that: The method comprises: Constructing a polymorphism database and generating a single nucleotide diversity feature matrix of the polymorphism database; Acquire HLA genotype data of the sample to be tested, reconstruct the HLA genotype data of the sample to be tested, and obtain a feature vector of the sample to be tested; Obtaining two alleles of the same gene locus from the polymorphism database, generating all heterozygous combinations of the two alleles, combining the single nucleotide diversity features in each of the heterozygous combinations, and obtaining an expected feature set of each heterozygous combination; Compare the feature vector of the sample to be tested with the expected feature set bit by bit, count the mismatch records between the feature vector of the sample to be tested and the data in the expected feature set, and record the number of mismatches of each of the heterozygous combinations; The heterozygous combination with the least number of mismatches was determined as the HLA typing result.

2. The method according to claim 1, characterized in that The step of constructing a polymorphism database and generating a single nucleotide diversity feature matrix of the polymorphism database comprises: Determine reference HLA genotypes and their corresponding reference sequences; Obtaining single nucleotide diversity difference sites between the exon sequences of the remaining HLA genotypes except the reference HLA genotype in the polymorphism database and the reference sequence; A single nucleotide diversity feature matrix of the polymorphism database is generated based on the single nucleotide diversity difference sites of each HLA genotype.

3. The method according to claim 1, characterized in that The step of reconstructing the HLA genotype data of the sample to be tested to obtain a feature vector of the sample to be tested includes: Sorting the exon region coordinates based on the HLA genotype data of the sample to be tested, and splicing to obtain a complete sequence; The heterozygous sites in the complete sequence are represented by IUPAC degenerate codes to obtain the characteristic vector of the sample to be tested.

4. The method according to claim 1, characterized in that: The step of obtaining HLA genotype data of the sample to be tested includes: For the HLA typing key site of the sample to be tested, two different probes are set for detection. If the detected single nucleotide diversity types are contradictory, the HLA typing key site is determined to be undetected and marked.

5. The method according to claim 4, characterized in that The counting of mismatch records between the feature vector of the sample to be tested and the data in the expected feature set includes: Statistics include mismatches caused by type conflict, heterozygous uncovered and site undetected; wherein, the type conflict includes the situation where the base of the sample to be tested does not match the base of the data in the expected feature set; the heterozygous uncovered includes the situation where the sample to be tested is homozygous, but the expected feature set includes the base of an allele; the site undetected includes the situation that is marked as undetected.

6. The method according to claim 1, characterized in that Determining the heterozygous combination with the least number of mismatches as the HLA typing result includes: When there are multiple heterozygous combinations with the same number of mismatches, the frequency of occurrence of heterozygous combinations with the same number of mismatches in the population is obtained based on an external database, and the heterozygous combination with the highest frequency is determined as the HLA typing result.

7. The method according to claim 1, characterized in that The method further comprises: The confidence of the current HLA typing result is determined based on the mismatch rate and the frequency of occurrence of the current HLA typing result in the population, and the confidence of the current HLA typing result is output.

8. An HLA genotyping device, characterized in that: The device comprises: A construction module is used to construct a polymorphism database and generate a single nucleotide diversity feature matrix of the polymorphism database; A reconstruction module, used to obtain HLA genotype data of the sample to be tested, reconstruct the HLA genotype data of the sample to be tested, and obtain a feature vector of the sample to be tested; A generation module, used for obtaining two alleles of the same gene locus from the polymorphism database, generating all heterozygous combinations of the two alleles, combining the single nucleotide diversity features in each of the heterozygous combinations, and obtaining an expected feature set of each heterozygous combination; A comparison module, used to compare the feature vector of the sample to be tested with the expected feature set bit by bit, count the mismatch records between the feature vector of the sample to be tested and the data in the expected feature set, and record the number of mismatches of each of the heterozygous combinations; A determination module is used to determine the heterozygous combination with the least number of mismatches as the HLA typing result.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and device for HLA genotyping, storage medium and processor

    CN110033827A

  • Blood type genotyping method and device and storage medium

    CN110942806A

  • Method for predicting HLA coincidence probability and mismatch types

    CN111613269A

  • HLA typing method based on next-generation sequencing data

    CN113409890A

  • HLA type determination method and device

    CN119028442A

Cited By

  • Human 80K molecular marker aiming at genome of Chinese population, solid-phase chip, kit and application of human 80K molecular marker

    CN122214491A