Hla genotyping methods, apparatuses, devices, and media

By constructing a polymorphism database and comparing feature matrices, the problems of insufficient efficiency and accuracy in HLA genotyping methods have been solved, achieving efficient and accurate HLA genotyping, which is suitable for clinical and research applications.

CN120148658BActive Publication Date: 2025-10-24SUZHOU LASSO BIOCHIP TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510629351.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-10-24
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

Existing HLA genotyping methods are inadequate in terms of efficiency and accuracy, especially in determining HLA genotypes efficiently and accurately from gene chip data, which remains a technical challenge.

Method used

By constructing a single nucleotide diversity feature matrix of a polymorphism database, HLA genotype data of the sample to be tested are obtained, feature vectors are reconstructed, and compared with the expected feature set bit by bit to count the number of mismatches. Finally, the heterozygous combination with the fewest mismatches is determined as the typing result.

Benefits of technology

It improves detection throughput, reduces costs, and has a high typing accuracy, making it suitable for clinical transplantation, immunology research and personalized treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148658B_ABST
    Figure CN120148658B_ABST
Patent Text Reader

Abstract

The application relates to an HLA genotyping method, device, equipment and medium. The method comprises the following steps: constructing a polymorphism database, generating an SNP feature matrix of the polymorphism database; obtaining HLA genotype data of a to-be-tested sample, reconstructing the HLA genotype data of the to-be-tested sample, and obtaining a feature vector of the to-be-tested sample; obtaining two alleles of the same gene site from the polymorphism database, generating all hybrid combinations of the two alleles, merging single nucleotide polymorphism features in each hybrid combination, and obtaining an expected feature set of each hybrid combination; comparing the feature vector of the to-be-tested sample with the expected feature set bit by bit, counting mismatch records of data in the feature vector of the to-be-tested sample and the expected feature set, and recording the number of mismatches of each hybrid combination; and determining the hybrid combination with the least number of mismatches as an HLA typing result. The method can efficiently and accurately perform gene typing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioengineering, and in particular to an HLA genotyping method, device, equipment and medium. BACKGROUND

[0002] The HLA (human leukocyte antigen) gene is of great significance in clinical detection due to its high polymorphism and close correlation with various diseases. There are currently various HLA typing methods, such as specific primer amplification technology (PCR-SSP), Sanger sequencing typing method (PCR-SBT) and high-throughput sequencing technology (NGS) and the like.

[0003] As a high-throughput detection platform, the gene chip detects SNP (single nucleotide polymorphism) information of the target region using the probes pre-designed on the chip, and can quickly obtain complete sequence information of the exon region of the sample, thereby providing a new idea for HLA typing. However, due to the extremely polymorphic HLA gene, only slight SNP differences exist between different types, and the current HLA genotyping method still has deficiencies in efficiency and accuracy. SUMMARY

[0004] Therefore, it is necessary to provide an efficient and accurate HLA genotyping method, which comprises the following steps:

[0005] Constructing a polymorphism database and generating a single nucleotide polymorphism feature matrix of the polymorphism database; obtaining HLA genotype data of a to-be-tested sample, reconstructing the HLA genotype data of the to-be-tested sample, and obtaining a feature vector of the to-be-tested sample; obtaining two alleles of the same gene site from the polymorphism database, generating all hybrid combinations of the two alleles, merging the single nucleotide polymorphism features in each hybrid combination, and obtaining an expected feature set of each hybrid combination; comparing the feature vector of the to-be-tested sample with the expected feature set bit by bit, counting the mismatch records of the data in the feature vector of the to-be-tested sample and the expected feature set, and recording the mismatch number of each hybrid combination; and determining the hybrid combination with the least mismatch number as the HLA typing result.

[0006] In an embodiment, the constructing the polymorphism database and generating the single nucleotide polymorphism feature matrix of the polymorphism database comprises: determining a reference HLA genotype and a corresponding reference sequence thereof; obtaining single nucleotide polymorphism difference sites of exon sequences of remaining HLA genotypes in the polymorphism database except for the reference HLA genotype and the reference sequence; and generating the single nucleotide polymorphism feature matrix of the polymorphism database based on the single nucleotide polymorphism difference sites of each HLA genotype.

[0007] In an embodiment, the reconstructing the HLA genotype data of the to-be-tested sample to obtain the feature vector of the to-be-tested sample comprises: sorting based on exon region coordinates of the HLA genotype data of the to-be-tested sample to splice a complete sequence; and representing heterozygous sites in the complete sequence in IUPAC degenerate code to obtain the feature vector of the to-be-tested sample.

[0008] In an embodiment, the obtaining the HLA genotype data of the to-be-tested sample comprises: for an HLA typing key site of the to-be-tested sample, setting two different probes for detection, and if a single nucleotide polymorphism type is detected to be contradictory, determining the HLA typing key site as undetected and marking.

[0009] In an embodiment, the counting mismatch records of the feature vector of the to-be-tested sample and data in the expected feature set comprises: counting mismatches caused by type conflicts, heterozygous non-coverage, and undetected sites; wherein the type conflict comprises a case that a base of the to-be-tested sample does not match a base of data in the expected feature set; the heterozygous non-coverage comprises a case that the to-be-tested sample is homozygous but a base of one allele is included in the expected feature set; and the undetected site comprises a case that is marked as undetected.

[0010] In an embodiment, the determining the heterozygous combination with the least number of mismatches as the HLA typing result comprises: in a case that there are multiple heterozygous combinations with the same number of mismatches, obtaining a frequency of the heterozygous combination with the same number of mismatches in a population based on an external database, and determining a heterozygous combination with the highest frequency as the HLA typing result.

[0011] In an embodiment, the method further comprises: determining a confidence of the current HLA typing result based on a mismatch rate and a frequency of the current HLA typing result in a population, and outputting the confidence of the current HLA typing result.

[0012] The application further provides an HLA genotyping device, which comprises: a construction module configured to construct a polymorphism database and generate a single nucleotide polymorphism feature matrix of the polymorphism database; a reconstruction module configured to obtain HLA genotype data of a sample to be tested, reconstruct the HLA genotype data of the sample to be tested, and obtain a feature vector of the sample to be tested; a generation module configured to obtain two alleles of a same genetic locus from the polymorphism database, generate all hybrid combinations of the two alleles, and combine single nucleotide polymorphism features in each of the hybrid combinations to obtain an expected feature set of each hybrid combination; a comparison module configured to compare the feature vector of the sample to be tested with the expected feature set bit by bit, count mismatch records of data in the feature vector of the sample to be tested and the expected feature set, and record the number of mismatches of each hybrid combination; and a determination module configured to determine a hybrid combination with the least number of mismatches as an HLA typing result.

[0013] The application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method of the above-mentioned embodiments when executing the computer program.

[0014] The application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the method of the above-mentioned embodiments.

[0015] The HLA genotyping method can accurately determine the specific type of HLA genes in a sample by detecting single nucleotide polymorphism data and comparing the data with a pre-constructed polymorphism database, determining a hybrid combination with the least number of mismatches as a typing result. The method has high typing accuracy while improving detection throughput and reducing costs, and is suitable for clinical transplantation, immunological research, individualized treatment and other application scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the application or the related art, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the application or the related art. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0017] Figure 1 FIG. 1 is a flowchart of an HLA genotyping method in an embodiment;

[0018] Figure 2 FIG. 1 is a flowchart of an HLA genotyping method in an embodiment;

[0019] Figure 3A structural block diagram of an HLA genotyping device in an embodiment;

[0020] Figure 4 An internal structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION

[0021] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.

[0022] The HLA (human leukocyte antigen) gene is of great significance in clinical detection due to its high polymorphism and close correlation with various diseases (such as autoimmune diseases, infectious diseases, tumors and organ transplant rejection). At present, there are various HLA typing methods, such as specific primer amplification of target HLA genes, specific primer amplification technology (PCR-SSP) for judging HLA genotypes according to electrophoresis results after amplification, Sanger sequencing typing method (PCR-SBT) for determining HLA genotypes by directly sequencing after PCR amplification and analyzing the base composition of the target sequence, and high-throughput sequencing technology (NGS) for determining genotypes by using NGS platform for deep sequencing of HLA genes and combining bioinformatics analysis.

[0023] However, the traditional methods more or less have some problems that are difficult to solve. For example, PCR-SSP (specific primer amplification technology) has low resolution, and the method has poor scalability and is difficult to detect new allelic types; PCR-SBT (Sanger sequencing typing method) has high resolution, but has low detection throughput, and since Sanger sequencing has problems such as low quality, hybrid base, and pollution signal, HLA typing based on the detection results relies on manual interpretation; NGS (high-throughput sequencing technology) has high detection throughput, but has high detection cost, and HLA typing based on NGS detection data relies on sequence alignment, which is time-consuming and laborious, and since HLA types have high similarity, there may be poor alignment results.

[0024] As a high-throughput detection platform, gene chips can quickly obtain complete sequence information of the exon region of the sample by detecting SNP information of the target region using probes pre-designed on the chip, thereby providing a new idea for HLA typing. However, since the HLA gene is highly polymorphic, there are only slight SNP differences between different types, and how to efficiently and accurately determine the HLA genotype of the sample from the gene chip data still has technical challenges.

[0025] To solve the above problems, the present invention proposes an HLA genotyping method, which can accurately determine the specific type of HLA gene in a sample by comparing gene chip detection data with a pre-constructed HLA genotype database.

[0026] Specifically, in one embodiment, Figure 1 As shown, an HLA genotyping method is provided. It is understood that the method can be applied to a terminal, a server, or a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0027] Step 110, constructing a polymorphism database and generating a single nucleotide diversity feature matrix of the polymorphism database;

[0028] Step 120, obtaining HLA genotype data of the sample to be tested, reconstructing the HLA genotype data of the sample to be tested, and obtaining a feature vector of the sample to be tested;

[0029] Step 130, obtaining two alleles of the same gene locus from the polymorphism database, generating all heterozygous combinations of the two alleles, and combining the single nucleotide diversity features in each of the heterozygous combinations to obtain an expected feature set for each heterozygous combination;

[0030] Step 140, comparing the feature vector of the sample to be tested with the expected feature set bit by bit, counting mismatch records between the feature vector of the sample to be tested and the data in the expected feature set, and recording the number of mismatches for each of the heterozygous combinations;

[0031] Step 150: determine the heterozygous combination with the least number of mismatches as the typing result.

[0032] For ease of expression, single nucleotide polymorphism (SNP) will be used in the following text to refer to single nucleotide polymorphism. HLA genes are highly polymorphic, and the differences between different alleles are generally concentrated in the protein-coding exon region.

[0033] For example, the polymorphism database of the HLA gene may be the IMGT / HLA database, which is a specialized database maintained by the International Immunogenetics Information System (IMGT) that stores detailed sequence and allele information of human leukocyte antigen (HLA) genes.

[0034] Exemplarily, the HLA genotype data of the sample to be detected can be detected based on a solid-phase chip or by using other means, which are not limited herein. Taking the solid-phase chip as an example, the solid-phase chip can fix the million-level oligonucleotide probes for detecting SNP sites on a glass slide by microarray technology, and capture and obtain the genetic information of the site to be detected by the probe and fluorescence labeling design according to the base complementary principle. In the present application, the sample to be detected is detected to obtain the HLA genotype data of the sample to be detected; and the HLA genotype data of the sample to be detected is reconstructed to obtain the SNP feature vector of the sample to be detected.

[0035] Two alleles of the same gene site (allowing the same type) are selected from the database, all possible hybrid combinations are generated, for each combination of two alleles, the SNP features are combined and marked as the expected feature set (SNP feature set) of the combination. The sample SNP feature vector is compared with the expected feature set bit by bit, and the number of various types of mismatches is counted.

[0036] All the hybrid combinations are arranged in ascending order according to the number of mismatches, and the hybrid combination with the least number of mismatches is determined as the typing result.

[0037] The embodiment has the following beneficial effects:

[0038] By comparing the detection data with the pre-constructed polymorphism database, the hybrid combination with the least number of mismatches is determined as the typing result, which can accurately determine the specific type of the HLA gene in the sample. This method has high typing accuracy while improving the detection throughput and reducing the cost, and is suitable for clinical transplantation, immunological research and individualized treatment application scenarios.

[0039] In one embodiment, the constructing the polymorphism database and generating the single nucleotide polymorphism feature matrix of the polymorphism database comprises: determining a reference HLA genotype and a reference sequence corresponding to the reference HLA genotype; acquiring single nucleotide polymorphism difference sites of exon sequences of remaining HLA genotypes in the polymorphism database except the reference HLA genotype and the reference sequence; and generating the single nucleotide polymorphism feature matrix of the polymorphism database based on the single nucleotide polymorphism difference sites of each HLA genotype.

[0040] Specifically, the HLA gene has high polymorphism, and the differences between different alleles are usually concentrated in the exon region (especially the antigen binding domain) coding protein. When constructing the HLA polymorphism database, the exon sequences of all alleles are extracted from the IMGT / HLA database, and the first HLA genotype is taken as a reference HLA genotype to obtain a reference sequence (Ref) corresponding to the reference HLA genotype.

[0041] Record the SNP difference site of each allele with Ref, generate SNP feature matrix, format is:

[0042] {allele: {position1: base, position2: base,...}}.

[0043] The SNP feature matrix pre-computation method provided in this embodiment takes the first allele as a reference to pre-compute the SNP difference sites of all HLA alleles with the reference sequence to form a structured database. The pre-computation offline reduces the computation amount of real-time comparison.

[0044] In one embodiment, the HLA genotype data of the to-be-tested sample is reconstructed to obtain a feature vector of the to-be-tested sample, including: based on the exon region coordinates of the HLA genotype data of the to-be-tested sample, the complete sequence is obtained by sorting and splicing; the heterozygous site in the complete sequence is represented by IUPAC degenerate code to obtain the feature vector of the to-be-tested sample.

[0045] In one embodiment, the HLA genotype data of the to-be-tested sample is obtained, including: for the typing key site of the to-be-tested sample, two different probes are set for detection, if the single nucleotide polymorphism type detected is contradictory, the typing key site is determined as undetected and marked.

[0046] Specifically, the SNP genotype data detected by the gene chip is input, sorted according to the exon region coordinates, and spliced into a complete sequence. Among them, the heterozygous site is represented by IUPAC degenerate code (such as Y=C / T), and the sample SNP feature vector is generated. Among them, IUPAC naming method is a method of systematically naming chemical substances stipulated by International Union of Pure and Applied Chemistry (IUPAC).

[0047] For the typing key site (i.e. the unique SNP site for distinguishing a single genotype from other genotypes), two different probes are set for detection, and if the detected SNP type is contradictory, the SNP site result is determined as undetected (represented by *).

[0048] The degenerate base encoding mapping method of this embodiment converts the heterozygous site (such as Y=C / T) detected by the chip into IUPAC code, and performs fuzzy matching with the allele combined SNP set in the database. Compatible with the ambiguity of chip data, improve the accuracy of heterozygous typing.

[0049] In one embodiment, the step of counting the mismatches between the feature vector of the sample to be tested and the data in the expected feature set comprises counting mismatches including type conflict, heterozygous uncovered and site undetected; wherein the type conflict comprises a case that the base of the sample to be tested does not match the base of the data in the expected feature set; the heterozygous uncovered comprises a case that the sample to be tested is homozygous, but the base of one allele is included in the expected feature set; and the site undetected comprises a case that is marked as undetected.

[0050] Specifically, two alleles of the same gene site (allowing the same type) are selected from the polymorphism database to generate all possible heterozygous combinations (such as *01:01 / *02:01). For each combination of two alleles, the SNP features of the two alleles are combined and marked as the expected feature set of the combination.

[0051] The SNP feature vector of the sample is compared with the expected feature set bit by bit, and the following types of mismatches are counted:

[0052] Type conflict: the base of the sample does not match the expected base at all (such as sample C and expected A).

[0053] Heterozygous uncovered: the sample is homozygous (such as C / C), but the combination contains only one base of the allele (such as C and T).

[0054] Site undetected: the base of the sample is undetected due to the existence of detection contradiction (*).

[0055] Weight optimization: higher penalty weight is given to the mismatch of the key functional domain of the exon (such as the antigen binding region).

[0056] This embodiment matches the dynamic allele combination and the hierarchical mismatch weight mechanism, and assigns different penalty weights to the SNP mismatches in different regions according to the importance of the HLA protein functional domain (such as the antigen binding groove) (such as key region mismatch weight = 3, non-key region = 1). Enhance clinical relevance and avoid false typing due to non-key mismatches.

[0057] In one embodiment, the step of determining the heterozygous combination with the least number of mismatches as the typing result comprises arranging all the heterozygous combinations in ascending order of the number of mismatches, and determining the heterozygous combination with the least number of mismatches as the typing result. In the case that there are multiple heterozygous combinations with the same number of mismatches, the frequency of the heterozygous combination with the same number of mismatches in the population is obtained based on an external database, and the heterozygous combination with the highest frequency is determined as the typing result.

[0058] Specifically, all combinations are arranged in ascending order of total mismatch number, and combinations with the least number of mismatches are preferentially selected. If multiple combinations have the same number of mismatches, the combination with the highest frequency in the population is selected by further comparing the frequency in the external database. Exemplarily, the external database here can select the CWD database (common and well-documented alleles). The typing result and the confidence score are output.

[0059] The embodiment preferentially arranges candidate combinations in ascending order of mismatch number by a two-factor sorting strategy, and if the mismatch numbers are the same, the combinations are arranged in descending order of the frequency of the allele combination in the population. The reliability of the result is improved by combining statistical and biological significance.

[0060] In one embodiment, the method further comprises determining the confidence of the current typing result based on the mismatch rate and the frequency of the current typing result in the population, and outputting the confidence of the current typing result.

[0061] Specifically, the embodiment also provides a confidence score model for calculating the comprehensive confidence of the typing result based on the mismatch rate, the matching degree of the key region, and the frequency in the population. The reliability of the result is quantified to assist clinical decision-making.

[0062] In summary, the embodiment can efficiently realize HLA genotyping. By precomputing the SNP feature matrix of HLA genotypes in the database, the genotyping time is shortened to 1 / 5 of that of the traditional method without relying on sequence alignment but directly based on SNP feature matrix matching. In addition, the embodiment also has the characteristics of high precision. The hybrid site fuzzy matching and hierarchical weight mechanism are introduced, and two probes are set for double detection of the key typing sites, which can improve the accuracy of HLA typing to more than 99%.

[0063] To illustrate the HLA genotyping method of the present application in detail, a most detailed embodiment is described below:

[0064] The HLA genotyping method of the present application is applied to an HLA genotyping system as shown in Figure 2 , for example,

[0065] Figure 2The system architecture shown includes: polymorphism database module 210, chip data processing module 220, dynamic matching engine 230. Among them, the polymorphism database module is used to preprocess the IMGT database full genotype information, each gene takes the first allelic genotype (*01:01:01:01) as reference, records all type differences with reference SNP difference sites; the chip data processing module is used to analyze the SNP genotype data output by the gene chip, reconstructs the complete sequence of the sample exon region, and marks the heterozygous site as a degenerate base (such as R=A / G). The dynamic matching engine is used for mismatch analysis based on the SNP characteristics of the allelic genotype combination and the sample SNP, and the optimal typing result is screened.

[0066] Taking HLA-A gene typing as an example, first, data input is performed, the gene chip detects that the HLA-A region of the sample to be tested has a total of 52 SNP sites, of which 3 are heterozygous (Y=C / T). Then, sequence reconstruction is performed to generate a sample SNP feature vector containing 52 sites, and the heterozygous site is marked as Y. Then, the sample SNP feature vector of the sample to be tested is mismatched with the alleles in the database. Taking the 1720 alleles of HLA-A*01:01:01 to HLA-A*74:13 in the database as an example, the heterozygous combination of all genotypes in the database is generated, and the sample SNP feature vector of the sample to be tested is mismatched with each combination to obtain the optimal combination HLA-A*02:01:01 / *03:01:01, and the total number of mismatches is 2 (both are non-critical regions). Finally, the typing result is output: the typing result is HLA-A*02:01:01 / *03:01:01, and the confidence score is 98.7%.

[0067] It should be understood that although each step in the flowchart involved in each embodiment as described above is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or steps or stages in other steps.

[0068] Based on the same inventive concept, the application further provides an HLA genotyping device for implementing the above-mentioned HLA genotyping method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, and therefore the specific limitations in one or more HLA genotyping device embodiments provided below can refer to the limitations of the HLA genotyping method described above, which will not be repeated here.

[0069] In one exemplary embodiment, as shown in Figure 3 An HLA genotyping device is provided, and the device comprises:

[0070] The construction module 310 is configured to construct a polymorphism database and generate a single nucleotide polymorphism feature matrix of the polymorphism database.

[0071] The reconstruction module 320 is configured to obtain HLA genotype data of a sample to be tested, reconstruct the HLA genotype data of the sample to be tested, and obtain a feature vector of the sample to be tested.

[0072] The generation module 330 is configured to obtain two alleles of the same gene locus from the polymorphism database, generate all heterozygous combinations of the two alleles, and combine single nucleotide polymorphism features in each heterozygous combination to obtain an expected feature set of each heterozygous combination.

[0073] The comparison module 340 is configured to compare the feature vector of the sample to be tested with the expected feature set bit by bit, count mismatch records of data in the feature vector of the sample to be tested and the expected feature set, and record the number of mismatches of each heterozygous combination.

[0074] The determination module 350 is configured to determine the heterozygous combination with the least number of mismatches as a typing result.

[0075] The construction module 310 is further configured to:

[0076] determine a reference HLA genotype and a reference sequence corresponding thereto, obtain single nucleotide polymorphism difference sites between exon sequences of remaining HLA genotypes in the polymorphism database and the reference sequence except for the reference HLA genotype, and generate a single nucleotide polymorphism feature matrix of the polymorphism database based on the single nucleotide polymorphism difference sites of each HLA genotype.

[0077] The reconstruction module 320 is further configured to:

[0078] sort based on exon region coordinates of the HLA genotype data of the sample to be tested, splice to obtain a complete sequence, represent heterozygous sites in the complete sequence in IUPAC degenerate code, and obtain the feature vector of the sample to be tested.

[0079] The reconstruction module 320 is further configured to:

[0080] For the typing key sites of the sample to be tested, two different probes are set to detect, if the single nucleotide polymorphism type detected is contradictory, the typing key site is determined as not detected and marked.

[0081] The comparison module 340 is further configured to:

[0082] Statistics include type conflict, hybrid coverage and site not detected mismatch; wherein the type conflict includes the case that the base of the sample to be tested does not match the base of the data in the expected feature set; the hybrid coverage includes the case that the sample to be tested is homozygous, but the expected feature set includes a base of one allele; the site not detected includes the case of being marked as not detected.

[0083] The determination module 350 is further configured to:

[0084] In the case that there are multiple heterozygous combinations with the same number of mismatches, the frequency of the heterozygous combination with the same number of mismatches in the population is obtained based on the external database, and the heterozygous combination with the highest frequency is determined as the typing result.

[0085] The determination module 350 is further configured to:

[0086] Based on the mismatch rate and the frequency of the current typing result in the population, the confidence of the current typing result is determined, and the confidence of the current typing result is output.

[0087] The above-mentioned various modules in the HLA gene typing device can be realized by software, hardware and their combinations in whole or in part. The above-mentioned various modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to call and execute the operations corresponding to the above-mentioned various modules by the processor.

[0088] In an exemplary embodiment, a computer device is provided, and the internal structure diagram of the computer device can be as shown in Figure 4As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. Among them, the processor, the memory and the input / output interface are connected through the system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with the external terminal in a wired or wireless manner. The wireless manner can be realized through WIFI, mobile cellular network, Near Field Communication (NFC) or other technologies. The computer program is executed by the processor to realize a laser radar angle calibration method.

[0089] Those skilled in the art can understand that, Figure 4 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0090] In one embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by the processor to realize the steps in each method embodiment described above.

[0091] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.

[0092] The technical features of the above embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.

[0093] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.

Claims

1. An HLA genotyping device, characterized by, The device comprises: a construction module for constructing a polymorphism database and generating a single nucleotide polymorphism feature matrix of the polymorphism database; the construction of the polymorphism database and the generation of the single nucleotide polymorphism feature matrix of the polymorphism database comprise: determining a reference HLA genotype and a corresponding reference sequence thereof; acquiring single nucleotide polymorphism difference sites of an exon sequence of a remaining HLA genotype in the polymorphism database other than the reference HLA genotype and the reference sequence; and generating a single nucleotide polymorphism feature matrix of the polymorphism database based on the single nucleotide polymorphism difference sites of each HLA genotype; a reconstruction module for acquiring HLA genotype data of a sample to be tested, reconstructing the HLA genotype data of the sample to be tested, and obtaining a feature vector of the sample to be tested; the acquisition of the HLA genotype data of the sample to be tested comprises: for an HLA typing key site of the sample to be tested, setting two different probes for detection, and if a single nucleotide polymorphism type is detected to be contradictory, determining the HLA typing key site as undetected and marking it; the reconstruction of the HLA genotype data of the sample to be tested to obtain the feature vector of the sample to be tested comprises: sorting based on exon region coordinates of the HLA genotype data of the sample to be tested to splice a complete sequence; and representing a heterozygous site in the complete sequence in IUPAC degenerate code to obtain the feature vector of the sample to be tested; a generation module for acquiring two alleles of a same gene site from the polymorphism database, generating all heterozygous combinations of the two alleles, and merging single nucleotide polymorphism features in each of the heterozygous combinations to obtain an expected feature set of each of the heterozygous combinations; a comparison module for comparing the feature vector of the sample to be tested with the expected feature set bit by bit, counting mismatch records of data in the feature vector of the sample to be tested and the expected feature set, and recording mismatch numbers of each of the heterozygous combinations; the counting of the mismatch records of the data in the feature vector of the sample to be tested and the expected feature set comprises: counting mismatches caused by type conflicts, heterozygous non-coverage, and undetected sites; the type conflicts comprise a case that a base of the sample to be tested does not match a base of data in the expected feature set; the heterozygous non-coverage comprises a case that the sample to be tested is homozygous but a base of one allele is included in the expected feature set; and the undetected sites comprise a case that is marked as undetected; a determination module for determining a heterozygous combination with the least mismatch number as an HLA typing result; and the determination of the heterozygous combination with the least mismatch number as the HLA typing result comprises: in a case that there are multiple heterozygous combinations with the same mismatch number, acquiring a frequency of the heterozygous combinations with the same mismatch number in a population based on an external database, and determining a heterozygous combination with the highest frequency as the HLA typing result.

2. The apparatus of claim 1, wherein, The determination module is further configured to determine the confidence of the current HLA typing result based on the mismatch rate and the frequency of the current HLA typing result in a population, and output the confidence of the current HLA typing result. 3.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is characterized in that, The processor implements the steps of the HLA genotyping method when executing the computer program. The HLA genotyping method comprises: constructing a polymorphism database and generating a single nucleotide diversity feature matrix of the polymorphism database; the constructing a polymorphism database and generating a single nucleotide diversity feature matrix of the polymorphism database comprises: determining a reference HLA genotype and a corresponding reference sequence thereof; acquiring single nucleotide diversity difference sites of an exon sequence of a remaining HLA genotype in the polymorphism database other than the reference HLA genotype and the reference sequence; and generating a single nucleotide diversity feature matrix of the polymorphism database based on the single nucleotide diversity difference sites of each HLA genotype; acquiring HLA genotype data of a to-be-tested sample, reconstructing the HLA genotype data of the to-be-tested sample, and obtaining a feature vector of the to-be-tested sample; the acquiring HLA genotype data of a to-be-tested sample comprises: for an HLA typing key site of the to-be-tested sample, setting two different probes for detection, and if a single nucleotide diversity type is detected to be contradictory, determining the HLA typing key site as undetected and marking it; the reconstructing the HLA genotype data of the to-be-tested sample comprises: sorting based on exon region coordinates of the HLA genotype data of the to-be-tested sample to splice a complete sequence; representing a heterozygous site in the complete sequence in IUPAC degenerate code to obtain the feature vector of the to-be-tested sample; acquiring two alleles of a same gene site from the polymorphism database, generating all heterozygous combinations of the two alleles, and merging single nucleotide diversity features in each heterozygous combination to obtain an expected feature set of each heterozygous combination; comparing the feature vector of the to-be-tested sample with the expected feature set bit by bit, counting mismatch records of data in the feature vector of the to-be-tested sample and the expected feature set, and recording mismatch numbers of each heterozygous combination; the counting mismatch records of data in the feature vector of the to-be-tested sample and the expected feature set comprises: counting mismatches caused by type conflicts, heterozygous non-coverage, and non-detection of sites; the type conflicts comprise a case that a base of the to-be-tested sample does not match a base of data in the expected feature set; the heterozygous non-coverage comprises a case that the to-be-tested sample is homozygous but a base of one allele is included in the expected feature set; and the non-detection of sites comprises a case that is marked as undetected. determine the hybrid combination with the least number of mismatches as the HLA typing result; and in the case that there are multiple hybrid combinations with the same number of mismatches, determine the hybrid combination with the highest frequency in the population as the HLA typing result based on the external database.

4. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the HLA genotyping method. The HLA genotyping method comprises: constructing a polymorphism database and generating a single nucleotide polymorphism feature matrix of the polymorphism database; the constructing a polymorphism database and generating a single nucleotide polymorphism feature matrix of the polymorphism database comprises: determining a reference HLA genotype and a corresponding reference sequence thereof; obtaining single nucleotide polymorphism difference sites of an exon sequence of a remaining HLA genotype in the polymorphism database other than the reference HLA genotype and the reference sequence; and generating a single nucleotide polymorphism feature matrix of the polymorphism database based on the single nucleotide polymorphism difference sites of each HLA genotype; obtaining HLA genotype data of a to-be-tested sample, reconstructing the HLA genotype data of the to-be-tested sample, and obtaining a feature vector of the to-be-tested sample; the obtaining HLA genotype data of a to-be-tested sample comprises: for an HLA typing key site of the to-be-tested sample, setting two different probes for detection, and if a single nucleotide polymorphism type is detected to be contradictory, determining the HLA typing key site as undetected and marking it; the reconstructing the HLA genotype data of the to-be-tested sample and obtaining a feature vector of the to-be-tested sample comprises: sorting based on exon region coordinates of the HLA genotype data of the to-be-tested sample to splice a complete sequence; and representing hybrid sites in the complete sequence in IUPAC degenerate code to obtain the feature vector of the to-be-tested sample; obtaining two alleles of the same gene site from the polymorphism database, generating all hybrid combinations of the two alleles, and merging single nucleotide polymorphism features in each hybrid combination to obtain an expected feature set of each hybrid combination; comparing the feature vector of the to-be-tested sample with the expected feature set bit by bit, counting mismatch records of data in the feature vector of the to-be-tested sample and the expected feature set, and recording the number of mismatches of each hybrid combination; the counting mismatch records of data in the feature vector of the to-be-tested sample and the expected feature set comprises: counting mismatches caused by type conflicts, hybrid non-coverage, and site undetected; the type conflict comprises a case that a base of the to-be-tested sample does not match a base of data in the expected feature set; the hybrid non-coverage comprises a case that the to-be-tested sample is homozygous but a base of one allele is included in the expected feature set; and the site undetected comprises a case that is marked as undetected. determining the hybrid combination with the least number of mismatches as the HLA typing result; and determining, in the case where there are multiple hybrid combinations with the same number of mismatches, the hybrid combination with the highest frequency in the population based on an external database to obtain the frequency of the hybrid combinations with the same number of mismatches, the hybrid combination with the highest frequency as the HLA typing result.

Citation Information

Patent Citations

  • HLA typing method based on next-generation sequencing data

    CN113409890A

  • Gene haplotype typing method and device based on sequencing data and medium

    CN119580843A