Model construction method and device for evaluating risk of occult hbv infection of subject and application

By detecting specific SNP sites in HLA genes and constructing models, the challenge of assessing the risk of occult HBV infection has been solved, achieving high accuracy and sensitivity in risk assessment and supporting clinical and blood donation screening.

CN120924724BActive Publication Date: 2026-04-07BEIJING HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively assess the risk of occult HBV infection, leading to an increased risk of transmission of occult HBV in blood transfusions or organ transplants, and there is a lack of effective genetic assessment methods.

Method used

By detecting specific SNP sites in HLA genes, a model for assessing the risk of occult HBV infection in subjects was constructed. The linkage disequilibrium and variation influence of SNP sites were used for screening, and a classifier model was constructed by combining the minor allele frequency and correlation coefficient.

Benefits of technology

It achieves high specificity and sensitivity in assessing the risk of occult HBV infection, with an accuracy of 85%, providing a scientific basis for clinical and blood donation screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120924724B_ABST
    Figure CN120924724B_ABST
Patent Text Reader

Abstract

The application discloses a model construction method and equipment for evaluating the risk of occult HBV infection of a subject and application. According to an embodiment of the application, 24 SNP sites in HLA gene regions are obtained by a specific screening method, and the 24 SNP sites are used as molecular markers of occult HBV infection (OBI) and are used for model construction, so that genetic risk evaluation of OBI is realized. Experimental results show that the evaluation index AUC of the application is 0.86, and the application has the characteristics of high specificity, high sensitivity and high accuracy, and can provide more comprehensive, accurate and individualized evaluation for the risk of OBI.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of occult hepatitis B virus infection evaluation, in particular to a model construction method and equipment and application for evaluating the risk of occult HBV infection of a subject. BACKGROUND

[0002] Occult hepatitis B virus infection (OBI) is a special state existing in the process of HBV infection, that is, the hepatitis B surface antigen (HBsAg) cannot be detected in the blood under the existing detection conditions, but the replication-competent HBV DNA exists in the liver tissue. The detection of rcDNA or cccDNA in liver tissue by puncture is the gold standard for the diagnosis of OBI, but it is limited by the invasiveness of liver puncture. In clinical practice, OBI is often diagnosed according to the detection of HBV DNA in the blood circulation, combined with long-term follow-up.

[0003] Adults infected with HBV generally have a 5%-10% chance of developing chronicity. Among them, the dominant HBV infection is positive for HBsAg and generally has a high level of HBV DNA (usually more than 200 IU / mL), so it is easier to detect whether it is clinical screening or blood donor detection. The occult HBV infection is negative for HBsAg and has very low viral load (less than 200 IU / mL or only fluctuation detection), but still increases the risk of HBV transmission caused by blood transfusion or organ transplantation, thereby posing a great threat to public health and blood safety.

[0004] A large number of studies have shown that host factors may play a more important role than viral factors in OBI. HBV-specific CD8+ cytotoxic T lymphocyte (CTL) plays a key role in the process of clearing HBV infection. In patients with acute HBV infection (AHB) who successfully clear the virus, HBV-specific CD8+ CTL response is strong, polyclonal and multispecific, while in chronic HBV infection (CHB) it is relatively weak, and compared with chronic HBV infection, CD4+ T cell proliferation is more significant in AHB patients. The glycoprotein encoded by MHC class I molecules can bind to viral polypeptides and present them to CD8+ CTL, which in turn induces the lysis and apoptosis of infected hepatocytes. The main function of MHC class II molecules is to bind to peptide antigens and present them to CD4+ T cells. In different populations, multiple studies have found that the polymorphism of HLA genes is closely related to HBV infection. However, whether there is a relationship between the polymorphism of HLA genes and OBI infection has not been reported.

[0005] The information in the background section is merely intended to illustrate the general background of the invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art. Summary of the Invention

[0006] To address at least some of the technical problems in the prior art, this invention provides a method, system, and application for constructing a model to assess the risk of occult HBV infection in subjects. Specifically, this invention includes the following.

[0007] A first aspect of the invention provides the use of a reagent in the preparation of a diagnostic agent for detecting occult HBV infection, said reagent comprising reagents for displaying information on the following SNP sites in HLA:

[0008] rs1136700, rs9256983, rs9260155, rs9260156, rs9260191, rs1050654, rs4997052, rs1050570, rs1050556, rs2596492, rs713032, rs1142324, rs1064944, rs1140313, rs707954, rs9269939, rs17886882, rs17884070, rs17884945, rs707957, rs41556512, rs138849995, rs1135085, and rs12722477.

[0009] In some embodiments, according to the use described in the first aspect, the reagent comprises primers and / or probes, or the reagent comprises sequencing reagents.

[0010] In some implementations, according to the use described in the first aspect, the SNP site information is as follows:

[0011] The rs1136700 locus is located at position 29943309 on human chromosome 6, with a base sequence of T / C; the rs9256983 locus is located at position 29943451 on human chromosome 6, with a base sequence of A / T,C,G; the rs9260155 locus is located at position 29943462 on human chromosome 6, with a base sequence of T / C; the rs9260156 locus is located at position 29943463 on human chromosome 6, with a base sequence of T / A,G; the rs9260191 locus is located at position 29944621 on human chromosome 6, with a base sequence of A / G; and the rs1050654 locus is located at position 31356323 on human chromosome 6, with a base sequence of G / T. The rs4997052 locus is located at position 31356367 on human chromosome 6, with a base sequence of T / G / A; the rs1050570 locus is located at position 31356772 on human chromosome 6, with a base sequence of T / C; the rs1050556 locus is located at position 31356809 on human chromosome 6, with a base sequence of C / T; the rs2596492 locus is located at position 31356934 on human chromosome 6, with a base sequence of A / G / C; the rs713032 locus is located at position 31271273 on human chromosome 6, with a base sequence of G / A / T; the rs1142324 locus is located at position 32641430 on human chromosome 6, with a base sequence of C / T. The rs1064944 locus is located at position 32641522 on human chromosome 6, with bases A / C / G; the rs1140313 locus is located at position 32664923 on human chromosome 6, with bases T / A; the rs707954 locus is located at position 32581724 on human chromosome 6, with bases A / C / T; the rs9269939 locus is located at position 32584101 on human chromosome 6, with bases A / G; the rs17886882 locus is located at position 32584171 on human chromosome 6, with bases G / T / C / A; and the rs17884070 locus is located at position 32584183 on human chromosome 6, with bases T / G. C; The rs17884945 locus is located at position 32584252 on human chromosome 6, with bases A / T; the rs707957 locus is located at position 32584282 on human chromosome 6, with bases G / A / T; the rs41556512 locus is located at position 3720387 on human chr6_GL000255v2_alt, with bases C / A; the rs138849995 locus is located at position 3720556 on human chr6_GL000255v2_alt, with bases G / A; the rs1135085 locus is located at position 3853092 on human chr6_GL000256v2_alt, with bases C / T;The rs12722477 locus is located at position 29828599 on human chromosome 6, and its base sequence is C / A.

[0012] A second aspect of the present invention provides a model construction method for assessing the risk of occult HBV infection in subjects, comprising the following steps:

[0013] (1) Obtain the SNP sites of each HLA gene of the subject to obtain the SNP site set, wherein the subject includes subjects with occult HBV infection and subjects with dominant HBV infection;

[0014] (2) Screening is performed among multiple associated SNP sites based on the linkage disequilibrium and the influence of variation of the SNP sites, and individual SNPs are screened based on the frequency of minor alleles to obtain a set of candidate SNP sites.

[0015] (3) In the candidate SNP locus set, frequency features are used to encode the SNP genotype features of each SNP locus, and the proportion of occult HBV infected subjects in each SNP genotype is used as the feature value of this SNP genotype.

[0016] (4) Calculate the correlation coefficient between the eigenvalues ​​of each SNP genotype and the classification label, and select eigenvalues ​​based on the correlation coefficients. Next, calculate the correlation coefficients between eigenvalues ​​within a gene, and select the eigenvalue set with the highest correlation coefficients to the classification label; and

[0017] (5) Based on the feature set, a classifier model is trained in the sample to obtain a model of the risk of occult HBV infection in the subject.

[0018] In some implementations, according to the model construction method for assessing the risk of occult HBV infection in subjects as described in the second aspect, the SNP sites of each HLA gene obtained from the subjects are quality controlled to obtain a set of SNP sites, wherein the quality control includes selecting SNP sites that simultaneously meet the following conditions:

[0019] 1) The SNP site coverage depth of the training set samples is greater than or equal to 5;

[0020] 2) The average coverage depth of SNP sites is greater than or equal to 10;

[0021] 3) The deletion rate of HLA-DRB3 / 4 / 5 gene variant sites is less than 0.5, and the deletion rate of other gene variant sites is less than 0.001;

[0022] 4) The number of occurrences of the second allele in the training set samples is greater than 3.

[0023] In some implementations, the model construction method for assessing the risk of occult HBV infection in subjects according to the second aspect includes screening among multiple associated SNP sites based on linkage disequilibrium and the influence of variation of the SNP sites, wherein when r 2 When the metric is above 0.8, one of the variables is retained in the order of high > medium > low impact of variation.

[0024] In some implementations, according to the model construction method for assessing the risk of occult HBV infection in subjects as described in the second aspect, when calculating the eigenvalue of an SNP genotype, if the SNP genotype contains alleles other than the major and minor alleles, and the frequency of the other alleles is less than 0.1, the eigenvalue of the SNP genotype is assigned a value of 0.5; if an SNP genotype not found in the training samples is detected in the test sample, its eigenvalue is assigned a value of 0.5.

[0025] In some implementations, according to the model construction method for assessing the risk of occult HBV infection in subjects as described in the second aspect, the correlation coefficient between the eigenvalues ​​of each SNP genotype and the classification label is calculated, and the eigenvalues ​​are selected based on the correlation coefficients, including:

[0026] When the impact level of the SNP is high or medium, the threshold for the correlation coefficient is 0.15, that is, the characteristic value of the SNP genotype with a correlation coefficient greater than 0.15 is selected.

[0027] When the impact level of the SNP is low, the threshold for the correlation coefficient is 0.2, that is, the characteristic value of the SNP genotype with a correlation coefficient greater than 0.2 is selected.

[0028] A third aspect of the present invention provides an apparatus for assessing the risk of occult HBV infection in a subject, comprising:

[0029] The data acquisition unit is configured to obtain a set of SNP sites for each HLA gene of the subjects, wherein the subjects include subjects with occult HBV infection and subjects with dominant HBV infection; and

[0030] A data processing unit configured to perform the method described in the second aspect.

[0031] A fourth aspect of the invention provides another device for assessing the risk of occult HBV infection in a subject, comprising a memory, a processor, and a program stored in the memory and capable of running on the processor, the program being configured to implement the method according to the second aspect.

[0032] This invention provides an OBI genetic risk assessment model based on HLA regional genetic variations. In practical applications, based on the OBI-related HLA molecular markers and genetic risk assessment results provided by this invention, it can determine whether an individual is prone to dominant or latent HBV infection when chronic HBV infection occurs, thereby strengthening clinical and blood donation screening of high-risk populations.

[0033] In an exemplary embodiment, this invention uses SNP sites in 24 HLA gene regions as molecular markers for OBI and employs them in model construction to assess the genetic risk of OBI. Experimental results show that the model assessment index AUC of this invention is 0.86, exhibiting high specificity and sensitivity. Among the 21 low-risk individuals with a risk coefficient below 0.528, 17 were HBV dominantly infected, achieving an accuracy of 80.95%. Among the 20 high-risk individuals with a risk coefficient above 0.528, 17 were OBI infected, achieving an accuracy of 85%. Therefore, the OBI genetic risk assessment model of this embodiment possesses high specificity, high sensitivity, and high accuracy, providing a more comprehensive, accurate, and individualized scientific basis for OBI risk assessment. Attached Figure Description

[0034] Figure 1 The application uses 1000 cross-validation ROC curves for training samples;

[0035] Figure 2 Genotypic heatmap of independently recruited samples. Detailed Implementation

[0036] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.

[0037] It should be understood that the terminology used in this invention is merely for describing particular embodiments and is not intended to limit the invention. Furthermore, with respect to numerical ranges in this invention, it should be understood that the upper and lower limits of the range and each intermediate value between them are specifically disclosed. Any stated value or intermediate value within a stated range, as well as each smaller range between any other stated value or intermediate value within said range, are also included in this invention. The upper and lower limits of these smaller ranges may be independently included or excluded from the range.

[0038] Unless otherwise stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. While only preferred methods and materials have been described herein, any methods and materials similar or equivalent to those described herein may be used in the implementation or testing of this invention. All references to this specification are incorporated by way of citation to disclose and describe methods and / or materials associated with those references. In the event of any conflict with any incorporated reference, the content of this specification shall prevail.

[0039] [use]

[0040] One aspect of the present invention provides the use of the reagent in the preparation of a diagnostic agent for detecting occult HBV infection. The reagents described in this invention include reagents for displaying the following SNP loci information in HLA: rs1136700, rs9256983, rs9260155, rs9260156, rs9260191, rs1050654, rs4997052, rs1050570, rs1050556, rs2596492, rs713032, rs1142324, rs1064944, rs1140313, rs707954, rs9269939, rs17886882, rs17884070, rs17884945, rs707957, rs41556512, rs138849995, rs1135085, and rs12722477. The 24 SNP sites mentioned above can also be referred to as diagnostic biomarkers for occult HBV infection, or simply as the biomarkers of this invention. The biomarkers of this invention are obtained through screening using the specific methods described in this invention.

[0041] The reagents used in this invention are not limited to any particular type, as long as they can display information about the markers of this invention. Exemplary examples include, but are not limited to, sequencing reagents, primers, and / or probes.

[0042] The primers of this invention typically comprise a first oligonucleotide and a second oligonucleotide, forming a primer pair suitable for PCR amplification. For this purpose, the length of the corresponding target sequence amplified by the first and second oligonucleotides is generally 25-350 nt, preferably 30-300 nt, more preferably 40-250 nt, and even more preferably 50-200 nt. The corresponding target sequence contains at least one, preferably two or more, markers of this invention.

[0043] The probe of this invention is an oligonucleotide that specifically binds to a target sequence and contains a fluorescent group. Its length is preferably suitable for hybridization with complementary DNA to provide a stable hybrid. Typically, the probe length is 10-50 nt, preferably 15-30 nt, such as 20 nt, 25 nt, etc. It should be noted that although the probe is preferably perfectly complementary to the DNA it hybridizes with, in some cases the probe may not be perfectly complementary.

[0044] In this invention, the fluorescent group in the probe is preferably labeled on nucleotide residues, and preferably, at least two of the fluorophore-labeled nucleotide residues are separated by at least two unlabeled nucleotide residues. Typically, 2, 3, 4, 5, or 6 unlabeled nucleotide residues separate the labeled nucleotide residues. Particularly preferably, two unlabeled residues separate the labeled residues. Even more preferably, all labeled nucleotide probes are separated by at least two unlabeled nucleotide residues. Preferably, the fluorescent groups are separated to avoid direct "contact" quenching between the fluorescent groups. Contact quenching occurs due to physical contact between fluorescent groups, and physical separation is typically used to avoid contact quenching, i.e., there are at least two nucleotide residues between the fluorophore-labeled bases.

[0045] In the probes of this invention, the nucleotide residues are often derived from natural nucleosides A, C, G, and T. However, nucleotide analogs may also be used at one or more sites on the probes of this invention. Such nucleotide analogs are modified, for example, at the base moiety and / or sugar moiety and / or phosphate bond. Base modifications (e.g., propynyl dU (dT analogs) and 2-amino dA) generally alter hybridization properties and make the application of oligonucleotides with fewer than 15 nucleotide residues more attractive. For oligonucleotides containing propynyl dU, their length is approximately 10 residues and varies depending on the desired melting temperature with the target sequence.

[0046] Examples of fluorescent groups in the probes of this invention include, but are not limited to, fluorescein-based fluorophores such as FAM (6-carboxyfluorescein), TET (tetrachlorofluorescein), and HEX (hexachlorofluorescein); rhodamine-based fluorophores such as ROX (6-carboxy-X-rhodamine) and TAMRA (6-carboxytetramethylrhodamine); and the Cy dye family, especially Cy3 and Cy5. Other fluoresceins, such as those with different emission spectra, such as NED and JOE, may also be used. Other fluorescent groups may be used, such as those from the Alexa, Atto, Dyomics, Dyomics Megastokes, and Thilyte dye families.

[0047] The reagents of this invention are preferably provided in the form of a kit. This kit typically includes primers, probes, and other reagents. In addition, the kit may include precautions related to the regulation of the manufacture, use, or sale of the diagnostic kit, as prescribed by government agencies. Furthermore, the kit may also include detailed instructions for use, storage, and troubleshooting. The kit may optionally be housed in a suitable device, preferably for high-throughput robotic operation.

[0048] In some embodiments, the components of the kit of the present invention (e.g., primer and probe combinations) are provided as dry powder. When the reagents and / or components are provided as dry powder, the powder can be restored to its original state by adding a suitable solvent. It is contemplated that the solvent may also be provided in another container. The container typically includes at least one vial, test tube, flask, bottle, syringe, and / or other container means in which the solvent is optionally placed in equal portions. The kit may also include means for containing a second container of sterile, pharmaceutically acceptable buffers and / or other solvents.

[0049] In some embodiments, the components of the kit of the present invention may be provided in solution form, such as an aqueous solution. When present in aqueous solution form, the concentration or content of these components can be readily determined by those skilled in the art according to different needs. For example, for storage purposes, the concentration of oligonucleotides may be higher, and when in working condition or for use, the concentration can be reduced to the working concentration, for example, by diluting the higher concentration solution.

[0050] The kit of the present invention may further contain other reagents or components. For example, DNA polymerase required for PCR, various dNTPs, and ions such as Mg²⁺. 2+ These other reagents or components are known to those skilled in the art and are readily available from publicly available publications such as Cold Spring Harbor's *Molecular Cloning: A Laboratory Manual*, fourth edition.

[0051] This invention provides a model construction method for assessing the risk of occult HBV infection in subjects, comprising the following steps:

[0052] (1) Obtain the SNP sites of each HLA gene of the subject to obtain the SNP site set, wherein the subject includes subjects with occult HBV infection and subjects with dominant HBV infection;

[0053] (2) Screening is performed among multiple associated SNP sites based on the linkage disequilibrium and the influence of variation of the SNP sites, and individual SNPs are screened based on the frequency of minor alleles to obtain a set of candidate SNP sites.

[0054] (3) In the candidate SNP locus set, frequency features are used to encode the SNP genotype features of each SNP locus, and the proportion of occult HBV infected subjects in each SNP genotype is used as the feature value of this SNP genotype.

[0055] (4) Calculate the correlation coefficient between the eigenvalues ​​of each SNP genotype and the classification label, and select eigenvalues ​​based on the correlation coefficients. Next, calculate the correlation coefficients between eigenvalues ​​within a gene, and select the eigenvalue set with the highest correlation coefficients to the classification label; and

[0056] (5) Based on the feature set, a classifier model is trained in the sample to obtain a model of the risk of occult HBV infection in the subject.

[0057] Step (1):

[0058] Step (1) of this invention involves obtaining the SNP loci of each HLA gene in the subject to obtain the SNP locus set. Here, the subject is usually multiple individuals, and multiple subjects typically form a sample set. The number of individuals in the sample set is, for example, 10 or more, 20 or more, 30 or more, 50 or more, 80 or more, 100 or more, 200 or more, etc. In some embodiments, the subjects in the sample set include both occult HBV-infected subjects and overt HBV-infected subjects.

[0059] In some embodiments, the SNP locus set of the present invention is a set after screening or quality control to exclude low-quality SNP loci. The quality control criteria include, but are not limited to:

[0060] 1) The SNP site coverage depth of the training set samples is greater than or equal to 5;

[0061] 2) The average coverage depth of SNP sites is greater than or equal to 10;

[0062] 3) The deletion rate of HLA-DRB3 / 4 / 5 gene variant sites is less than 0.5, and the deletion rate of other gene variant sites is less than 0.001;

[0063] 4) The number of occurrences of the second allele in the training set samples is greater than 3.

[0064] Step (2):

[0065] Step (2) of the present invention is a step of obtaining a set of candidate SNP sites, which includes a first screening among associated SNP sites and a second screening for a single SNP site.

[0066] In some implementations, the first screening is based on the linkage disequilibrium and variation effect of the SNP loci. Preferably, this includes the step of calculating the linkage disequilibrium and variation effect of the SNP loci. Linkage disequilibrium refers to the degree of non-random association between alleles at different loci in the genome. Linkage disequilibrium is typically expressed as r... 2 The statistics are quantified. 2 The value range is from 0 to 1. Where, r 2 =0 indicates that the sites are completely independent, r 2 =1 indicates that there is a complete association between the sites. 2 The closer the value is to 1, the stronger the linkage. This invention selects a specific r... 2 The value serves as a threshold for whether a chain is established. For example, when r... 2 A value above 0.8 indicates linkage, requiring the selection of one of the loci. Selection can be based on the degree of variation.

[0067] In this invention, the degree of influence of variation can be confirmed or calculated according to methods known in the art. Exemplarily, the degree of influence of variation therefore includes whether it affects the structure and function of a protein, whether it regulates gene expression, its correlation with disease susceptibility, drug response variability, its impact on physiological characteristics, and genetic linkage and population genetics, etc. In this invention, affecting the structure and function of a protein means that the SNP occurs in the gene region encoding the protein, causing changes in the encoded amino acids, thereby affecting the protein's structure, stability, activity, or interaction with other molecules, and thus affecting the corresponding physiological function. In this invention, regulating gene expression means that the SNP is located in the regulatory region of a gene, such as a promoter or enhancer, affecting the transcription and expression levels of the gene. This may lead to an increase or decrease in gene expression, thereby affecting cell function and phenotype. In this invention, being correlated with disease susceptibility means that certain SNPs may increase an individual's susceptibility or resistance to a specific disease. In this invention, drug response variability means that certain SNPs can affect an individual's response to a drug, including the drug's efficacy and adverse reactions. This is because SNPs may lead to differences in drug-metabolizing enzymes, drug targets, etc., thereby affecting the metabolism and action of the drug in the body. This invention addresses the impact on physiological characteristics, referring to the potential correlation between certain SNPs and individual physiological traits such as height, weight, and skin color. While the influence of a single SNP on these physiological characteristics is typically small, the cumulative effect of multiple SNPs can lead to significant individual differences. In this invention, genetic linkage and population genetics refer to the crucial role of SNPs in population genetics research, enabling the tracking of human migration and evolution, as well as the study of genetic structure and diversity within populations.

[0068] In some implementations, the second screening is a screening of individual SNP loci based on minor allele frequency. Minor allele frequency (MAF) refers to the frequency of a low-frequency allele at a given locus (such as an SNP locus) within a specific population. Minor allele frequency can be calculated using the following formula: Minor allele frequency = (Total copy number of the minor allele) ÷ (Total copy number of all alleles at that locus). For example, in a population of 100 individuals (200 chromosomes in total), if the alleles of a certain SNP are A and T, where A occurs 150 times and T occurs 50 times, then T is the minor allele, and its MAF is 50 / 200 = 0.25 (i.e., 25%). To ensure the accuracy of the assessment, this invention requires selecting SNP loci with high minor allele frequencies. For example, the screening threshold for minor allele frequency is, for instance, 0.2, meaning that SNP loci with a minor allele frequency greater than 0.2 are retained, while loci with excessively low minor allele frequencies are ignored because they would negatively impact the assessment results.

[0069] Step (3):

[0070] Step (3) of the present invention is a feature encoding step, which includes encoding the SNP genotype features of each SNP site in the candidate SNP site set using frequency features, and using the proportion of occult HBV infected subjects in each SNP genotype as the feature value of this SNP genotype.

[0071] For each SNP in the candidate SNP locus set, feature encoding is performed, i.e., a value is assigned to each SNP locus as a feature value. For example, in a sample, the proportion of individuals with occult HBV infection in a particular SNP genotype out of all individuals in the sample is analyzed as the feature value of the SNP genotype. For instance, for a genotype of AA, AT, and TT, if the number of individuals with the AA genotype is 800, and the number of individuals with occult HBV infection is 100, then the feature value of this AA genotype SNP is 80 / 200 = 0.4.

[0072] In certain specific embodiments, the present invention includes assigning values ​​to genotypes at specific SNP loci. For example, when calculating the eigenvalue of an SNP genotype, if the SNP genotype contains alleles other than the major and minor alleles, and the frequency of these other alleles is less than 0.1, then the eigenvalue of that SNP genotype is assigned a value of 0.5. In some embodiments, if an SNP genotype not present in the training samples is detected in the test samples, then its eigenvalue is assigned a value of 0.5.

[0073] Step (4):

[0074] Step (4) of the present invention is a step of screening feature values ​​and constructing a feature value set, which includes calculating the correlation coefficient between the feature value of each SNP genotype and the classification label, and selecting feature values ​​according to the correlation coefficient. Next, the correlation coefficient between feature values ​​within a gene is calculated on a gene-by-gene basis, and the feature with the highest correlation coefficient with the classification label is selected to form a feature value set.

[0075] In this invention, the correlation coefficient can be calculated using any method within the art, and is not particularly limited herein. After the correlation coefficient is calculated, it is typically screened according to different scenarios (e.g., based on influence level). For example, when the influence level of a SNP is high or medium, the threshold for the correlation coefficient between the feature value and the classification label in the feature value set is generally small, for example, 0.15. That is, if the correlation coefficient is greater than this threshold, the corresponding SNP is included in the feature value set; otherwise, it is excluded. Conversely, when the influence level of a SNP is low, the threshold for the selected correlation coefficient is generally high, for example, 0.2. That is, only when the correlation coefficient is greater than this threshold is the SNP included in the feature value set.

[0076] In a preferred implementation, the method further includes calculating the correlation coefficient between the feature values ​​corresponding to each SNP site within a gene, for example, using HLA-A or HLA-B as the unit. If the correlation coefficient is greater than a certain threshold, it indicates that the SNP sites are highly correlated. In this case, the feature value with the highest correlation coefficient with the classification label is retained, and the other feature values ​​are excluded.

[0077] Step (5):

[0078] Step (5) of the present invention is a model building step, which includes training a classifier model in the sample based on the obtained feature set to obtain a model of the risk of occult HBV infection in the subject.

[0079] In this invention, any known classifier can be used when building the model, such as an ensemble logistic regression classifier.

[0080] In another aspect, the present invention provides an apparatus for assessing the risk of occult HBV infection in a subject, also referred to as a first apparatus, comprising:

[0081] The data acquisition unit is used to acquire the SNP sites of each HLA gene of the subject to obtain the SNP site set, wherein the subject includes subjects with occult HBV infection and subjects with dominant HBV infection;

[0082] A data processing unit is configured to perform the method described in the first aspect of the present invention.

[0083] In some implementations, the data processing unit performs the following processing:

[0084] The candidate SNP loci are selected by screening multiple associated SNP loci based on linkage disequilibrium and variation influence, and by screening individual SNPs based on the frequency of minor alleles.

[0085] In the candidate SNP locus set, frequency features were used to encode the SNP genotype features of each SNP locus, and the proportion of subjects with occult HBV infection in each SNP genotype was used as the feature value of this SNP genotype.

[0086] Calculate the correlation coefficient between the eigenvalues ​​of each SNP genotype and the classification label, and select eigenvalues ​​based on the correlation coefficients. Next, on a gene-by-gene basis, calculate the correlation coefficients between eigenvalues ​​within a gene, and select the eigenvalue set with the highest correlation coefficients to the classification label; and

[0087] Based on the feature set, a classifier model is trained in the sample to obtain a model of the risk of latent HBV infection in the subject.

[0088] Genotype feature encoding method: The proportion of OBI samples in a SNP genotype is used as the feature value of that SNP genotype. That is, the higher the proportion of OBI-infected individuals (case group: 1) in samples detecting this genotype, the closer the feature value is to 1, and vice versa. 0.5 represents that this genotype has no effect on HBV infection status or the effect is uncertain. If a genotype not found in the training samples is detected in the test samples, the feature value is assigned to 0.5. Considering the randomness of the OBI proportion due to the low frequency of genotype occurrences in the training samples (e.g., a genotype appearing only once in extreme cases), the phenotype of the corresponding sample is highly likely to be accidental), if the genotype contains alleles other than the major and minor alleles, and the genotype frequency is less than 0.1, the feature value is assigned to 0.5.

[0089] In this invention, the data acquisition unit is capable of acquiring sequencing data, for example, during or after sequencing. Optionally, the data acquisition unit can be communicatively connected to a sequencer to acquire data during the sequencing process. Optionally, the data acquisition unit can be communicatively connected to a database to acquire sequencing data from the database. The database may be known public data.

[0090] In another aspect, the present invention provides an apparatus for assessing the risk of occult HBV infection in a subject, referred to herein as a second apparatus, which includes a memory, a processor, and a program stored on the memory and capable of running the method described in the first aspect of the present invention on the processor.

[0091] In some embodiments, the second device of the present invention is a computer device, which may be a server. The computer device includes a processor, memory, and a network interface connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores sequencing data to be processed. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the method described in the first aspect of the present invention.

[0092] In addition to the aforementioned units, devices, and components, the first or second device of the present invention may further include other units or devices, exemplarily including a display unit or device, etc. Each unit in the first or second device of the present invention may be implemented wholly or partially through software, hardware, or a combination thereof. Each unit may be embedded in or independent of any computer device in hardware form, or may be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned units.

[0093] Exemplarily, the first or second device of the present invention further includes one or more output devices that can be connected to, for example, a computer system. Examples of output devices include displays, printers, communication devices such as modems, and audio outputs. Optionally, it further includes one or more input devices that can also be connected to the computer system. Examples of input devices include keyboards, mice, pens and writing tablets, communication devices, and data input devices such as sensors.

[0094] Example

[0095] I. Collection of case and control samples

[0096] Based on informed consent, blood samples were collected from voluntary blood donors who were initially positive for HBsAg or HBV DNA from blood collection and supply institutions in 14 provinces of China. After the test results were verified, HBV-related biomarkers were tested multiple times over the next two years for screening. Samples that met the following criteria were included in the training sample set.

[0097] Case group (OBI): At least two follow-up results over 2 years (with an interval of no less than 6 months), with each test result being negative for HBsAg and positive for HBV DNA (viral load <200 IU / mL). If HBV DNA is found to be sometimes positive and sometimes negative in multiple follow-ups, this group is also included (because OBI-infected individuals have the characteristic of fluctuating HBV DNA detection).

[0098] Control group (HBV symptomatic infection): At least two follow-up results over 2 years (with an interval of no less than 6 months), each test result was positive for both HBsAg and HBV DNA (clinically recognized).

[0099] Ultimately, 53 samples were collected from both the case group and the control group.

[0100] II. Whole exome sequencing of case and control samples

[0101] Genomic DNA was extracted from blood, whole-exome data were obtained to build a genome sequencing library, and the sequencing library was targeted for enrichment using whole-exome capture probes. High-throughput sequencing was performed with an average sequencing depth of 100X.

[0102] III. HLA gene region SNP detection and pretreatment

[0103] HLA gene region SNPs were identified from 106 training set samples using known methods. Exemplary methods include, for example, those described in Chinese Patent Application No. 2025105248340. Next, the obtained SNPs underwent quality control, and retained SNPs were required to simultaneously meet the following requirements:

[0104] 1) The SNP site coverage depth of the training set samples is greater than or equal to 5;

[0105] 2) The average coverage depth of SNP sites is greater than or equal to 10;

[0106] 3) The deletion rate of HLA-DRB3 / 4 / 5 gene variant sites is less than 0.5, and the deletion rate of other gene variant sites is less than 0.001;

[0107] 4) The number of minor alleles in the training set samples was greater than 3. SNP sites in 1,1868 HLA gene regions were included in subsequent analysis.

[0108] Based on the r² measure of linkage disequilibrium (LD) among the 1868 HLA SNPs, unique SNPs with an r² value greater than or equal to 0.8 were retained, and one SNP was retained in order of priority based on its impact (high > moderate > low > modifier). 916 SNPs were included in subsequent analyses.

[0109] SNPs with a minor allele frequency (MAF) greater than 0.2 in the training samples were retained. 402 SNPs were included in subsequent analyses.

[0110] IV. Feature Coding

[0111] Considering the polymorphism of SNPs in HLA regions, a frequency encoding strategy is used to encode SNP genotype features to fully reflect the relationship between different genotypes and OBIs. The proportion of OBI samples in an SNP genotype is used as the feature value of that SNP genotype. To ensure reliability, if a genotype contains alleles other than the major and minor alleles and the genotype frequency is less than 0.1, the feature value is assigned to 0.5. In addition, if a genotype not found in the training samples is detected in the test samples, the feature value is assigned to 0.5.

[0112] V. Feature Filtering

[0113] The Pearson correlation coefficient between the feature values ​​and the classification label (case: 1; control: 0) was calculated. Feature values ​​included in the analysis had to meet one of the following conditions: 1) Features with a high or moderate SNP influence level had a correlation coefficient greater than 0.15 with the classification label; 2) Features with a low SNP influence level had a correlation coefficient greater than 0.2 with the classification label. 45 HLA SNP features were included in the subsequent analysis.

[0114] Using each gene in the HLA family (HLA-A, HLA-B, etc.) as a unit, SNP features with a Pearson correlation coefficient greater than 0.8 were retained, and those features had the highest correlation coefficient with the classification label. 41 HLA SNP features were included in subsequent analyses.

[0115] L2 regularized logistic regression was used to recursively eliminate low-importance features from the 41 HLA SNP features. The logistic regression coefficient was used as the metric for feature importance. A feature was removed if its removal resulted in no change or an increase in the training set AUC; otherwise, it was retained. This process continued until removing any feature caused a decrease in the training set AUC. The training set AUC was calculated using a leave-one-out method. In special cases, such as when multiple features have similar importance scores, features with more gene features were prioritized for removal (e.g., if HLA-A has 10 SNP features and HLA-B has 5 SNP features, then HLA-A was removed). The values ​​of 24 HLA SNP features were incorporated into the model (Tables 1 and 2).

[0116] Table 1. 24 HLA SNP loci associated with OBI genetic risk

[0117]

[0118]

[0119] Note: a. Genomic location reference is GRCH38.p14; b. Correlation coefficient between features and classification.

[0120] Table 2. SNP Genotype Characteristics Comparison Table

[0121]

[0122]

[0123]

[0124] VI. Establishment of the OBI Genetic Risk Assessment Model

[0125] Based on the 24 SNP features of the training samples, an ensemble logistic regression classifier was trained. To avoid overfitting, established methods were used in this process (e.g., see Teschendorff, AE, Avoiding common pitfalls in machine learning omic data science. Nat Mater, 2019, 18(5): p.422-427.). Specifically, the training samples were randomly divided into two groups in a 1:1 ratio: one group was used for model fitting, and the other group was used for model validation. The model fitting set was used to train the L2 regularized logistic regression classifier, and the model score of the model validation set was calculated. This process was repeated 1000 times by randomly splitting the training set, and the model score of each sample was averaged to produce the final result. Therefore, the risk coefficient of the final OBI genetic risk assessment model is the average of the 1000 L2 regularized logistic regression classifiers built on different splits of the training set.

[0126] The approximation index is calculated based on the scoring results of the training samples. The risk coefficient corresponding to the maximum approximation index (0.528 in this embodiment) is used as the risk judgment threshold of the model. When the risk coefficient is lower than 0.528, the individual is judged as low-risk; when the risk coefficient is higher than 0.528, the individual is judged as high-risk.

[0127] VII. Verification

[0128] The sample was divided into model fitting data and model validation data in a 1:1 ratio, and 1000 cross-validations were performed. The receiver operating characteristic (ROC) curve (an ROC curve is a curve with the true positive rate on the y-axis and the false positive rate on the x-axis) is shown below. Figure 1 As shown, the model evaluation index AUC (area under the ROC curve, which can be used to evaluate the overall ability of the model; the larger the AUC value, the higher the model classification accuracy) is 0.86, indicating that the evaluation model of the present invention has the characteristic of high accuracy.

[0129] An additional 41 samples were independently recruited, and their risk was assessed using the aforementioned 24 HLA SNP loci and the established evaluation model, compared with HBV infection status obtained during follow-up. Detailed results regarding the genotypic heatmap information of the samples can be found in [link to heatmap]. Figure 2 The experimental results showed that among the 21 low-risk individuals with a risk coefficient below 0.528, 17 were symptomatic HBV infection, with an accuracy of 80.95%. Among the 20 high-risk individuals with a risk coefficient above 0.528, 17 were OBI infection, with an accuracy of 85%.

[0130] VIII. Advantages of the method of the present invention:

[0131] Traditional GWAS studies only consider the major and minor alleles of SNP loci for differential testing of alleles (2 types) or genotypes (3 types). However, HLA gene regions are highly variable regions of the human genome, especially exons 2 and 3, which can present antigenic peptides to T cells and exhibit extremely high polymorphism. Considering only major / minor alleles will miss other alleles associated with the phenotype, such as the SNP locus rs1049066 located in exon 2 of the DQB1 gene. In the Chinese population, the frequencies of the T / A / C alleles are 0.367 / 0.361 / 0.272, respectively. Traditional GWAS studies would ignore allele C and related genotypes. To effectively utilize alleles other than major / minor alleles, this application introduces a frequency encoding strategy for SNP genotype feature encoding, and then performs feature screening based on feature values ​​after genotype feature encoding. Compared with traditional GWAS methods, the method of this application has superior technical performance.

[0132] Table 3 Comparison of p-values ​​between this application and conventional case-control GWAS analysis

[0133]

[0134] Although the invention has been described with reference to exemplary embodiments, it should be understood that the invention is not limited to the disclosed exemplary embodiments. Various adjustments or changes may be made to the exemplary embodiments described in this specification without departing from the scope or spirit of the invention. The scope of the claims should be interpreted in the broadest possible sense to cover all modifications and equivalent structures and functions.

Claims

1. The use of the reagent in the preparation of a diagnostic agent for detecting occult HBV infection in HBV infection, characterized in that, The reagents include those for displaying information on the following SNP sites in HLA: rs1136700, rs9256983, rs9260155, rs9260156, rs9260191, rs1050654, rs4997052, rs1050570, rs1050556, rs2596492, rs713032, rs1142324, rs1064944, rs1140313, rs707954, rs9269939, rs17886882, rs17884070, rs17884945, rs707957, rs41556512, rs138849995, rs1135085, and rs12722477.

2. The use according to claim 1, characterized in that, The reagents include primers and / or probes, or the reagents include sequencing reagents, wherein the primers include a first oligonucleotide and a second oligonucleotide for amplifying a corresponding target sequence, and the probes are used to specifically bind to a corresponding target sequence, wherein the corresponding target sequence contains at least one of the SNP sites.

3. The use according to claim 1, characterized in that, The SNP site information is as follows: The rs1136700 locus is located at position 29943309 on human chromosome 6, with a base sequence of T / C; the rs9256983 locus is located at position 29943451 on human chromosome 6, with a base sequence of A / T,C,G; the rs9260155 locus is located at position 29943462 on human chromosome 6, with a base sequence of T / C; the rs9260156 locus is located at position 29943463 on human chromosome 6, with a base sequence of T / A,G; the rs9260191 locus is located at position 29944621 on human chromosome 6, with a base sequence of A / G; and the rs1050654 locus is located at position 31356323 on human chromosome 6, with a base sequence of G / T. The rs4997052 locus is located at position 31356367 on human chromosome 6, with a base sequence of T / G / A; the rs1050570 locus is located at position 31356772 on human chromosome 6, with a base sequence of T / C; the rs1050556 locus is located at position 31356809 on human chromosome 6, with a base sequence of C / T; the rs2596492 locus is located at position 31356934 on human chromosome 6, with a base sequence of A / G / C; the rs713032 locus is located at position 31271273 on human chromosome 6, with a base sequence of G / A / T; the rs1142324 locus is located at position 32641430 on human chromosome 6, with a base sequence of C / T. The rs1064944 locus is located at position 32641522 on human chromosome 6, with bases A / C / G; the rs1140313 locus is located at position 32664923 on human chromosome 6, with bases T / A; the rs707954 locus is located at position 32581724 on human chromosome 6, with bases A / C / T; the rs9269939 locus is located at position 32584101 on human chromosome 6, with bases A / G; the rs17886882 locus is located at position 32584171 on human chromosome 6, with bases G / T / C / A; and the rs17884070 locus is located at position 32584183 on human chromosome 6, with bases T / G. C; The rs17884945 locus is located at position 32584252 on human chromosome 6, with bases A / T; the rs707957 locus is located at position 32584282 on human chromosome 6, with bases G / A / T; the rs41556512 locus is located at position 3720387 on human chr6_GL000255v2_alt, with bases C / A; the rs138849995 locus is located at position 3720556 on human chr6_GL000255v2_alt, with bases G / A; the rs1135085 locus is located at position 3853092 on human chr6_GL000256v2_alt, with bases C / T;The rs12722477 locus is located at position 29828599 on human chromosome 6, and its base sequence is C / A.