Hepatitis B virus infection rapid diagnosis system combining machine learning and peripheral serum MALDI-TOF MS
By combining MALDI-TOF MS and machine learning algorithms, the HBV_pred system was constructed, which solves the problems of high cost, complex operation and insufficient accuracy of existing HBV infection detection, and realizes rapid, low cost and high accuracy HBV infection diagnosis.
Patent Information
- Application Number
- CN202511703591.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-03-06
AI Technical Summary
Existing HBV infection detection methods suffer from high costs, complex operation, reliance on human visual recognition of spectral variations, and insufficient accuracy and consistency. Furthermore, MALDI-TOF MS lacks effective application in HBV infection diagnosis.
Using MALDI-TOF MS combined with machine learning algorithms, a diagnostic model for HBV infection was constructed by collecting protein fingerprint profiles of peripheral blood serum from patients and employing LightGBM, DNN, and RF algorithms. The profile data was standardized, binned, and normalized within samples. The learning rate was optimized, stable characteristic peak regions were screened, and the HBV_pred system was constructed.
It enables rapid, accurate, and low-cost diagnosis of HBV infection, with short testing time and low cost, reducing reliance on professional personnel and improving diagnostic accuracy and consistency.
Smart Images

Figure CN121617482A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical technology, specifically relating to a rapid diagnostic system for hepatitis B virus infection that combines machine learning and peripheral serum MALDI-TOF MS. Background Technology
[0002] Hepatitis B virus (HBV) is a hepatotropic DNA virus that can cause acute and chronic liver infections, often leading to liver fibrosis, cirrhosis, and hepatocellular carcinoma. Studies indicate that timely antiviral treatment in the early stages of infection can effectively prevent HBV complications and reduce disease-related mortality. Therefore, only by promoting early HBV screening can patients with potential disease progression risk who require treatment be identified, enabling early antiviral therapy and regular monitoring, thus reducing the economic burden on patients and society as a whole. Based on this situation, developing simple, rapid, accurate, and inexpensive HBV infection detection methods will help to better promote HBV infection screening, improve the early detection rate of infected patients, improve the prognosis of infected patients, and improve the overall health of the Chinese population.
[0003] Studies indicate that HBV primarily exists in the human bloodstream, and serum, as a component of blood, directly reflects infection status, making it the most commonly used sample type for HBV testing. Serum contains classic HBV biomarkers, including serological markers such as hepatitis B surface antigen (HBsAg) and its antibody (anti-HBs), hepatitis B e antigen (HBeAg) and its antibody (anti-HBe), and hepatitis B core antibody (anti-HBc), as well as molecular biological markers such as HBV DNA. Detecting these biomarkers is crucial for assessing HBV infection status, viral replication levels, and disease progression. Generally, HBsAg is the marker of HBV infection and the standard for determining the presence of HBV in the body. It appears 1-12 weeks after HBV infection and can persist for life if the patient does not receive treatment. Anti-HBs is a protective antibody, and a concentration ≥ 10 mIU / ml indicates immunity. HBeAg is an important indicator of viral replication. It appears slightly later than HBsAg after HBV infection and has a good correlation with HBV DNA. Its positivity indicates active viral replication and strong infectivity. Anti-HBe appears after HBeAg becomes negative, indicating reduced HBV replication and infectivity. Anti-HBc IgM appears 2-4 weeks after HBsAg positivity and is a marker of acute and chronic HBV infection activity. Anti-HBc IgG appears later but persists for a long time or even remains positive for life, and is a marker of current or past infection. HBV DNA is a direct marker of viral replication and infectivity, reflecting the level of HBV replication activity and infectivity. It is also the most important indicator for selecting indications and judging the efficacy of antiviral therapy.
[0004] Currently, the detection of serum biomarkers in HBV-infected patients is mainly achieved through enzyme-linked immunosorbent assay (ELISA) and quantitative real-time polymerase chain reaction (qPCR). ELISA detects serological markers such as HBsAg, HBeAg, and anti-HBc in serum, utilizing the principle of specific antigen-antibody binding to determine the presence of specific antigens or antibodies in the blood through a colorimetric reaction. While ELISA is the most commonly used method for clinical HBV infection detection, the quality of the antibodies used is a major factor affecting the accuracy and sensitivity of the test. Therefore, testing laboratories need to be equipped with excellent antibody transportation, storage, and quality control technologies to ensure antibody effectiveness, which significantly increases the cost of testing. On the other hand, qPCR, by quantitatively detecting HBV DNA levels in serum samples, allows for dynamic monitoring of HBV replication and is an ideal choice for assessing patient infectivity and disease progression. This technology uses fluorescent dyes or fluorescently labeled specific probes to quantitatively label and track PCR products, offering high detection sensitivity. However, qPCR detection relies on reagents such as Taq polymerase, which require cryopreservation, and necessitates a highly clean laboratory environment. Therefore, it is difficult to perform in laboratories with limited resources, equipment, and skilled personnel. Thus, developing novel diagnostic methods that are simple to use and inexpensive is of significant practical importance.
[0005] Matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF MS) is a novel bio-soft ionization organic mass spectrometry technique. It represents one of the most significant technological breakthroughs in clinical microbiology identification in recent years and is widely used in clinical hospitals, disease control centers, and customs systems worldwide. In microbial detection, this technology uses high-throughput measurement of the precise mass-to-charge ratio (m / z) of peptides and proteins to directly generate specific biomarker molecular profiles (protein fingerprints) from the entire sample within minutes, enabling pathogen identification. It offers advantages such as simple operation, rapid identification, and low cost. Currently, MALDI-TOF MS technology has also been introduced into liver disease diagnostic research. Specifically, Sun Fei et al. used MALDI-TOF MS technology to analyze proteins in patient serum and established a liver cancer diagnostic model by detecting characteristic serum marker proteins of liver cancer. Liangpunsakul et al. used MALDI-TOF MS technology to compare serum protein profiles of volunteers before and after alcohol consumption, finding that α-fibrinogen can serve as a specific biomarker for alcoholic liver disease. Liu C et al. used MALDI-TOF MS technology to compare serum protein profiles of patients with hepatocellular carcinoma, patients with other liver diseases, and healthy individuals, finding that protein peaks at 4471, 8936, 11670, and 13752 m / z can serve as specific biomarkers for identifying hepatocellular carcinoma. However, currently, there is still a lack of applications for the direct identification of HBV infection using MALDI-TOF MS. In principle, MALDI-TOF MS has the potential to obtain fingerprints of specific HBV proteins or peptides from patient serum, thereby assisting in virus identification and typing, and revealing its genetic variations and evolutionary relationships. Therefore, it has the potential to be developed into a novel detection method for diagnosing HBV infection.
[0006] However, there are still many challenges in developing a diagnostic system for HBV infection based on MALDI-TOF MS. The reasons for the low diagnostic capability of MALDI-TOF MS for HBV infection may be: (1) Incomplete database: At present, the MALDI-TOF MS database mainly focuses on the protein spectrum of microorganisms such as bacteria and fungi, while the protein spectrum data for viruses or serum of patients with viral infection is relatively scarce, especially the specific protein spectrum data for HBV. (2) Limited image feature recognition capability relying on human eyes: The development of traditional MALDI-TOF MS diagnostic systems is based on the human eye to identify and collect protein fingerprint spectrum data. After statistically analyzing each of the identified spectrum peaks, fingerprint peaks with statistical differences are selected to form a diagnostic system. However, due to the differences in interpretation of the same mass spectrometry spectrum by different operators, the accuracy and consistency of the diagnostic results are affected, and the reproducibility of the diagnostic results is also poor. In order to overcome these technical bottlenecks, modern MALDI-TOF MS diagnostic systems are being developed towards intelligent applications based on big data.
[0007] In recent years, researchers have actively introduced machine learning to analyze complex fingerprint data and have successfully built fast and efficient identification mechanisms, thereby significantly improving the recognition performance of feature fingerprints. Machine learning, simply put, empowers computers to learn autonomously based on established algorithms and existing data, thereby constructing models capable of predicting unknown data. The core of machine learning lies in algorithms, which define how computers learn from data. Different data characteristics require different algorithms for summarization and generalization. In the MALDI-TOF MS research field, commonly used algorithms include Random Forest (RF), Deep Neural Network (DNN), and Lightweight Gradient Boosting Machine (LightGBM). Specifically, the RF algorithm is an ensemble learning strategy that significantly improves the accuracy and stability of the model by constructing numerous decision trees and combining their predictions. DNN (Dual Neural Network) is a type of machine learning algorithm inspired by the structure of the human brain. It consists of multiple layers of neurons, each receiving the output of the previous layer as input and calculating its own output through nonlinear transformations and weight adjustments. DNNs continuously adjust network parameters using backpropagation and gradient descent, thus automatically extracting key features from raw data and significantly improving the model's generalization ability. LightGBM, as an efficient implementation of gradient boosting decision trees, minimizes the loss function by iteratively adding new trees, significantly optimizing memory usage and computational efficiency. This algorithm supports multi-threaded and parallel training, making it fast in training and accurate prediction when handling large datasets. It also allows training to stop when performance on the validation set no longer improves, preventing overfitting. The application of different machine learning algorithms in building the MALDI-TOF MS prediction model is affected by various factors, therefore, it is necessary to flexibly select the most suitable algorithm based on the actual situation. Based on this, using a serum protein fingerprint database of HBV-infected patients, this study attempts to construct a diagnostic model using multiple algorithms, which is expected to achieve rapid, accurate, and cost-effective diagnosis of HBV infection, demonstrating broad application prospects. Summary of the Invention
[0008] To overcome the shortcomings of existing technologies, this invention provides an accurate, rapid, simple, and low-cost rapid diagnostic system for HBV infection.
[0009] The first objective of this invention is to provide a method for constructing an HBV infection diagnostic model, comprising the following steps:
[0010] (1) Collect fresh peripheral blood serum from patients and perform HBsAg testing to determine patient grouping;
[0011] (2) Protein fingerprints of peripheral blood serum from patients were collected using MALDI-TOF MS. The fingerprints were classified into two categories based on the sample source. Protein fingerprints of HBV-infected patients were marked as “1”, and protein fingerprints of patients without HBV infection were marked as “0”. All labeled fingerprints were randomly divided into training and testing sets to form a binary classification dataset for subsequent modeling.
[0012] (3) The binary classification dataset is standardized by the R package MALDIquant to collect protein fingerprints, including quality control, variance stabilization, spectrum smoothing, baseline calibration and intensity normalization. After processing, the peak positions and intensities of the protein fingerprints are recorded.
[0013] (4) The standardized binary classification dataset is binned according to the mass-to-charge ratio in the range of 2000-20000 m / z. Different groups are set with binsize of 3, 5, 10 and 15 m / z. Each sample is processed independently.
[0014] (5) Perform intra-sample standardization on the binned data: In each sample, sum the intensity values of all bins, and then divide each bin value by the sum to complete the total normalization and generate a fixed-length proportional feature vector for subsequent modeling;
[0015] (6) Call the Python packages lightgbm, keras, and scikit-learn, and use the LightGBM, DNN, and RF algorithms to perform preliminary model training on the dataset. The optimal modeling conditions are obtained by screening with the F1 Score index: LightGBM algorithm and binsize = 15. Retrain under this combination to obtain a predictive model for HBV infection diagnosis.
[0016] (7) Optimize the learning rate of the prediction model to obtain the optimized diagnostic model;
[0017] (8) Robust feature selection: Fix the hyperparameters of the optimized diagnostic model, perform 100 Bootstrap resampling trainings, calculate the absolute mean of each m / z-bin using SHAP each time, and retain the peak regions where the SHAP value is non-zero in ≥66 iterations; finally determine the stable fingerprint spectrum feature peak regions;
[0018] (9) The final model for HBV infection diagnosis, HBV_pred, is constructed based on the characteristic peak region of the stable fingerprint spectrum.
[0019] Preferably, in step (6), the retraining under this combination to obtain a predictive model for HBV infection diagnosis is as follows: first, the mass spectrometry data is binned according to 15 m / z and normalized within the sample ratio, and then trained using the LightGBM single algorithm. The hyperparameters are locked with the highest AUC and F1 as the criteria, and the obtained model is the predictive model for HBV infection diagnosis.
[0020] Preferably, in step (7), the learning rate optimization of the prediction model specifically involves adjusting the hyperparameter of learning rate to 0.4152.
[0021] Preferably, in step (8), the characteristic peak regions of the stable fingerprint spectrum are: 2015~2030, 2030~2045, 2705~2720, 2735~2750, 3125~3140, 3920~3935, 4115~4130, 4790~4805, 6455~6470, 7565~7580, 7790~8050, 8150~8165 m / z.
[0022] Preferably, in step (9), the final model HBV_pred for HBV infection diagnosis is constructed based on the stable fingerprint spectrum feature peak region as follows: only the intensity value of the stable fingerprint spectrum feature peak region is extracted, in-sample normalization is performed again, and the model is retrained with the same hyperparameters to obtain the final model HBV_pred for HBV infection diagnosis.
[0023] The second objective of this invention is to provide an HBV infection diagnostic model constructed according to the above-described construction method.
[0024] A third objective of this invention is to provide the application of the above-mentioned HBV infection diagnostic model in the identification of HBV infection for non-disease diagnosis and treatment purposes.
[0025] Preferably, the application includes the following steps: using MALDI-TOF MS to collect protein fingerprint profiles from a single serum sample of peripheral blood of the patient to be tested at least 16 times, standardizing the collected profiles by summing the values, calculating the predict value for each sample by calling an infection diagnosis model, determining whether the sample to be tested is infected with HBV based on the predict value, and calculating the reliability of the diagnostic results.
[0026] Preferably, the step of determining whether the sample to be tested is infected with HBV based on the predict value specifically involves: inputting the protein fingerprint of a single serum sample into the infection diagnosis model, and then calculating the predict value; in the HBV infection diagnosis model, when predict ≥ 0.5, the input pattern is considered to indicate that the sample comes from an HBV-infected patient (positive); when predict < 0.5, the input pattern is considered to indicate that the sample comes from an uninfected patient (negative).
[0027] This invention utilizes 226 peripheral blood serum samples from HBV-infected individuals and 196 peripheral blood serum samples from healthy individuals without HBV infection to establish a biological database containing 5424 HBV-infected and 4704 normal MALDI-TOF MS protein fingerprint profiles. Based on this database, we successfully developed a system called HBV_pred using machine learning technology. This system has the ability to comprehensively evaluate the fingerprint results, enabling rapid and accurate diagnosis of HBV infection.
[0028] Compared with existing HBV infection diagnostic systems, the advantages of this invention are:
[0029] (1) Using the HBV fingerprint spectrum obtained by MALDI-TOF MS in this invention, users can obtain the diagnosis result of HBV infection in a single test within seconds. The test cost is low (about RMB 1 for this invention, about RMB 28 for HBsAg test by hospital ELISA, and about RMB 56 for HBV DNA quantitative test), and the efficiency is significantly higher than that of existing traditional test methods (about RMB 1 for this invention, about RMB 2 for ELISA test by hospital, and about RMB 3 for HBV DNA quantitative test).
[0030] (2) When using the present invention to test samples, there is no need for cumbersome and complicated pretreatment steps. The operation is simple and significantly reduces the consumption of reagents and consumables, and greatly reduces the dependence on experienced personnel.
[0031] (3) The machine learning algorithm is used to identify the characteristic fingerprint spectrum of HBV. Compared with visual identification, it can capture the characteristic peaks more sensitively and accurately, thereby improving the accuracy of the system in diagnosing HBV infection. Attached Figure Description
[0032] Figure 1 The F1 score is a preliminary model for HBV infection diagnosis and prediction constructed using different algorithms.
[0033] Figure 2 This refers to the learning rate adjustment of the HBV infection diagnosis prediction model.
[0034] Figure 3It is the AUC of the HBV infection diagnosis prediction model.
[0035] Figure 4 The result is that the SHAP value is non-zero after constructing the model using 100 different random numbers.
[0036] Figure 5 This is an analysis of the contribution of the SHAP value to the diagnostic model of HBV infection. Detailed Implementation
[0037] The following embodiments are further illustrations of the present invention, but not limitations thereof.
[0038] The reagents involved in the following examples are as follows:
[0039] 70% Formic Acid Solution: 700 μL of pure formic acid solution and 300 μL of sterile deionized water, stored at room temperature for later use.
[0040] Standard solution for protein spectrum detection: 10 mg α-cyano-4-hydroxycinnamic acid, 500 μL acetonitrile, 25 μL trifluoroacetic acid, and 475 μL sterile deionized water. After shaking and mixing, store at room temperature away from light.
[0041] Example 1: Acquisition and database construction of peripheral blood serum MALDI-TOF MS profiles
[0042] (1) Establishment of clinical cohorts of HBV-infected and HBV-non-infected patients:
[0043] Between November 2023 and April 2024, fresh peripheral blood serum samples were collected from patients undergoing physical examinations at a tertiary hospital in Guangdong Province. HBV infection status was diagnosed by detecting HBsAg using the ARCHITECT HBV surface antigen qualitative assay kit (Abbott, Ireland). A total of 422 serum samples were collected, of which 226 were HBsAg positive and 196 were HBsAg negative.
[0044] Patient information collection revealed that the median age of patients in the HBV-infected group was 51 years (range 18–82 years), and the median age of patients in the non-HBV-infected group was 50 years (range 18–88 years). There was no statistically significant difference in age between the two groups (P = 0.4038). The male-to-female ratio of patients in the HBV-infected group was 141 / 85, and the male-to-female ratio of patients in the non-HBV-infected group was 120 / 76. There was no statistically significant difference in gender composition between the two groups (P = 0.8059).
[0045] (2) MALDI-TOF MS detection:
[0046] 1 μL of peripheral blood serum sample was added to the target plate and allowed to dry. Then, 1 μL of 70% formic acid solution was added, followed by further drying. Finally, 1 μL of standard protein chromatography solution was added, and the sample was allowed to dry completely before analysis. MALDI-TOF MS was performed on a Microflex instrument. ® The experiment was completed on the smart (Bruker, Leipzig, Germany) platform: The experimental instruments were calibrated regularly using the Bacterial Test Standard (Bruker) according to the specifications; after calibration, among the instruments that meet the specifications, a plate-bombardment strategy of bombarding the target with a 200Hz smartbeam solid-state laser 240 times was selected, and the spectrum in the range of 2000~20000 m / z was collected in cation linear mode to form a serum protein fingerprint spectrum, and a file package containing the spectrum file (.fid) was saved.
[0047] (3) Construction of HBV infection protein fingerprint database:
[0048] Protein fingerprinting was performed on 422 fresh serum samples collected in this study using the method described above. Eight target sites were coated on each sample, and three fingerprint images were collected from each target site. A total of 10,128 protein fingerprint images were collected using this method, and the collected protein fingerprint dataset is shown in Table 1.
[0049] Table 1. Statistics on protein fingerprint data collection
[0050] Example 2: Construction of HBV_pred, a machine learning-based HBV infection diagnostic system
[0051] This invention constructs the HBV_pred HBV infection diagnostic system based on machine learning by comparing the differences in protein fingerprint profiles of HBV and non-HBV. The system consists of three main steps: data preprocessing, diagnostic model development, and diagnostic system establishment.
[0052] (1) Preprocessing of serum protein fingerprint data:
[0053] All raw data from serum protein fingerprints will undergo a two-step preprocessing process: standardization and binning.
[0054] ①Standardization:
[0055] The collected serum protein fingerprints were standardized using the R package MALDIquant. The main preprocessing steps included: <1> Quality control: Test all mass spectrometry data to ensure they contain the same number of data points and are not empty, exclude blank mass spectrometers and calibration mass spectrometers, and ensure that the data quality of different sites reaches a similar level. <2> Variance stabilization: The mass spectrum intensity values are transformed by square root to overcome the potential dependence of variance on the mean. <3> Spectral smoothing: The Savitzky-Golay-Filter method is used to smooth the peak values. <4> Baseline calibration: The peak baseline is calibrated through 20 iterations of the SNIP algorithm. <5> Intensity normalization: Total ion current calibration was used to balance the peak intensity values, and the peak size range was trimmed to 2000~20000 m / z. The processed data yielded a .csv file containing the peak positions and intensities of the serum protein fingerprint.
[0056] ②Separate boxes:
[0057] Binning refers to summing the intensity of peaks within a fixed range in a fingerprint spectrum based on different bin sizes. This invention uses Python to divide protein peaks in the 2000–20000 m / z range into non-overlapping bins of equal size. To select the optimal classification effect, we set different bin sizes of 3, 5, 10, and 15 m / z, and standardized the summation values of these bins for each individual sample for downstream machine learning tasks.
[0058] The mass spectrometry data, after being binned and standardized from the raw data, will be used for subsequent model building.
[0059] (2) Development of a diagnostic model for HBV infection protein fingerprinting:
[0060] The development process of the HBV infection protein fingerprint diagnostic model in this invention is divided into three steps: dataset construction, model evaluation, and model optimization.
[0061] ①Database construction:
[0062] Based on whether the protein fingerprint data originated from HBV, the classification information of the fingerprint data was binarized. Serum protein fingerprint data from HBV-infected patients were marked as "1" (positive), while serum protein fingerprint data from non-HBV-infected patients were marked as "0" (negative). All data in the fingerprint database were randomly divided into training datasets (60%) and test datasets (40%). Simultaneously, it was ensured that different fingerprint data from the same sample did not partially exist in both the training and test sets to avoid information leakage.
[0063] ② Model evaluation:
[0064] This invention uses three classification algorithms with different computational capabilities to build a predictive model for diagnosing HBV infection: Random Forest (RF), Deep Neural Network (DNN), and Light Gradient Boosting Machine (LightGBM). All three algorithms are based on open-source Python packages: lightgbm, keras, and scikit-learn.
[0065] The area under the receiver operator characteristic curve (AUC) is an evaluation metric for the predictive classification performance of this model; the closer the AUC is to 1, the better the classification performance. The F1 score defines the accuracy of the predicted data; the closer the F1 score is to 1, the better the classification performance. We first evaluated the diagnostic efficacy of models built using different algorithms using AUC and F1 score, and the results are shown in Table 2. Table 2 shows that using binsize = 15 during preprocessing and combining it with the LightGBM algorithm achieves the best diagnostic efficacy in HBV infection diagnosis, with an AUC of 0.94 and an F1 score of 0.85. Figure 1 ).
[0066] Table 2. Impact of different algorithms and binsizes on the performance of the MALDI-TOF-based HBV detection and discrimination model.
[0067] Note: The data is in the upper right corner. * Indicates the statistical significance level. # Indicates marginal saliency, where * P < 0.05, ** P < 0.01, *** P < 0.001, # P < 0.1, and all tests were obtained by Nemenyi post-hoc test or Wilcoxon signed-rank test.
[0068] ③ Model optimization:
[0069] After selecting LightGBM to build the HBV infection diagnosis prediction model, we optimized the hyperparameter of the prediction model, namely the learning rate. The results showed that adjusting the parameter to 0.4152 achieved the optimal performance for the HBV infection diagnosis model. Figure 2 After parameter adjustment, the AUC of the HBV infection diagnosis prediction model was 0.94. Figure 3 ), while the F1Score is 0.87.
[0070] To improve the interpretability of the classifier model, this study utilizes Shapley Additive Explanation (SHAP) to analyze model features. SHAP values are a method in machine learning used to explain model decisions; they assign weights based on the importance of model features, helping to understand the important features of the model. Subsequently, the model's generalization ability is enhanced by constructing a model using 100 sets of random numbers, selecting results where at least 66 features have non-zero SHAP values, to stabilize the features and optimize the prediction model. Figure 4 Its specific characteristic peak regions are as follows: Figure 5 As shown, subsequent studies mainly construct the diagnostic system HBV_pred using these characteristic peak regions.
[0071] (3) Establish a diagnostic system for HBV infection protein fingerprinting HBV_pred
[0072] The HBV infection diagnostic system HBV_pred is constructed by calling a model. The system's usage and calculation method are as follows:
[0073] The protein fingerprint of a single serum sample is input into an optimized model, and then the predicted value is calculated. In the HBV infection diagnostic model, when predict ≥ 0.5, the input fingerprint is considered to indicate that the sample came from an HBV-infected patient (positive); when predict < 0.5, the input fingerprint is considered to indicate that the sample came from an uninfected patient (negative).
[0074] Example 3: HBV_pred System Diagnostic Efficacy Evaluation Results
[0075] (1) Diagnosis of HBV infection using the HBV_pred system
[0076] Six peripheral blood serum samples with unknown HBV infection status (Test 1-6) were randomly collected according to Example 1. Serum MALDI-TOF MS was performed on the peripheral blood serum samples according to the protocol in Example 2, along with ELISA and HBV DNA quantitative qPCR. The detection time, cost, number of reagents, and storage conditions of the three detection technologies were compared, and the results are shown in Table 3.
[0077] In addition, to assess the specificity of the diagnostic system, two peripheral blood serum samples were collected, one infected with syphilis (Test 7) and the other with hepatitis C virus (Test 8), but confirmed to be HBV-free. Furthermore, purified HBV virus particles (Test 9-11) were prepared by concentrating the supernatants of HepG2.2.15 cells, HBV-infected HepG2-NTCP cells, and Huh7-NTCP cells, respectively, for subsequent comprehensive evaluation of the diagnostic system's performance. The results are shown in Table 4.
[0078] The results indicate that, in terms of detection time, ELISA takes approximately 2 hours, qPCR takes approximately 3 hours, while the HBV_pred system completes diagnosis in approximately 1 minute. This demonstrates that the HBV_pred system of this invention is superior to traditional serological and molecular biological detection methods in diagnosing HBV infection, offering the advantage of shorter detection time.
[0079] Regarding testing costs, an ELISA test kit costs approximately 28 yuan per sample, a qPCR test kit costs approximately 56 yuan per sample, while the HBV_pred system only costs 1 yuan per sample for diagnosis. Compared to existing testing methods, the HBV_pred system of this invention can significantly reduce the cost of testing.
[0080] Regarding the quantity of reagents and storage conditions, the ELISA kit includes nine consumables: microplate, six standard tubes, sample diluent, HRP antibody, 20× wash buffer, substrate A, substrate B, stop solution, and sealing film. The kit must be stored at 4°C. qPCR requires at least three reagents: qPCR SuperMix, RNase-free H2O, and ROX II. The kit must be stored at -20°C and protected from repeated freeze-thaw cycles. Furthermore, DNA extraction is required before qPCR. The DNA extraction kit includes eight consumables: lysis buffer, wash buffer I, wash buffer II, RNase-free H2O, proteinase K, proteinase K preparation solution, carrier RNA, and silica gel adsorption column kit. Proteinase K and carrier RNA must be stored at -20°C and protected from repeated freeze-thaw cycles, while proteinase K preparation solution must be stored at 4°C. The HBV_pred system only requires 70% formic acid solution and standard solutions for protein proteomics detection, and can be stored at room temperature.
[0081] Regarding detection specificity, all selected samples showed negative results. Specifically, no characteristic peaks were detected in pure HBV virus particles, and the detection signals for syphilis and hepatitis C viruses showed no cross-reactivity with the HBV characteristic peaks. This demonstrates that the HBV_pred system of this invention possesses strong specificity and accuracy. All of the above indicates that the HBV_pred system of this invention is easy to operate, highly specific, consumes few samples, saves significant costs, and requires no stringent, clean testing environment, making it very user-friendly even for non-professionals.
[0082] Table 3. Diagnostic results of HBV infection using different identification methods
[0083] Table 4. Accuracy verification of different identification schemes
[0084] The above detailed description is a specific description of the embodiments of the present invention. These embodiments are not intended to limit the patent scope of the present invention. All equivalent implementations or modifications that do not depart from the present invention should be included in the patent scope of this case.
Claims
1. A method for constructing a diagnosis model of HBV infection, characterized in that, The method comprises the following steps: (1) Collecting fresh peripheral blood serum of patients, and performing HBsAg detection to determine patient grouping; (2) Collecting protein fingerprint of peripheral blood serum of patients by using MALDI-TOF MS, and marking the protein fingerprint as "1" for HBV infected patients and "0" for non-HBV infected patients according to sample sources; dividing all the labeled graphs into a training set and a test set at random to form a binary classification data set for subsequent modeling; (3) Standardizing the collected protein fingerprint of the binary classification data set by using the R package MALDIquant, including quality control, variance stabilization, graph smoothing, baseline calibration and intensity normalization, and obtaining a record file of protein fingerprint peak position and intensity after processing; (4) Processing the binary classification data set after standardization by binning in the range of 2000-20000 m / z, setting different groups with binsize of 3, 5, 10 and 15 m / z, and performing independently for each sample; (5) Performing in-sample standardization on the binned data: in each sample, summing the intensity values of all bins, and then dividing each bin value by the total sum to complete the total normalization, generating a fixed-length proportion feature vector for subsequent modeling; (6) Calling Python packages lightgbm, keras and scikit-learn, and using LightGBM, DNN and RF algorithms to train the data set, and selecting the optimal modeling condition as LightGBM algorithm and binsize = 15 by F1 Score index, and retraining under the combination to obtain a prediction model for HBV infection diagnosis; (7) Learning rate optimization of the prediction model to obtain an optimized diagnosis model; (8) Robust feature screening: fixing the hyperparameters of the optimized diagnosis model, performing 100 times of Bootstrap resampling training, calculating the absolute mean of each m / z-bin by SHAP each time, and retaining the peak region with non-zero SHAP value in ≥66 iterations; finally determining the stable fingerprint feature peak region; (9) Constructing a final model HBV_pred for HBV infection diagnosis based on the stable fingerprint feature peak region.
2. The construction method of claim 1, wherein, In step (6), the retraining under the combination to obtain a prediction model for HBV infection diagnosis is specifically: first binning the mass spectrum data by 15 m / z and performing in-sample proportion normalization, and then training by using LightGBM algorithm, and locking the hyperparameters according to the highest AUC and F1 to obtain the prediction model for HBV infection diagnosis.
3. The construction method of claim 1, wherein, In step (7), the learning rate optimization of the prediction model is specifically: adjusting the learning rate hyperparameter to 0.4152.
4. The construction method of claim 1, wherein, In step (8), the stable fingerprint characteristic peak region is: 2015-2030, 2030-2045, 2705-2720, 2735-2750, 3125-3140, 3920-3935, 4115-4130, 4790-4805, 6455-6470, 7565-7580, 7790-8050, 8150-8165 m / z.
5. The construction method of claim 1, wherein, Step (9) is specifically: only extract the intensity value of the stable fingerprint characteristic peak region, re-normalize within the sample, retrain with the same hyperparameters, and obtain the final model HBV_pred for HBV infection diagnosis.
6. The HBV infection diagnosis model obtained by the construction method according to any one of claims 1-5.
7. The use of the HBV infection diagnosis model of claim 6 in identifying HBV infection for non-disease diagnosis and treatment purposes.
8. Use according to claim 7, characterized in that, The method comprises the following steps: At least 16 independent protein fingerprint acquisitions of a single serum sample of peripheral blood of a patient to be tested are performed using MALDI-TOF MS, the acquired fingerprints are subjected to sum value standardization, the infection diagnosis model is called, the predict value of each sample is calculated, and it is determined whether the sample to be tested is infected with HBV according to the predict value, and the reliability of the diagnosis result is calculated.
9. Use according to claim 8, characterized in that, The determination of whether the sample to be tested is infected with HBV according to the predict value is specifically: the protein fingerprint of a single serum sample is input into the infection diagnosis model, and then the predict value is calculated; in the HBV infection diagnosis model, when predict ≥ 0.5, it is considered that the determination result predicted by the input fingerprint for the model is that the sample comes from a patient infected with HBV; when predict < 0.5, it is considered that the determination result predicted by the input fingerprint for the model is that the sample comes from a patient not infected with HBV.