Hepatitis B diagnostic model, its construction method, and hepatitis B diagnostic chip
By constructing a hepatitis B diagnostic model and chip, and combining peptide chip data and clinical indicators, the problem of tracking the progression of hepatitis B patients has been solved, enabling early and accurate diagnosis and staging of hepatitis B.
Patent Information
- Application Number
- CN202210830821.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-15
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-07-15
AI Technical Summary
Existing technologies struggle to provide a simple and rapid method to track the disease progression of hepatitis B patients at different stages, especially when early symptoms are not obvious and accurate diagnostic methods are lacking.
By acquiring peptide chip data and target clinical indicator data from different types of hepatitis B patients, a hepatitis B diagnostic model is established using machine learning methods. Combining differential peptides on the peptide chip with clinical indicators, a hepatitis B diagnostic chip is constructed to achieve staged diagnosis of hepatitis B patients.
It improves the diagnostic efficiency and accuracy of hepatitis B staging, enabling early identification of disease progression in hepatitis B patients and reducing misdiagnosis and missed diagnosis.
Smart Images

Figure BDA0003748255080000091 
Figure BDA0003748255080000101 
Figure BDA0003748255080000102
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence application technology, and more specifically, to a hepatitis B diagnostic model, its construction method, apparatus, and hepatitis B diagnostic device. Background Technology
[0002] Liver cancer is a common cancer, and hepatitis and liver cancer are closely related. Clinically, liver cancer patients often have a history of progression from acute hepatitis to chronic hepatitis to cirrhosis to liver cancer. Chronic hepatitis B is an important basis for the development of cirrhosis and liver cancer; a small percentage of hepatitis B patients will develop cirrhosis or even liver cancer. Chronic hepatitis B can also lead to liver cancer directly without going through the cirrhosis stage.
[0003] Because diagnosed hepatitis B patients have strong liver compensatory function, the disease usually progresses very slowly in the early stages with subtle symptoms, making it difficult for patients to detect. By the time symptoms of liver damage appear, the optimal period for treatment and control has often been missed. Patients diagnosed with chronic hepatitis need to undergo regular, complex testing.
[0004] Currently, the diagnosis of hepatitis, cirrhosis, and liver cancer is mostly based on a combination of clinical characteristics and laboratory biochemical and pathological methods. Patients need to cooperate with numerous related examinations, including blood biochemistry, imaging, and liver biopsy. Even patients in the early stages of chronic hepatitis need regular or irregular testing of liver function indicators, complete blood count, quantitative hepatitis B virus gene detection, and alpha-fetoprotein (AFP) testing. Furthermore, doctors and patients need to monitor the disease progression regularly using a combination of these multiple indicators.
[0005] Structural changes in proteins within tumor cells can generate tumor-associated antigens (TAAs), which can stimulate the immune system to produce tumor-associated autoantibodies (AABs). AABs can reflect tumor development and progression as well as the body's immune status. AABs have a long half-life and appear in serum months or even years before clinical diagnosis. AABs targeting TAAs change accordingly with alterations in the biological behavior of tumor cells, and thus have potential value in detecting tumor progression and predicting prognosis.
[0006] Serum hepatitis B surface antigen (HBsAg), HBV DNA quantification, and hepatitis B e antigen (HBeAg) are the main standards for diagnosing patients with chronic hepatitis B. The diagnosis of liver cancer using autoimmune antibodies typically employs enzyme-linked immunosorbent assay (ELISA). This method can be used for single AAB detection, but the positive rate for a single AAB in liver cancer is generally low (10%–30%). For example, anti-P53 AAB has a positive rate of 83.3% in AFP-positive HCC patients, but it is a pan-cancer AAB and lacks specificity for liver cancer. High-throughput protein detection technology can utilize multiple TAAs to screen for combinations of AABs, improving the detection rate of liver cancer. Establishing diagnostic models of AAB combinations can also predict the occurrence of liver cancer.
[0007] AAB combinations are more sensitive than single AABs in the diagnosis of HCC. However, using a positive result for all AABs as the criterion for a positive result results in low sensitivity. The composition of AABs varies significantly among different AAB combinations, leading to differences in diagnostic performance. Some TAAs are already present in benign diseases. Using pre-existing diseases such as hepatitis or cirrhosis as controls increases bias from non-study factors, resulting in significant differences in the diagnostic performance of the same AAB among different researchers.
[0008] Currently, there is no single detection method to accurately track the disease progression of hepatitis B patients. Therefore, providing a simple and rapid method to track hepatitis B patients at different stages is a challenge in hepatitis B diagnosis. Summary of the Invention
[0009] To address the aforementioned problems and achieve the tracking and diagnosis of hepatitis B patients at different stages, the primary objective of this invention is to provide a method for constructing a hepatitis B diagnostic model, which includes:
[0010] Acquire peptide chip data and target clinical indicator data from samples of different types of hepatitis B patients. The peptide chip data includes characteristic signal data of target differentially expressed peptides, which are selected from at least five peptides with sequences shown in SEQ ID NO.1 to SEQ ID NO.25.
[0011] A hepatitis B diagnostic model was established using machine learning methods based on peptide chip data, target clinical indicator data, and sample information from multiple samples of different types of hepatitis B patients.
[0012] In one implementation of the present invention, the sample type includes at least two of the three types: chronic hepatitis, cirrhosis, and liver cancer.
[0013] In one implementation of the present invention, the sample is a blood sample.
[0014] In one implementation of the present invention, the target clinical indicator data includes numerical data of at least one of the following indicators: gender, age, infection cycle, liver function, HBV, liver fibrosis, AFP, AFP isoform, autoimmune disease, and diabetes.
[0015] In one implementation of the present invention, the liver function indicators include at least one of the following: ALB, A / G, AST, ALT, GGT, ALP, PALB, CHE, TBIL, DBIL, IDBIL, TBA, MYO, and UA.
[0016] HBV indicators include at least one of the following: HBsAg, Anti-HBs, HBeAg, Anti-HBe, HBcAb-IgM, Anti-HBII, and HBV-DNA.
[0017] Liver fibrosis markers include TP (thrombocytopenic purpura) levels.
[0018] In one implementation of the present invention, the target differential peptide is selected from the five peptides shown in SEQ ID NO.1 to SEQ ID NO.5; or
[0019] The target differentially expressed peptides are selected from the 10 peptides shown in SEQ ID NO.1 to SEQ ID NO.10; or
[0020] The target differentially expressed peptides are selected from the 15 peptides shown in SEQ ID NO.1 to SEQ ID NO.15; or
[0021] The target differentially expressed peptides are selected from the 20 peptides shown in SEQ ID NO.1 to SEQ ID NO.20; or
[0022] The target differentially expressed peptides were selected from the 25 peptides shown in SEQ ID NO.1 to SEQ ID NO.25.
[0023] In one implementation of the present invention, the machine learning method employs any of the following algorithms: logistic regression, linear discriminant analysis, support vector machine, and random forest;
[0024] Preferably, the machine learning method used is the random forest algorithm.
[0025] The second objective of this invention is to provide a method for screening peptides for preparing hepatitis B diagnostic chips, specifically comprising:
[0026] Acquire peptide chip data from samples of different types of hepatitis B patients, and determine a peptide set including differentially expressed peptides among different types of hepatitis B patients based on the peptide chip data from samples of different types of hepatitis B patients.
[0027] A first machine learning model was established based on peptide chip data of differentially expressed peptides, target clinical indicator data, and sample information from different types of hepatitis B patient samples.
[0028] The target differentially expressed peptides in the peptide set used to construct a hepatitis B diagnostic model were determined based on the AUC value of the first machine learning model.
[0029] In one implementation of the present invention, obtaining polypeptide chip data from samples of different types of hepatitis B patients, and determining a polypeptide set including differentially expressed peptides among different types of hepatitis B patients based on the polypeptide chip data of the samples of different types of hepatitis B patients specifically includes:
[0030] We acquired peptide chip data from samples of different types of hepatitis B patients and performed differential analysis on the peptide chip data of samples of different types of hepatitis B patients. Based on the differential analysis results, we screened a peptide set that included differentially expressed peptides among different types of hepatitis B patients according to a preset threshold.
[0031] In one implementation of the present invention, before acquiring polypeptide chip data from different types of hepatitis B patient samples, the method further includes:
[0032] We acquired peptide chip detection data from samples of different types of hepatitis B patients, and performed gridding processing on the peptide chip detection data to extract peptide chip signal intensity data from each hepatitis B patient sample.
[0033] Peptide chip data for each hepatitis B patient sample was generated based on the peptide chip signal intensity data of each patient sample.
[0034] In one implementation of the present invention, after performing gridding processing on the peptide chip detection data to extract the peptide chip signal intensity data of each hepatitis B patient sample, the method further includes:
[0035] Sample quality control and system stability quality control were performed on the peptide chip signal intensity data of each hepatitis B patient sample;
[0036] Samples or chips whose peptide chip signal intensity data do not meet the quality control standards should be retested.
[0037] In one implementation of the present invention, sample quality control includes at least one of sample signal oversaturation quality control, sample signal distribution quality control, sample gridding location quality control, sample outlier quality control, and sample CV value quality control; and / or
[0038] System stability quality control includes at least one of the following: quality control of reference materials and quality control of reference material CV values.
[0039] In one implementation of the present invention, generating polypeptide chip data for each hepatitis B patient sample based on the polypeptide chip signal intensity data specifically includes:
[0040] The raw peptide chip data for each hepatitis B patient sample is obtained by matrixing the peptide chip signal intensity data of each sample. The raw peptide chip data for each hepatitis B patient sample is then standardized to obtain the peptide chip data for each sample.
[0041] In one implementation of the present invention, the standardization processing of the raw peptide chip data of each hepatitis B patient sample to obtain the peptide chip data of each hepatitis B patient sample specifically includes:
[0042] The raw peptide chip data of each hepatitis B patient sample is logarithmically transformed after adding a preset constant to the raw peptide chip data of each hepatitis B patient sample.
[0043] The peptide chip data for each hepatitis B patient sample is obtained by subtracting the median of the peptide chip conversion data for each array from the peptide chip conversion data of each array.
[0044] In one implementation of the present invention, the target clinical indicator data includes numerical data of at least one of the following indicators: gender, age, infection cycle, liver function, HBV, liver fibrosis, AFP, AFP isoform, autoimmune disease, and diabetes.
[0045] In one implementation of the present invention, the liver function indicators include at least one of the following: ALB, A / G, AST, ALT, GGT, ALP, PALB, CHE, TBIL, DBIL, IDBIL, TBA, MYO, and UA.
[0046] HBV indicators include at least one of the following: HBsAg, Anti-HBs, HBeAg, Anti-HBe, HBcAb-IgM, Anti-HBII, and HBV-DNA.
[0047] Liver fibrosis markers include TP (thrombocytopenic purpura) levels.
[0048] A third objective of this invention is to provide a method for constructing a consistent hepatitis B diagnostic model, specifically including:
[0049] The study aims to acquire target clinical indicator data from different types of hepatitis B patient samples. These target clinical indicator data include numerical data for at least one of the following: gender, age, infection cycle, liver function, HBV, liver fibrosis, AFP, AFP isoform, autoimmune disease, and diabetes. Based on the target clinical indicator data and sample information from multiple different types of hepatitis B patient samples, a hepatitis B diagnostic model will be established using machine learning methods. The sample information will include the sample type, which includes at least two of the following three types: chronic hepatitis, cirrhosis, and liver cancer.
[0050] In one embodiment of the present invention, liver function indicators include at least one of the following: ALB, A / G, AST, ALT, GGT, ALP, PALB, CHE, TBIL, DBIL, IDBIL, TBA, MYO, and UA; HBV indicators include at least one of the following: HBsAg, Anti-HBs, HBeAg, Anti-HBe, HBcAb-IgM, Anti-HBII, and HBV-DNA; and liver fibrosis indicators include TP.
[0051] In one embodiment of the present invention, the machine learning method employs any of the following algorithms: logistic regression, linear discriminant analysis, support vector machine, and random forest.
[0052] In one embodiment of the present invention, the machine learning method employs the random forest algorithm.
[0053] The fourth objective of this invention is to provide a hepatitis B diagnostic model constructed using the above-described construction method.
[0054] The fifth objective of this invention is to provide a hepatitis B diagnostic chip, wherein the peptides fixed on the chip are selected from at least five peptides with sequences shown in SEQ ID NO.1 to SEQ ID NO.25.
[0055] The sixth objective of this invention is to provide a hepatitis B diagnostic device, comprising:
[0056] Data acquisition module: used to acquire peptide chip data and / or target clinical indicator data of hepatitis B patient samples to be tested. The peptide chip data includes characteristic signal data of target differentially expressed peptides. The target differentially expressed peptides are selected from at least five peptides with sequences shown in SEQ ID NO.1 to SEQ ID NO.25.
[0057] Hepatitis B staging diagnosis module: used to input the characteristic signal data of the target differential peptide segment and the target clinical indicator data of the hepatitis B patient sample to be tested into the above-mentioned hepatitis B diagnosis model, or to input the target clinical indicator data of the hepatitis B patient sample to be tested into the above-mentioned hepatitis B diagnosis model, and determine the type of the hepatitis B patient sample to be tested based on the output of the hepatitis B diagnosis model.
[0058] In one implementation of the present invention, the type of hepatitis B patient sample to be tested includes any one of three sample types: chronic hepatitis, cirrhosis, and liver cancer.
[0059] A fourth objective of this invention is to provide a method for diagnosing hepatitis B, the method comprising:
[0060] Acquire peptide chip data and / or target clinical indicator data from hepatitis B patient samples to be tested. The peptide chip data includes characteristic signal data of target differentially expressed peptides, which are selected from at least five peptides with sequences shown in SEQ ID NO.1 to SEQ ID NO.25.
[0061] The characteristic signal data of the target differential peptide segment and the target clinical indicator data of the hepatitis B patient sample to be tested are input into the above hepatitis B diagnostic model, or the target clinical indicator data of the hepatitis B patient sample to be tested are input into the above hepatitis B diagnostic model, and the type of the hepatitis B patient sample to be tested is determined according to the output of the hepatitis B diagnostic model.
[0062] The present invention also relates to a computer device, including a memory and a processor, wherein the memory stores a computer program, characterized in that the processor executes the computer program to implement the steps of the above-described hepatitis B diagnosis method.
[0063] The present invention also relates to a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the above-described hepatitis B diagnosis method.
[0064] The present invention also relates to a computer program product, comprising a computer program, characterized in that, when executed by a processor, the computer program implements the steps of the hepatitis B diagnosis method described above.
[0065] This invention provides a method for constructing a hepatitis B diagnostic model. Based on target clinical indicators, this method improves the efficiency and accuracy of hepatitis B staging diagnosis. This invention also provides another method for constructing a hepatitis B diagnostic model. This method combines target differentially expressed peptides with clinical indicators among different types of hepatitis B patients to construct a diagnostic model, further improving the model's diagnostic efficiency and accuracy. Attached Figure Description
[0066] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0067] Figure 1 A flowchart of a method for constructing a hepatitis B diagnostic model provided by the present invention;
[0068] Figure 2 This is a structural block diagram of a hepatitis B diagnostic device provided by the present invention;
[0069] Figure 3 An internal structural diagram of a computer device provided in an embodiment of the present invention;
[0070] Figure 4 This is a heatmap of characteristic signals of differentially expressed peptides from three groups of hepatitis B patients in Example 1 of the present invention, wherein the horizontal and vertical axes in the figure represent different grouping situations;
[0071] Figure 5 This is a schematic diagram of the confusion matrix of the hepatitis B diagnostic model constructed from 25 target differential peptide features and 29 target clinical indicator features in Embodiment 1 of the present invention. In the figure, the horizontal axis represents the number of predictions for different sample types, and the vertical axis represents the number of actual samples for different sample types.
[0072] Figure 6 This is the ROC curve of the hepatitis B diagnostic model constructed from 25 target differentially expressed polypeptide features and 29 target clinical indicator features in Example 1 of the present invention. In the figure, the horizontal axis represents the false positive rate and the vertical axis represents the true negative rate.
[0073] Figure 7 A confusion matrix diagram of a classification model based on 29 clinical data points is shown, where the horizontal axis represents the predicted number of different sample types and the vertical axis represents the actual number of different sample types.
[0074] Figure 8 The ROC curve of the classification model is based on 29 clinical data. In the figure, the horizontal axis represents the false positive rate and the vertical axis represents the true negative rate. Class 1 represents CHB, class 2 represents HBC, and class 3 represents HCC.
[0075] Figure 9 This is a heatmap of the characteristic signals of the target differential peptides in three groups of hepatitis B patient samples in Embodiment 1 of the present invention, wherein the horizontal and vertical axes in the figure represent different grouping situations;
[0076] Figure 10This is a schematic diagram showing the ranking of the characteristic importance of the target differentially expressed peptides in the hepatitis B diagnostic model of Embodiment 1 of the present invention. In the figure, the horizontal axis represents the target differentially expressed peptides of different sequences, and the vertical axis represents the degree of importance. Detailed Implementation
[0077] Reference will now be made to detailed embodiments of the present invention, one or more of which are described below. Each example is provided for explanation and not for limitation of the invention. In fact, it will be apparent to those skilled in the art that various modifications and variations can be made to the invention without departing from its scope or spirit. For example, features described or illustrated as part of one embodiment may be used in another embodiment to produce further embodiments.
[0078] Therefore, this invention is intended to cover such modifications and variations falling within the scope of the appended claims and their equivalents. Other objects, features, and aspects of the invention are disclosed in or will be apparent from the following detailed description. It will be understood by those skilled in the art that this discussion is merely a description of exemplary embodiments and is not intended to limit the broader aspects of the invention.
[0079] Peptide chip: A chip based on a substrate material, which includes features with pre-designed quantity, location and sequence. Each feature is a cluster of peptides with the same sequence. The peptide sequences between features are often different. These features form a high-density peptide array.
[0080] Peptide chip technology is a detection technology based on peptide chips. It utilizes the contact between a variety of peptides on a peptide chip and the sample, then employs image acquisition technology to capture various characteristic signals from the peptide chip (specifically, fluorescence images carrying these signals). The signal intensity of each feature in the chip is then output, representing the peptide chip detection result data. Based on this data, analysis of analytes in samples bound to peptides on the chip, and other sample analyses, can be performed.
[0081] Peptide microarray technology can detect serum immune responses induced by viral infection. It has been used for antibody identification and validation, research on autoimmune diseases, tumor markers, allergens, and infectious diseases. Compared to other antibody detection techniques, this technology offers high stability and design flexibility. Purely chemically synthesized peptide molecules possess protective groups on their surface, significantly extending the shelf life of peptide microarrays. Peptide microarrays stored at room temperature for over two years retain full biological activity. The design of peptide microarrays is flexible, enabling high density and high throughput. Probe proteins on the microarray can be designed according to specific needs, allowing for the simultaneous measurement of tens of thousands of protein-peptide biochemical reactions.
[0082] Currently, there are no applications of peptide chip technology in the classification of hepatitis B patients.
[0083] In order to at least partially solve the above-mentioned technical problems, such as Figure 1 As shown, the first aspect of the present invention provides a method for constructing a hepatitis B diagnostic model, the method comprising:
[0084] S10: Obtain peptide chip data and target clinical indicator data from samples of different types of hepatitis B patients. The peptide chip data includes characteristic signal data of target differentially expressed peptides. The target differentially expressed peptides are selected from at least five peptides with sequences shown in SEQ ID NO.1 to SEQ ID NO.25.
[0085] Specifically, peptide chip data from hepatitis B patient samples refers to data obtained based on the characteristic signals generated by the binding of proteins and multiple peptide segments on the peptide chip in hepatitis B patient samples. It can be peptide chip detection data composed of TIFF image files obtained through scanning imaging, raw peptide chip data obtained by converting peptide chip detection data into a data matrix, or standardized peptide chip data obtained by standardizing raw peptide chip data.
[0086] Understandably, in order to facilitate modeling, standardized peptide chip data is prioritized for model building. The peptide chip data can be obtained by matrix transformation and standardization of peptide chip detection data, or it can be obtained by standardization of raw peptide chip data.
[0087] The characteristic signal of the target differentially expressed peptide refers to the characteristic signal generated by the binding of the target differentially expressed peptide in the hepatitis B patient sample and on the peptide chip. Accordingly, the characteristic signal data of the target differentially expressed peptide in the peptide chip data can be the characteristic signal displayed in the TIFF image file obtained from scanning imaging, the characteristic signal intensity of the target differentially expressed peptide in the data matrix, or the standardized data of the characteristic signal intensity after standardization. The standardized characteristic signal of the target differentially expressed peptide is preferentially obtained for model construction.
[0088] In some specific embodiments, the feature signal can be a fluorescence signal.
[0089] In some specific embodiments, the sample types include at least two of the three sample types: chronic hepatitis, cirrhosis, and liver cancer, to achieve classification of at least two of the three sample types. It is understood that the hepatitis B patients in this invention refer to patients infected with HBV. Based on the degree of disease progression in each patient, hepatitis B patients can be divided into three types: chronic hepatitis, cirrhosis, and liver cancer. Correspondingly, hepatitis B patient samples are divided into chronic hepatitis samples, cirrhosis samples, and liver cancer samples according to the patient's disease progression.
[0090] In some specific embodiments, the target differential peptides of the present invention include at least 5 of 25 peptides, the specific sequences of which are shown in SEQ ID NO.1 to SEQ ID NO.25.
[0091] Among them, SEQ ID NO.1: NGALYLSYASG;
[0092] SEQ ID NO.2: SLSVVLG;
[0093] SEQ ID NO.3: ALKLSLKVF;
[0094] SEQ ID NO.4: VGAFALVFG;
[0095] SEQ ID NO.5: RFVFLFLFS;
[0096] SEQ ID NO.6: VFVGVGLFG;
[0097] SEQ ID NO.7: VQVHLLFG;
[0098] SEQ ID NO.8: SALYLKVLF;
[0099] SEQ ID NO.9: SFVLSVG;
[0100] SEQ ID NO.10: ALFFHVKFD;
[0101] SEQ ID NO.11: KLFFAFVG;
[0102] SEQ ID NO.12: VRLFVLVFS;
[0103] SEQ ID NO.13: QFYLQVYFG;
[0104] SEQ ID NO.14: VFNVVLYFG;
[0105] SEQ ID NO.15: QVYFYVYFSE;
[0106] SEQ ID NO.16: LAWVLVVSG;
[0107] SEQ ID NO.17: YLGLLFWFSG;
[0108] SEQ ID NO.18: PPAFVFYARFS;
[0109] SEQ ID NO.19: YFYLYLSAQVL;
[0110] SEQ ID NO.20: NLAWYFQVSG;
[0111] SEQ ID NO.21: VFVFARFVLF;
[0112] SEQ ID NO.22: FGVAVAVFS;
[0113] SEQ ID NO.23: LYQVVLLFG;
[0114] SEQ ID NO.24: PLWQVFVVFS;
[0115] SEQ ID NO.25: NVFVVLVFS.
[0116] In some implementations, the target differential peptide is selected from the five peptides shown in SEQ ID NO.1 to SEQ ID NO.5.
[0117] In some implementations, the target differential peptide is selected from the 10 peptides shown in SEQ ID NO.1 to SEQ ID NO.10.
[0118] In some implementations, the target differential peptide is selected from the 15 peptides shown in SEQ ID NO.1 to SEQ ID NO.15.
[0119] In some implementations, the target differential peptide is selected from the 20 peptides shown in SEQ ID NO.1 to SEQ ID NO.20.
[0120] In some implementations, the target differential peptide is selected from the 25 peptides shown in SEQ ID NO.1 to SEQ ID NO.25.
[0121] To obtain the aforementioned peptide information with diagnostic value for hepatitis B staging, a second aspect of the present invention also provides a method for screening peptides for preparing a hepatitis B diagnostic chip, comprising:
[0122] S1: Obtain peptide chip data from samples of different types of hepatitis B patients, and determine a peptide set including differentially expressed peptides among different types of hepatitis B patients based on the peptide chip data from samples of different types of hepatitis B patients.
[0123] Specifically, peptide chip data refers to data obtained by processing the characteristic signals generated from the specific binding of various peptides on the peptide chip with proteins in the sample, using peptide chips to detect hepatitis B patient samples.
[0124] The differential peptide set refers to the set of peptides in which there are significant differences in peptide characteristic signals among samples of different types of hepatitis B patients, obtained by differential analysis of peptide chip data from samples of different types of hepatitis B patients.
[0125] The first machine learning model is a random forest model, which combines differentially expressed peptides and target clinical indicators. The random forest model is trained using feature signal data of differentially expressed peptides combined with peptide chips from multiple different types of hepatitis B patients, target clinical indicator data, and sample information. The importance ranking of differentially expressed peptides is obtained, and the RFE (Recursive Feature Elimination) method is used to screen peptides that achieve high AUC values for all three categories in the model.
[0126] In some implementations, in order to obtain peptide chip data, the process includes the following steps before acquiring peptide chip data from multiple different types of hepatitis B patient samples in step S1:
[0127] M10: Acquire peptide chip detection data from multiple hepatitis B patient samples of different types, and perform gridding processing on the peptide chip detection data to extract peptide chip signal intensity data from each hepatitis B patient sample.
[0128] M20: Generate peptide chip data for each hepatitis B patient sample based on the peptide chip signal intensity data of each hepatitis B patient sample.
[0129] Specifically, peptide chip detection data refers to TIFF image files obtained by detecting hepatitis B patient samples using peptide chips. This involves the specific binding of various peptides on the peptide chip with proteins in the sample to generate characteristic signals, followed by scanning and imaging of these signals. Mesh processing refers to using conventional mesh processing software, such as the mesh processing software included with HealthTell, to mesh the TIFF images of the test samples. Furthermore, the raw fluorescence signal intensity of all peptides is extracted from the meshed TIFF images; these are the fluorescence signal values, which yield the peptide chip signal intensity data for each hepatitis B patient sample.
[0130] In some specific implementations, the signal intensity data, i.e. the fluorescence signal value, for each peptide segment ranges from 0 to 65535.
[0131] Furthermore, for each patient's test sample, after extracting the characteristic signal intensities of all peptides, one GPR5 data file and one corner image file are output. The raw peptide chip data for each patient sample includes both the GPR5 data file and the corner image file. The GPR5 file contains all the information for a patient's peptide chip test sample, such as the array location and chip number, as well as the signal intensity information of all peptides.
[0132] In some implementation schemes, to ensure the accuracy of the test results, after performing gridding processing on the peptide chip detection data to extract the peptide chip signal intensity data of each hepatitis B patient sample, the following steps are also included:
[0133] M101: Perform sample quality control and system stability quality control on the peptide chip signal intensity data of each hepatitis B patient sample. Samples or chips whose peptide chip signal intensity data do not meet the quality control standards are retested.
[0134] Specifically, sample quality control includes at least one of the following: sample signal oversaturation quality control, sample signal distribution quality control, sample gridding positioning quality control, sample outlier quality control, and sample CV value quality control, in order to ensure that the error of the sample detection data is within a reasonable range.
[0135] The signal oversaturation quality control is used to determine the proportion of features (i.e., peptides) whose original fluorescence intensity (FG) values for a single sample exceed the detection limit of the imager. For chips with an oversaturation ratio greater than the quality control standard for a single sample signal, the exposure time needs to be adjusted and the chip rescanned.
[0136] Sample signal distribution quality control is used to plot the frequency density map of the fluorescence intensity signal (LFG) after logarithmic transformation of a single sample, and to determine whether the distribution of the blank control, standard sample, and test sample is normal. The sample signal distribution quality control standards are as follows: For the blank control, the signal should show a narrow pulse width and high peak value distribution, with the LFG distribution peak value within 3; for the standard sample and the test serum sample, the signal should show an overall positively skewed distribution, and the peak value should be much larger than that of the blank control. If the blank control and standard sample are abnormal, the quality control of the corresponding sample within the chip fails, and internal testing needs to be repeated; if the test sample is abnormal, the sample quality needs to be confirmed, and re-loading should be considered.
[0137] The grid-based localization quality control is used to obtain the corrected coefficient of difference and root mean square (RMS) of each sample except the negative control through Mask Analysis. If the corrected coefficient of difference and RMS of a sample are both less than the quality control standard, the grid-based localization preview of the image is manually checked. If the grid-based localization is manually confirmed to be incorrect, the grid-based localization quality control of that sample fails and needs to be retested.
[0138] Outlier quality control is used to count the number of samples on a chip (excluding all samples of the negative control) whose outlier values represent a feature that accounts for ≤2%. If the number of such samples exceeds the quality control standard, all samples on that chip need to be tested again.
[0139] The sample CV value quality control is used to statistically analyze the mean CV of peptide signal intensity within a single sample. If the mean CV exceeds the quality control standard, the sample needs to be tested again.
[0140] System stability quality control includes at least one of the following: relevant quality control of standard samples and CV value quality control of standard samples, to ensure that the systematic error of the test data is within a reasonable range.
[0141] System stability quality control is used to perform correlation and CV value analysis on the signal strength of all standards tested in this batch (one standard per chip, i.e., 1 standard / all samples) to control system stability.
[0142] Among them, the correlation coefficient of the standard is used to calculate the correlation coefficient of the signal intensity of all standard samples tested in this batch. If the correlation coefficient is greater than the quality control standard, the batch of samples needs to be tested again.
[0143] The standard sample CV value quality control is used to statistically analyze the CV values of the signal intensity of all standard samples tested in this batch. For CV values greater than the quality control standard, the batch of samples needs to be tested again.
[0144] In some implementation schemes, to perform differential analysis on the raw peptide chip data, the specific methods for generating peptide chip data for each hepatitis B patient sample based on the peptide chip signal intensity data of each hepatitis B patient sample include:
[0145] M201: The raw peptide chip data of each hepatitis B patient sample is obtained by matrix processing of the peptide chip signal intensity data of each hepatitis B patient sample, and the raw peptide chip data of each hepatitis B patient sample is obtained by standardization processing of the raw peptide chip data of each hepatitis B patient sample.
[0146] Specifically, raw peptide chip data refers to the data obtained by matrixing the acquired peptide chip signal intensity data. Raw peptide chip data can be stored in a GPR5 format file, including detection information of all peptide chips for at least one sample. In addition to the characteristic signal intensity information of all peptides, it also includes array position and chip number, etc.
[0147] Furthermore, the raw peptide chip data exhibited a Log-Norm distribution. Standardization processes included logarithmic transformation and median standardization to obtain standardized peptide chip data for differential analysis, thereby identifying differentially expressed peptides among different types of hepatitis B patients.
[0148] In some specific implementation schemes, differential analysis is performed on standardized peptide microarray data from a large number of hepatitis, cirrhosis, and liver cancer patient samples to obtain a peptide set consisting of 2599 differentially expressed peptides. The total number of hepatitis, cirrhosis, and liver cancer patient samples can be 180, with roughly the same number of patients in each group, specifically approximately 60 patients per group.
[0149] In some specific implementation plans, the standardization processing of the raw peptide chip data of each hepatitis B patient sample to obtain the peptide chip data of each hepatitis B patient sample specifically includes:
[0150] M2011: The raw peptide chip data of each hepatitis B patient sample is logarithmically transformed after adding a preset constant to the raw peptide chip data of each hepatitis B patient sample.
[0151] M2012: Subtract the median of the peptide chip conversion data of the corresponding array from the peptide chip conversion data of each array of hepatitis B patient samples to obtain the peptide chip data of each hepatitis B patient sample.
[0152] Specifically, the logarithmic transformation process includes: adding a constant to the original peptide chip data (LG) and then performing a logarithmic transformation to obtain the transformed peptide chip data LFG (Log-FG), in order to improve the homoscedasticity of the transformed peptide chip data, so that the measurement accuracy of the transformed peptide chip data is roughly proportional to the intensity; further, subtracting the median of all peptide signals LFG of the array from the peptide signal LFG of each array to standardize the LFG, and obtaining the standardized peptide chip data NLFG as the peptide chip data.
[0153] Furthermore, peptide chip data from samples of different types of hepatitis B patients were obtained. Based on the peptide chip data from these samples, a peptide set including differentially expressed peptides among different types of hepatitis B patients was determined, specifically including:
[0154] S11: Obtain peptide microarray data from samples of different types of hepatitis B patients, and perform differential analysis on the peptide microarray data of samples of different types of hepatitis B patients.
[0155] S12: Based on the differential analysis results, a set of peptides including differentially expressed peptides among different types of hepatitis B patients is selected according to a preset threshold.
[0156] Specifically, differential analysis of peptide characteristic signals can be performed using the ANOVA test (i.e., the F-test) to screen differentially expressed peptides among samples from different types of hepatitis B patients. The screening process used the Bonferroni method to correct the p-value, and the screening threshold was set to the top 0.5% of the corrected FDR.
[0157] Furthermore, after performing differential analysis on the peptide chip data, differentially expressed peptides with significant differences in characteristic signals between samples are screened. The peptide chip data of differentially expressed peptides can be directly used to construct the first machine learning model in combination with target clinical indicators and to screen target differentially expressed peptides.
[0158] Specifically, the target clinical indicator data includes numerical data of at least one of the following indicators: gender, age, infection cycle, liver function, HBV, liver fibrosis, AFP, AFP isoform, autoimmune disease, and diabetes. These data are used to train the first machine learning model in combination with the characteristic signal data of the target differential peptide to construct a hepatitis B diagnostic model.
[0159] Specifically, liver function indicators include 14 test indicators such as ALB and AST, HBV indicators include test indicators such as HBsAg and HBeAg, and target clinical indicators include a total of 29 test indicators.
[0160] In some specific implementation plans, liver function indicators include at least one of the following: ALB, A / G, AST, ALT, GGT, ALP, PALB, CHE, TBIL, DBIL, IDBIL, TBA, MYO, and UA.
[0161] HBV indicators include at least one of the following: HBsAg, Anti-HBs, HBeAg, Anti-HBe, HBcAb-IgM, Anti-HBII, and HBV-DNA.
[0162] Liver fibrosis markers include TP (thrombocytopenic purpura) levels.
[0163] In some specific implementations, the process includes the following steps before obtaining target clinical indicator data for different types of hepatitis B patients in step S1:
[0164] We acquired clinical indicator data for hepatitis B diagnosis from multiple patients with different types of hepatitis B. Based on these data and the corresponding patient types, we established a second machine learning model. We then determined the target clinical indicators for constructing the hepatitis B diagnostic model based on the classification performance of the second machine learning model.
[0165] Understandably, in order to facilitate machine learning, the target clinical indicators of this invention are all numerical indicators, and indicators that are irrelevant to the diagnosis of hepatitis B patients have been removed, so as to build a hepatitis B diagnostic model.
[0166] Specifically, the machine learning methods used in the second machine learning model include any of the following algorithms: logistic regression (LR), linear discriminant analysis (LDA), support vector machine (SVM), and random forest (RF).
[0167] In some specific implementations, the machine learning method employs the Random Forest (RF) algorithm.
[0168] A random forest (RF) model is constructed using the random forest algorithm. It involves creating a forest in a random manner, which consists of many decision trees. Each decision tree is independent of the others. Once the forest is formed, when a new input sample is introduced, each decision tree in the forest makes a judgment to determine which class the sample should belong to. The class that is selected the most is then used to predict the class of the sample.
[0169] The random forest model constructed separately for the 29 target clinical indicator features of this invention also has good classification performance.
[0170] S2: Establish the first machine learning model based on peptide chip data, target clinical indicator data, and sample information from multiple different types of hepatitis B patient samples;
[0171] Specifically, a first machine learning model was constructed using peptide chip data of multiple different types of hepatitis B patient samples combined with differentially expressed peptides, target clinical indicator data, and sample information, in order to screen peptide features with hepatitis B staging diagnostic effects among differentially expressed peptides.
[0172] Specifically, the machine learning methods used in the first machine learning model include any of the following algorithms: logistic regression (LR), linear discriminant analysis (LDA), support vector machine (SVM), and random forest (RF).
[0173] In some specific implementations, the machine learning method employs the Random Forest (RF) algorithm.
[0174] S3: Determine the target differentially expressed peptides in the peptide set for constructing a hepatitis B diagnostic model based on the AUC value of the first machine learning model.
[0175] Specifically, the recursive feature elimination method was used to optimize the differentially expressed peptide features used in the first machine learning model constructed from differentially expressed peptides and target clinical indicators, in order to obtain peptide features that achieve the best results in hepatitis B staging diagnosis.
[0176] In some specific implementation schemes, a random forest model is built using 2599 differentially expressed peptides and 29 clinical indicators to obtain a feature importance ranking. The classification accuracy (AUC) of the initial feature subset is obtained using 10-fold cross-validation. The peptide features with the lowest ranking are removed one by one, and a new importance ranking is obtained by re-modeling and calculating the corresponding classification accuracy (AUC). The remaining peptides when the classification AUCs of all three categories reach high values are selected as the target differentially expressed peptides. Furthermore, a hepatitis B diagnostic model is constructed based on the peptide chip-standardized data of the target differentially expressed peptides and the target clinical indicator data.
[0177] S20: Based on multiple hepatitis B patient samples of different types, combined with peptide chip data of target differentially expressed peptides, target clinical indicator data, and corresponding hepatitis B patient types, a hepatitis B diagnostic model is established using machine learning methods. Sample information includes the sample type.
[0178] In some implementation schemes, the machine learning methods used to build hepatitis B diagnostic models include any of the following algorithms: logistic regression (LR), linear discriminant analysis (LDA), support vector machine (SVM), and random forest (RF).
[0179] In some specific implementations, the machine learning method employs the Random Forest (RF) algorithm.
[0180] In some implementation schemes, the target differentially expressed peptides and target clinical indicators are combined to form a feature. A random forest model is trained based on peptide chip data of the target differentially expressed peptides and target clinical indicator data obtained from peptide chip detection of multiple different types of hepatitis B patient samples. This constructs a hepatitis B diagnostic model, enabling the classification of hepatitis B patients with different disease progressions and improving the accuracy of classifying different types of hepatitis B patients.
[0181] It is understandable that the peptide chip data from hepatitis B patient samples can be peptide chip detection data containing characteristic signals of the target differentially expressed peptides, raw peptide chip data containing characteristic signals of the target differentially expressed peptides, or standardized peptide chip data containing characteristic signals of the target differentially expressed peptides. When the peptide chip data from hepatitis B patient samples is either peptide chip detection data containing characteristic signals of the target differentially expressed peptides, or raw peptide chip data containing characteristic signals of the target differentially expressed peptides, the corresponding peptide chip data can be converted into standardized peptide chip data for use in constructing a hepatitis B diagnostic model. The specific conversion method has been detailed above and will not be repeated here.
[0182] In some specific implementations, the present invention uses the target differential peptides comprising 25 peptide segments to construct a hepatitis B diagnostic model. The present invention constructs a hepatitis B diagnostic model by combining the target differential peptides of different types of hepatitis B patients with clinical indicators, thereby excluding interference from other immune features unrelated to hepatitis B classification, which can improve the diagnostic efficiency and accuracy of the model.
[0183] It is understood that this invention creatively combines the target differential peptides detected by peptide chips with non-invasive and easily accessible target clinical indicators, thereby enabling the acquisition of a hepatitis B diagnostic model with superior classification performance.
[0184] A second aspect of the present invention also provides a method for constructing a hepatitis B diagnostic model, the method comprising:
[0185] To obtain target clinical indicator data for different types of hepatitis B patient samples, the target clinical indicator data includes numerical data of at least one of the following indicators: gender, age, infection cycle, liver function, HBV, liver fibrosis, AFP, AFP isoform, autoimmune disease, and diabetes.
[0186] Based on the target clinical indicator data and sample information of different types of hepatitis B patients, a hepatitis B diagnostic model is established using machine learning methods. The sample information includes the sample type.
[0187] The machine learning model constructed separately based on the target clinical indicator features of this invention also has good classification performance.
[0188] In some implementation schemes, the sample types include at least two of the three types: chronic hepatitis, cirrhosis, and liver cancer.
[0189] In some implementation plans, liver function indicators include at least one of the following: ALB, A / G, AST, ALT, GGT, ALP, PALB, CHE, TBIL, DBIL, IDBIL, TBA, MYO, and UA.
[0190] HBV indicators include at least one of the following: HBsAg, Anti-HBs, HBeAg, Anti-HBe, HBcAb-IgM, Anti-HBII, and HBV-DNA.
[0191] Liver fibrosis markers include TP (thrombocytopenic purpura) levels.
[0192] In some implementation schemes, the machine learning method employs any of the following algorithms: logistic regression, linear discriminant analysis, support vector machine, random forest;
[0193] In some specific implementation schemes, the machine learning method employed the random forest algorithm.
[0194] Therefore, a third aspect of the present invention provides a hepatitis B diagnostic model constructed by the above-described construction method to improve the accuracy and efficiency of classifying and diagnosing hepatitis B patients with different disease progressions.
[0195] A fourth aspect of the present invention provides a hepatitis B diagnostic chip, wherein the peptide segments fixed on the chip are selected from at least five peptide segments of the sequences shown in SEQ ID NO.1 to SEQ ID NO.25, for detecting characteristic signals of target differential peptide segments between different types of patients, thereby realizing diagnosis of different types of patients by combining the characteristic signals with the above-mentioned hepatitis B diagnostic model.
[0196] In some specific implementations, the peptides immobilized on the hepatitis B diagnostic chip are selected from the five peptides shown in SEQ ID NO.1 to SEQ ID NO.5.
[0197] In some specific implementations, the peptides immobilized on the hepatitis B diagnostic chip are selected from the 10 peptides shown in SEQ ID NO.1 to SEQ ID NO.10; or
[0198] In some specific implementations, the peptides immobilized on the hepatitis B diagnostic chip are selected from the 15 peptides shown in SEQ ID NO.1 to SEQ ID NO.15.
[0199] In some specific implementations, the peptides immobilized on the hepatitis B diagnostic chip are selected from the 20 peptides shown in SEQ ID NO.1 to SEQ ID NO.20.
[0200] In some specific implementations, the peptides immobilized on the hepatitis B diagnostic chip are selected from the 25 peptides shown in SEQ ID NO.1 to SEQ ID NO.25.
[0201] A fifth aspect of the present invention provides a hepatitis B diagnostic device, comprising:
[0202] Data acquisition module 100: used to acquire peptide chip data and target clinical indicator data of the hepatitis B patient sample to be tested. The peptide chip data includes characteristic signal data of the target differential peptides. The target differential peptides are selected from at least five peptides with sequences shown in SEQ ID NO.1 to SEQ ID NO.25.
[0203] Hepatitis B staging diagnosis module 200: It is used to input the characteristic signal data of the target differential peptide segment and the target clinical indicator data of the hepatitis B patient sample to be tested into the above hepatitis B diagnosis model, and determine the type of the hepatitis B patient sample to be tested based on the output of the hepatitis B diagnosis model.
[0204] In some implementation plans, the types of hepatitis B patient samples to be tested include any one of three sample types: chronic hepatitis, cirrhosis, and liver cancer.
[0205] A sixth aspect of the present invention provides a method for diagnosing hepatitis B, the method comprising:
[0206] Acquire peptide chip data and target clinical indicator data from hepatitis B patient samples to be tested. The peptide chip data includes characteristic signal data of target differentially expressed peptides, which are selected from at least five peptides with sequences shown in SEQ ID NO.1 to SEQ ID NO.25.
[0207] The characteristic signal data of the target differentially expressed peptide and the target clinical indicator data of the hepatitis B patient sample to be tested are input into the above hepatitis B diagnostic model, and the type of the hepatitis B patient sample to be tested is determined according to the output of the hepatitis B diagnostic model.
[0208] Each module in the aforementioned hepatitis B diagnostic device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0209] In some implementations, a computer device is provided, which can be a server 104 or a terminal 102, and its internal structure diagram can be as follows. Figure 3 As shown. The computer device includes a processor, memory, and communication interface connected via a system bus. When the computer device is a terminal, it also includes a display screen and input devices connected to the system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements at least one of the methods for diagnosing hepatitis B. The display screen of the computer device can be a liquid crystal display (LCD) or an electronic ink display. The input devices of the computer device can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse, etc.
[0210] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0211] This application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method for constructing a hepatitis B diagnostic model.
[0212] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the above-described hepatitis B diagnostic model construction method.
[0213] This application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the above-described hepatitis B diagnostic model construction method.
[0214] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0215] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0216] The embodiments of the present invention will now be described in detail with reference to examples.
[0217] Example 1
[0218] 1. Experimental Sample Information
[0219] This embodiment used samples from 180 HBV-infected patients. In addition to routine clinical examinations, patients with chronic hepatitis and cirrhosis underwent four liver fibrosis tests to confirm their condition, while patients with liver cancer also underwent liver biopsy and related imaging examinations. As shown in Table 1, among the HBV-infected patients, there were 57 cases of chronic hepatitis, 63 cases of cirrhosis, and 60 cases of liver cancer. The corresponding clinical indicators of the samples are shown in Table 2. Gender, autoimmune indicators, and diabetes indicators were all binarized, for example, male was 1 and female was 0; those with autoimmune diseases and diabetes were treated as having a value of 1, and those without autoimmune diseases and diabetes were treated as having a value of 0. Other indicators are the mean ± variance SD.
[0220] Table 1
[0221]
[0222]
[0223] Table 2
[0224]
[0225] 2. Peptide chips generate antibody profile data.
[0226] The detailed process of generating peptide chip data is as follows:
[0227] 2.1 Experimental Design
[0228] One 96-well plate constitutes one testing unit. Before the experiment begins, design the experiment carefully, calculating the required number of chips and determining the chip numbering and sample arrangement based on the number of test samples, blank controls, and standards. Standards, blank controls, and test samples are randomly distributed across all used chips.
[0229] 2.2 Experimental Procedure
[0230] 1) Sample preparation
[0231] Serum or plasma samples were diluted twice with 1% D-mannitol solution in a 96-well deep plate, resulting in a 625-fold diluted sample plate for later use.
[0232] 2) Chip hydration and assembly
[0233] The chip was placed in a chip hydration apparatus, and ultrapure water was added to cover the chip. Hydration was carried out on a track shaker at 55±5 rpm / min for 20 min. Then, isopropanol was sprayed onto the chip surface, and the chip was centrifuged to dry. The dried chips were then assembled into an assay cassette according to the experimental design.
[0234] 3) Sample and chip incubation combined
[0235] Add the diluted sample at a rate of 90 μL / well to the assembled chip and incubate it on a constant temperature shaker for 1 hour.
[0236] 4) Sample cleaning
[0237] The analysis kit was placed in a plate washer for cleaning.
[0238] 5) Incubation with fluorescent secondary antibody
[0239] Prepare a 2 nM fluorescent secondary antibody solution using 0.75% casein solution, add 40 μL per well to the analysis kit, and incubate with shaking on a constant temperature shaker for 1 hour.
[0240] 6) Secondary antibody washing
[0241] Same as step 3).
[0242] 7) Imaging
[0243] After the chip in the analysis cassette is disassembled, cleaned, and dried, it is assembled into the imaging cassette and placed into Molecular Devices' ImageXpress micro4 imager for scanning and imaging. Each sample ultimately yields a TIFF image file, which is the raw data.
[0244] 2.3 Data Preprocessing
[0245] 1) Using a self-developed MIAMI pipeline, the TIFF images generated for each sample were meshed (it should be noted that existing meshing software can perform meshing on the TIFF images generated for each sample, such as the meshing software included with HealthTell). Then, the fluorescence intensity values of the features (here, features refer to the set of differentially expressed peptides) were extracted, and one GPR5 data file and one image localization result file were output. The GPR5 file contains all the information of a sample and the fluorescence intensity information of all features.
[0246] 2) Extract the fluorescence intensity information of features from the GPR5 data files of all samples to generate the original fluorescence intensity (FG, foreground) data matrix. Then, perform logarithmic transformation on the data of each sample to obtain the LFG (log-transferred foreground) data matrix and perform median normalization to obtain the NLFG (normalized and log-transferred foreground) data matrix. This step also generates a sample chip information file, which includes information such as sample array position and chip number used.
[0247] 2.4 Quality Control
[0248] (1) Single Sample Quality Control
[0249] 1) Oversaturation Analysis
[0250] Quality control index: The proportion of features (i.e., peptides) in the single sample's original fluorescence intensity value (FG) that exceed the upper detection limit of the imager.
[0251] Quality control standard: The above proportion ≤ 1% is considered qualified.
[0252] Treatment method when quality control fails: Adjust the exposure time and rescan the chip until the proportion ≤ 1%.
[0253] 2) Fluorescence Signal Distribution Analysis
[0254] Quality control index: Plot the frequency density diagram for the logarithmically transformed fluorescence intensity signal (LFG) of a single sample, and respectively determine whether the distributions of the blank control, standard sample, and test sample are normal.
[0255] Quality control standard: For the blank control, confirm that the signal shows a narrow pulse width and high peak value distribution, and the peak value of the LFG distribution is within 3; for the standard sample and test serum sample, confirm that the overall signal shows a positive skewed distribution, and the peak value is much larger than that of the blank control.
[0256] Treatment method when quality control fails: If the blank control and standard are abnormal, the quality control of the test samples in the corresponding chip fails; if the test sample is abnormal, the quality of the sample needs to be confirmed, and re-sampling should be considered.
[0257] 3) Grid Location Accuracy Quality Control
[0258] Quality control index: The corrected determination coefficient and root mean square (RMS) obtained by performing Mask Analysis on each sample other than the negative control.
[0259] Quality control standard: The corrected determination coefficient ≥ 0.3 and RMS ≥ 0.3 are considered qualified.
[0260] Treatment method when quality control is unqualified: Manually check the grid situation of the image positioning result. If it is manually confirmed that the grid is incorrect, the quality control of this sample fails and it needs to be redetected.
[0261] 4) Outlier analysis
[0262] Quality control index: The number of samples in which the proportion of features represented by outliers on one chip (24 samples, excluding negative controls) ≤ 2%.
[0263] Quality control standard: If the above number of samples does not exceed 2, the chip passes the quality control.
[0264] Treatment method when quality control is unqualified: All samples on this chip need to be retested.
[0265] 5) Analysis of the mean value of quality control peptides (coefficient of variation, CV)
[0266] Quality control index: The CV mean value of the signal intensity of quality control peptides within a single sample.
[0267] Quality control standard: The above CV mean value ≤ 1% is qualified.
[0268] Treatment method when quality control is unqualified: This sample needs to be retested.
[0269] (2) Quality control of system stability
[0270] Perform correlation and CV value analysis on the signal intensities of all standard products in this batch of detections (one standard product per chip, that is, 1 standard product / 24 samples) to conduct quality control of system stability.
[0271] 1) Standard product correlation
[0272] Quality control index: The correlation coefficient of the signal intensities of all standard samples in this batch of detections.
[0273] Quality control standard: The above correlation coefficient ≥ 0.8 is qualified.
[0274] Treatment method when quality control is unqualified: This batch of samples needs to be retested.
[0275] 2) CV mean value of standard samples
[0276] Quality control index: The CV mean value of the signal intensities of all standard samples in this batch of detections.
[0277] Quality control standard: The above CV mean value ≤ 4% is qualified.
[0278] Treatment method when quality control is unqualified: This batch of samples needs to be retested.
[0279] 3. Filter differential signals
[0280] a. Obtain raw data FG
[0281] Using the V13 peptide chip technology, the sample was tested according to the above standard procedure, and the signal values of 126,501 peptide segments were obtained from the V13 chip. Each peptide segment signal value is called a feature, and its value range is 0 to 65535. The raw data is called FG (foreground) and stored in a GPR5 format file.
[0282] b. Perform data correction
[0283] The raw data FG for each peptide is extracted from the data matrix in GPR5 format. Since the raw fluorescence signal data follows a Log-Norm distribution, FG is logarithmically transformed by adding a constant of 100 to obtain LFG (Log-FG) to improve homoscedasticity. The measurement accuracy of the data is roughly proportional to the intensity. The normalized data NLFG is obtained by subtracting the median of all peptides in the array from the peptide signal LFG of each array.
[0284] c. Statistical tests to obtain differentially expressed peptides
[0285] Following HBV infection, the human immune system clears the virus, resulting in an increase in specific antibody signals. These antibody signals change as the disease progresses. To screen for characteristic peptides that can characterize all three disease groups, ANOVA (F-test) was used to analyze the differences in all 180 samples from chronic hepatitis, cirrhosis, and liver cancer, identifying peptides with significant differences among the three groups. The Bonferroni method was used to correct the p-value. The screening threshold was set to the top 0.5% of the corrected FDR. 2599 peptides meeting the criteria were selected as characteristic peptides. Figure 4 Heatmaps of characteristic peptides for the three patient types after screening are shown. White represents characteristic peptides of CHB patients, gray represents characteristic peptides of HBC patients, and black represents characteristic peptides of HCC patients. After clustering the peptides and samples, it can be seen that the differential peptide sets are sufficient to reflect the immune characteristics of the three disease groups.
[0286] 4. Screening clinical indicators useful for modeling
[0287] 1) Screening and Preprocessing of Clinical Indicators: To enhance the clinical value of the obtained model, the following clinical indicators were screened: 1. Unquantifiable clinical examinations were discarded, such as imaging results, which are based on physician descriptions and cannot be quantified; 2. Clinical indicators related to the four liver fibrosis markers and AFP isoforms were removed, as these require invasive procedures such as biopsy; 3. Liver cancer-specific indicators CEA and CA199 were removed, as these tests are generally not performed on patients in other disease groups; 4. Blood glucose levels were removed, as blood glucose is generally considered unrelated to hepatitis B. The remaining 29 clinical indicators (see Table 2) were used for modeling. These indicators are all numerical, readily available clinically, and relatively easy to detect.
[0288] 2) Clinical indicator evaluation: A classification model was established based on 29 clinical indicators. The model was trained and evaluated using the RF algorithm and 10-fold cross-validation. The final classification evaluation results of the model are shown in Table 3.
[0289] Table 3
[0290]
[0291]
[0292] 5. Establish a hepatitis B diagnostic model based on differentially expressed peptides and clinical indicators.
[0293] The 2599 peptides and 29 clinical indicators were combined, and the Recursive Feature Elimination (RFE) method was used to screen peptides based on the model's AUC value. Specifically, an RF model was first built using the 2599 peptides and 29 clinical indicators to obtain a feature importance ranking. The classification accuracy AUC of the initial feature subset was obtained using 10-fold cross-validation. Then, the peptide features with the lowest ranking were removed one by one, and a new importance ranking was obtained. The corresponding classification accuracy AUC was calculated. Finally, the set of 25 peptides (as shown in Table 4) remaining after the classification AUC values for CHB, HBC, and HCC all reached high values were selected. The AUC values are as follows: Figure 6 As shown, based on the importance of the features returned by the model, the importance of each peptide in Table 4 decreases sequentially from SEQ ID NO.1 to SEQ ID NO.25.
[0294] Table 4
[0295]
[0296] The classification evaluation results of the model are shown in Table 5, and the model confusion matrix is as follows: Figure 5 As shown, the ROC results of the model are as follows: Figure 6As shown in the figure. The heatmap results of the signals of 25 important peptide features across all samples are as follows. Figure 9 As shown.
[0297] Figure 10 The left half represents the ranking of the importance of clinical indicator features in the hepatitis B diagnostic model constructed in this invention, and the right half represents the ranking of the importance of polypeptide features. The clinical indicator features and polypeptide features are ranked separately, and it does not mean that the clinical indicator is ranked before the polypeptide.
[0298] Macro-average is a simple and direct metric for evaluating multi-class classification models. This method averages the evaluation metrics (Precision / Recall / F1-score) for different classes, ensuring equal weighting for each class. The hepatitis B diagnostic model constructed in this embodiment includes 25 peptide features and 29 clinical indicator features. The importance ranking of all features in the model is as follows: Figure 8 As shown in the figure, compared with models that only use clinical indicators as features, the overall macro-average of all classification performance indicators of the hepatitis B diagnostic model constructed in this embodiment is improved, especially the accuracy and f1 score, and the AUC is also improved. As can be seen from the figure, the 25 newly added peptide features are beneficial to improving the model's classification performance.
[0299] Table 5
[0300] Evaluation parameters CHB HBC HCC micro-average accuracy 0.788889 0.744444 0.855556 0.796296 f1 score 0.677966 0.616667 0.786885 0.694444 False detection rate 0.298246 0.412698 0.200000 0.305556 False Negative Rate 0.344262 0.350877 0.225806 0.305556 False positive rate 0.142857 0.211382 0.101695 0.152778 Negative predictive value 0.829268 0.829060 0.883333 0.847222 Positive predictive value 0.701754 0.587302 0.800000 0.694444 accuracy 0.701754 0.587302 0.800000 0.694444 Recall rate 0.655738 0.649123 0.774194 0.694444 Sensitivity 0.655738 0.649123 0.774194 0.694444 Specificity 0.857143 0.788618 0.898305 0.847222 True negative rate 0.857143 0.788618 0.898305 0.847222 True rate 0.655738 0.649123 0.774194 0.694444
[0301] Furthermore, based on the importance ranking of peptide features in Table 4, peptide features ranked in the top 20, 15, 10, and 5 were examined. Combined with the 29 clinical indicators identified above, a hepatitis B subtyping diagnostic model was constructed, and the model performance is shown in Tables 6 and 7.
[0302] Table 6
[0303] Evaluation parameters top 20 top 15 top 10 top 5 accuracy 0.785185 0.785185 0.80 0.792593 f1 score 0.677778 0.677778 0.70 0.688889 False detection rate 0.322222 0.322222 0.30 0.311111 False Negative Rate 0.322222 0.322222 0.30 0.311111 False positive rate 0.161111 0.161111 0.15 0.155556 Negative predictive value 0.838889 0.838889 0.85 0.844444 Positive predictive value 0.677778 0.677778 0.70 0.688889 accuracy 0.677778 0.677778 0.70 0.688889 Recall rate 0.677778 0.677778 0.70 0.688889 Sensitivity 0.677778 0.677778 0.70 0.688889 Specificity 0.838889 0.838889 0.85 0.844444 True negative rate 0.838889 0.838889 0.85 0.844444 True rate 0.677778 0.677778 0.70 0.688889 AUC 0.826286 0.838190 0.842755 0.850759
[0304] Note: The values in Table 6 represent the average values of the corresponding assessment parameters for the three categories: CHB, HBC, and HCC.
[0305] Table 7
[0306] AUC for each category top 20 top 15 top 10 top 5 CHB 0.82 0.83 0.84 0.85 HBC 0.76 0.78 0.78 0.79 HCC 0.91 0.91 0.92 0.92
[0307] Among them, the top 20 refers to the hepatitis B diagnostic model constructed using peptide signal data from SEQ ID NO.1 to SEQ ID NO.20 and 29 clinical indicators; the top 15 refers to the hepatitis B diagnostic model constructed using peptide signal data from SEQ ID NO.1 to SEQ ID NO.15 and 29 clinical indicators; the top 10 refers to the hepatitis B diagnostic model constructed using peptide signal data from SEQ ID NO.1 to SEQ ID NO.10 and 29 clinical indicators; and the top 5 refers to the hepatitis B diagnostic model constructed using peptide signal data from SEQ ID NO.1 to SEQ ID NO.5 and 29 clinical indicators.
[0308] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0309] The above embodiments merely illustrate several implementation methods of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A method for constructing a hepatitis B diagnostic model, characterized in that, The method includes: Acquire peptide chip data and target clinical indicator data from samples of different types of hepatitis B patients. The peptide chip data includes characteristic signal data of target differentially expressed peptides, which are the five peptides shown in SEQ ID NO.1 to SEQ ID NO.5; or The target differentially expressed peptides are the 10 peptides shown in SEQ ID NO.1 to SEQ ID NO.10; or The target differentially expressed peptides are the 15 peptides shown in SEQ ID NO.1 to SEQ ID NO.15; or The target differentially expressed peptides are the 20 peptides shown in SEQ ID NO.1 to SEQ ID NO.20; or The target differentially expressed peptides are the 25 peptides shown in SEQ ID NO.1 to SEQ ID NO.25; A hepatitis B diagnostic model was established using machine learning methods based on peptide chip data, target clinical indicator data, and sample information from multiple hepatitis B patient samples of different types. The sample information includes the sample type.
2. The construction method according to claim 1, characterized in that, The sample types include at least two of the three types: chronic hepatitis, cirrhosis, and liver cancer.
3. The construction method according to claim 1, characterized in that, The sample is a blood sample.
4. The construction method according to claim 1, characterized in that, The target clinical indicator data includes numerical data of at least one of the following indicators: gender, age, infection cycle, liver function, HBV, liver fibrosis, AFP, AFP isoform, autoimmune disease, and diabetes.
5. The construction method according to claim 1, characterized in that, Liver function indicators include at least one of the following: ALB, A / G, AST, ALT, GGT, ALP, PALB, CHE, TBIL, DBIL, IDBIL, TBA, MYO, and UA. HBV indicators include at least one of the following: HBsAg, Anti-HBs, HBeAg, Anti-HBe, HBcAb-IgM, Anti-HBII, and HBV-DNA. Liver fibrosis markers include TP (thrombocytopenic purpura) levels.
6. The construction method according to any one of claims 1 to 5, characterized in that, The machine learning method employs any of the following algorithms: logistic regression, linear discriminant analysis, support vector machine, or random forest.
7. The construction method according to claim 6, characterized in that, The machine learning method described uses the random forest algorithm.
8. A method for screening peptides for preparing hepatitis B diagnostic chips, characterized in that, include: Acquire peptide chip data from samples of different types of hepatitis B patients, and determine a peptide set including differentially expressed peptides among different types of hepatitis B patients based on the differential analysis of peptide chip data from samples of different types of hepatitis B patients. A first machine learning model was established based on peptide chip data of differentially expressed peptides, target clinical indicator data, and sample information from different types of hepatitis B patient samples. The target differential peptides for constructing a hepatitis B diagnostic model in the polypeptide set are determined based on the AUC value of the first machine learning model. The target differential peptide is the target differential peptide as described in claim 1.
9. A hepatitis B staging diagnostic model, characterized in that, Constructed according to the construction method according to any one of claims 1 to 7.
10. A hepatitis B diagnostic chip, characterized in that, The peptides fixed on the chip are the target differential peptides as described in claim 1.
11. A hepatitis B diagnostic device, characterized in that, include: Data acquisition module: used to acquire polypeptide chip data and / or target clinical indicator data of hepatitis B patient samples to be tested, wherein the polypeptide chip data includes characteristic signal data of target differential peptides, and the target differential peptides are the target differential peptides as described in claim 1; Hepatitis B staging diagnosis module: used to input the characteristic signal data of the target differential peptide segment and the target clinical indicator data of the hepatitis B patient sample to be tested into the hepatitis B diagnosis model of claim 9, and determine the type of the hepatitis B patient sample to be tested based on the output results of the hepatitis B diagnosis model.
12. The apparatus according to claim 11, characterized in that, The types of hepatitis B patient samples to be tested include any one of three sample types: chronic hepatitis, cirrhosis, and liver cancer.
13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements all the steps for hepatitis B staging diagnosis, including: Acquire peptide chip data and / or target clinical indicator data of a hepatitis B patient sample to be tested, wherein the peptide chip data includes characteristic signal data of a target differentially expressed peptide, and the target differentially expressed peptide is the target differentially expressed peptide as described in claim 1; The characteristic signal data of the target differentially expressed peptide and the target clinical indicator data of the hepatitis B patient sample to be tested are input into the hepatitis B diagnostic model of claim 9, and the type of the hepatitis B patient sample to be tested is determined according to the output of the hepatitis B diagnostic model.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements all the steps for hepatitis B staging diagnosis, including: Acquire peptide chip data and / or target clinical indicator data of a hepatitis B patient sample to be tested, wherein the peptide chip data includes characteristic signal data of a target differentially expressed peptide, and the target differentially expressed peptide is the target differentially expressed peptide as described in claim 1; The characteristic signal data of the target differentially expressed peptide and the target clinical indicator data of the hepatitis B patient sample to be tested are input into the hepatitis B diagnostic model of claim 9, and the type of the hepatitis B patient sample to be tested is determined according to the output of the hepatitis B diagnostic model.
15. A computer program product, comprising a computer program, characterized in that, When executed by a processor, this computer program performs all the steps for hepatitis B staging diagnosis, including: Acquire peptide chip data and / or target clinical indicator data of a hepatitis B patient sample to be tested, wherein the peptide chip data includes characteristic signal data of a target differentially expressed peptide, and the target differentially expressed peptide is the target differentially expressed peptide as described in claim 1; The characteristic signal data of the target differentially expressed peptide and the target clinical indicator data of the hepatitis B patient sample to be tested are input into the hepatitis B diagnostic model of claim 9, and the type of the hepatitis B patient sample to be tested is determined according to the output of the hepatitis B diagnostic model.
Citation Information
Patent Citations
Non-creative scoring model for ULN chronic hepatitis B and hepatic fibrosis and establishment method therefor
CN104866720A