Method for rapidly analyzing and identifying helicobacter pylori infection and antibody typing by using serum Raman spectrum

By combining surface-enhanced Raman spectroscopy with machine learning algorithms, a nanosilver particle Raman-enhanced substrate was prepared, and the SERS spectra of serum samples were collected. Cluster analysis and multiple machine learning algorithm processing were performed, which solved the problems of long detection time, complex operation and low sensitivity in the existing technology, and achieved rapid and accurate Helicobacter pylori infection and antibody typing detection.

CN120668631APending Publication Date: 2025-09-19GUANGDONG GENERAL HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410310211.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies for detecting Helicobacter pylori antibodies in serum have problems such as long detection time, complex operation, and low sensitivity, and machine learning is rarely used in this field.

Method used

Combining surface-enhanced Raman spectroscopy with machine learning algorithms, a fast and accurate detection method was established by preparing a nanosilver particle Raman-enhanced substrate, collecting SERS spectra of serum samples, performing cluster analysis and data processing with multiple machine learning algorithms.

Benefits of technology

It achieves rapid and accurate identification of Helicobacter pylori infection and antibody typing, improves detection efficiency and sensitivity, and optimizes the detection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120668631A_ABST
    Figure CN120668631A_ABST
Patent Text Reader

Abstract

The invention provides a method for rapidly analyzing and identifying helicobacter pylori infection and antibody typing by using a serum Raman spectrum, which solves the problems of tedious detection process, long consumed time and the like in a manner of combining a surface enhanced Raman spectrum preparation method and a machine learning algorithm. On the other hand, a Raman spectrum database of different helicobacter pylori infection and antibody typing is established through machine learning, the newly collected Raman spectrum is directly imported into the database for judgment, and the accuracy of the identification result can be better guaranteed through comparison and combination of different machine learning methods and evaluation indexes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Raman spectroscopy detection, and specifically to a method for rapidly detecting antibody content in serum samples using Raman spectroscopy, analyzing the distribution of spectral characteristic peaks, and comparing the performance of different machine learning models in detecting type I and type II infection in negative and positive serum samples, as well as in positive serum samples. Background Art

[0002] Helicobacter pylori is a curved, spiral-shaped, Gram-negative, microaerophilic bacillus and the only bacterium designated a Class I carcinogen by the International Agency for Research on Cancer. It is primarily distributed in the gastric antrum and duodenum, where it can cause chronic inflammation of the gastric mucosa and, in severe cases, even duodenal ulcers and gastric cancer. Therefore, accurate, rapid, and convenient detection of H. pylori infection is crucial for human health. Unless eradicated with medication or eventually developing atrophic gastritis or gastric cancer, H. pylori infection persists for life. Therefore, a positive serological test for H. pylori antibodies should be considered active infection. Enzyme-linked immunosorbent assay (ELISA) is a commonly used clinical serological test, using purified, partially purified, and crude antigens. This method has a sensitivity and specificity approaching 95%. However, due to the significant phenotypic heterogeneity of H. pylori, the preparation of cytotoxic antigens requires the use of a mixture of multiple strains, particularly isolates from the study population. In recent years, surface-enhanced Raman spectroscopy (SERS) combined with machine learning algorithms has been applied to identify a variety of bacterial pathogens with rapid and accurate accuracy. However, there are currently few studies combining SERS with machine learning to detect antibodies produced by Helicobacter pylori infection in serum samples.

[0003] Existing technologies for identifying Helicobacter pylori antibodies in serum mostly use enzyme-linked immunosorbent assay (ELISA) and latex-enhanced turbidimetry. ELISA is characterized by long testing times, complex procedures, and poor reproducibility, while latex-enhanced immunoturbidimetry suffers from low sensitivity. Using artificial intelligence analysis methods such as machine learning can further reduce testing costs and improve efficiency. Summary of the Invention

[0004] The purpose of the present invention is to address the deficiencies of the above-mentioned existing technologies and provide a method for rapid analysis and identification of Helicobacter pylori infection and antibody typing using serum Raman spectroscopy. By combining surface light-enhanced Raman spectroscopy with machine learning algorithms, the problems of cumbersome and time-consuming detection processes are solved.

[0005] The technical solutions of the present invention are as follows:

[0006] A method for rapid analysis and identification of Helicobacter pylori infection and antibody typing using serum Raman spectroscopy comprises the following steps:

[0007] (1) Sample collection: Serum samples of subjects were collected in the clinic;

[0008] (2) Sample identification: The serum sample obtained in step (1) is tested for Helicobacter pylori antibodies using a quantum dot immunofluorescence assay;

[0009] (3) Sample division: The serum samples identified in step (2) were divided according to the results of the three antibodies Urease, cagA, and vacA;

[0010] (4) Sample pretreatment: adding sodium citrate solution to a silver nitrate solution heated to boiling to react with stirring, centrifuging the resulting reaction solution, and resuspending the resulting precipitate in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles. The serum sample identified in step (2) is mixed with the Raman substrate solution to obtain a high-sensitivity surface-enhanced serum Raman sample;

[0011] (5) SERS detection: The high-sensitivity surface-enhanced serum Raman sample obtained in step (4) is subjected to multiple, multi-site signal sampling to obtain the SERS spectrum of each serum sample, and antibody positive and negative databases for data analysis, as well as Helicobacter pylori type I and type II infection SERS databases, are constructed respectively;

[0012] (6) Quality assessment: The serum SERS spectra obtained in step (5) were subjected to repeatability and variability analysis;

[0013] (7) Cluster analysis: The Raman spectral data obtained after evaluation in step (6) were subjected to orthogonal partial least squares discriminant analysis (OPLS-DA) cluster analysis. The differences between patients infected with Helicobacter pylori and the types of infected patients were compared, and abnormal samples were eliminated according to the discreteness of the sample points.

[0014] (8) Data identification: The data after evaluation and abnormal samples in steps (6) and (7) are analyzed using a variety of machine learning algorithms. The Raman spectroscopy data set is divided into a training set, a validation set, and a test set by uniform random sampling. The optimal model parameters are found by grid search and tested using K-fold cross validation, where K is an integer between 1 and 10. The sample data are classified and labeled by the trained classifier and stored in the database.

[0015] (9) Data evaluation: Select test set samples to test model differentiation and prediction ability, use precision, recall rate, ROC curve and confusion matrix to modify the final judgment model, evaluate the performance of different machine learning algorithms, and select the best judgment model, including:

[0016]

[0017]

[0018]

[0019] Where: Precision represents precision, Recall represents recall, and AUC represents the area under the ROC curve, which is considered a performance indicator; TP, FP, TN, and FN represent the number of true positives, false positives, true negatives, and false negatives, respectively. These four cases constitute the confusion matrix.

[0020] (10) Performance verification: Randomly recruit a certain number of subjects to test the performance of the optimal model selected in step (9).

[0021] In the present invention, in step (1), the subject fasts the night before the sampling, and the serum sample is collected as follows:

[0022] (a) The subject sits or lies down with their arm resting on a cushion on the table, exposing the puncture site. Disinfection is performed in a circular pattern from the inside out, with the puncture site as the center. The disinfection area is 5 cm in diameter and is repeated 1 to 3 times. The disinfectant must remain in contact with the skin for at least 30 seconds to take effect. The disinfectant should be allowed to dry naturally before puncturing.

[0023] (b) Hold the subject's arm below the puncture site. Position your thumb 2.5 to 5.0 cm below the puncture site, pulling down on the skin to secure the vein. Avoid contact with the disinfected area. Keeping the needle bevel upward, insert the lancet into the vein at an angle of approximately 30° to the arm. Stop sampling when the blood volume reaches the 4 mL mark.

[0024] In a preferred embodiment, the serum sample collection process is as follows:

[0025] (a) The subject sits with their arm placed on a cushion on the table, exposing the puncture site. Disinfection is performed in a circular pattern from the inside out, with the puncture site as the center. The disinfection area is 5 cm in diameter and disinfected twice. The disinfectant must remain in contact with the skin for at least 30 seconds to take effect. The disinfectant should be allowed to dry naturally before puncturing.

[0026] (b) Hold the subject's arm below the puncture site. Position your thumb 3.5 cm below the puncture site to pull down on the skin to secure the vein, avoiding contact with the sterile area. Keeping the needle bevel upward, insert the lancet into the vein at an angle of approximately 30° to the arm. Stop sampling when the blood volume reaches near the 4 mL mark.

[0027] In the present invention, in step (2), the infection status and typing identification of serum samples are quantitatively detected using a quantum dot immunofluorescence Helicobacter pylori detection kit, and the specific process is as follows:

[0028] (a) taking a coagulated blood sample and centrifuging it at a low speed of 3000 to 5000 rpm / min, preferably 4000 rpm / min, for 5 minutes to obtain a serum sample;

[0029] (b) Take 60-90 μL of serum sample from step (a) and vertically add it to the sample well, preferably 70 μL of serum sample. Insert the test card into the instrument, select automatic timing, complete the test in 15 minutes, and record the test results.

[0030] In the present invention, in step (3), the samples are divided according to the three antibodies of Urease, cagA and vacA. The specific process is as follows:

[0031] (a) The presence or absence of Helicobacter pylori infection or past infection was determined based on whether the Urease value was ≥8;

[0032] (b) The positive samples infected with Helicobacter pylori in step (a) were classified and classified. Samples that met at least one of the following conditions: cytotoxin-associated gene A protein (cagA) ≥ 6 and vacuolating toxin (vacA) ≥ 4 were determined to be infected with Helicobacter pylori type I, while samples with cagA < 6 and vacA < 4 were determined to be infected with Helicobacter pylori type II.

[0033] When preparing a Raman enhancement substrate, silver nitrate is dissolved in water and stirred and heated to boiling, wherein the concentration of silver nitrate in the silver nitrate solution is 0.5 to 1.5 mol / L, preferably 1 mol / L. The reducing agent is sodium citrate, and the sodium citrate is prepared into a sodium citrate solution with a concentration of 0.5 to 1.5 wt%, preferably 1 wt%. The serum sample after centrifugation in step (2) is mixed with an equal amount of the silver nitrate solution.

[0034] In a preferred embodiment, the preparation method of the Raman enhanced substrate is as follows:

[0035] (a) A 1 mmol / L silver nitrate solution was heated to boiling, and 8 mL of a 1 wt% sodium citrate solution was added while stirring. The mixture was stirred at 600-800 rpm for 30-60 min to obtain a negatively charged silver nanoparticle reaction solution.

[0036] (b) centrifuging the silver nanoparticle reaction solution obtained in step (a) at 6000-8000 rpm for 5-10 minutes, discarding the supernatant, and resuspending the resulting precipitate in deionized water to obtain a negatively charged silver nanoparticle solution;

[0037] (c) 10 to 30 μL of the silver nanoparticle solution obtained in step (b) was mixed evenly with an equal amount of serum sample. The mixed sample was then dropped onto a clean silicon wafer to form circular spots, which were then naturally dried in a safety cabinet for testing.

[0038] In a more preferred embodiment, the preparation method of the Raman enhanced substrate is as follows:

[0039] (a) A 1 mmol / L silver nitrate solution was heated to boiling, and 8 mL of a 1 wt% sodium citrate solution was added while stirring. The mixture was stirred at 650 r / min for 40 min to obtain a negatively charged silver nanoparticle reaction solution.

[0040] (b) Centrifuging 1 mL of the silver nanoparticle reaction solution obtained in step (a) at 7000 rpm for 7 min, discarding the supernatant, and resuspending the resulting precipitate in 100 μL of deionized water to obtain a negatively charged silver nanoparticle solution, which serves as the Raman enhancement substrate.

[0041] (c) 20 μL of the silver nanoparticle solution obtained in step (b) was mixed with an equal amount of serum sample. Then, 30 μL of the mixed sample was dropped onto a clean silicon wafer to form circular spots. The solution was dried naturally in a safety cabinet and a SERS fingerprint was obtained.

[0042] Preferably, in step (5), the number of sites in the multi-site Raman spectrum sampling is 100 to 150 sites. In order to improve data balance, it is further preferably 120 sites.

[0043] Preferably, in step (5), the SERS detection is performed using a Cora-100 handheld Raman spectrometer to collect data. The sampling parameters are as follows: the excitation wavelength of the Raman spectrometer is 785 nm, the detector type is a high-sensitivity CCD array, the exposure time is 10 s, and the wavelength scanning range is 400-2300 cm -1 .

[0044] Preferably, SERS detection is performed using a silicon wafer at 520 cm before spectrum acquisition. -1 The Raman peak at is used as the reference peak for wavenumber calibration, and the dark current is deducted within the same integration time.

[0045] For the present invention, in step (6), sample quality assessment is used to check the data quality and spectral differences collected in step (5), which is achieved by averaging the spectrum and spectral deconvolution.

[0046] Preferably, the SERS average Raman spectra of samples of Helicobacter pylori infection-negative and -positive patients, as well as those carrying type I and type II antibodies in positive patients are plotted, and the overall data quality is checked by standard deviation.

[0047] Preferably, the spectral deconvolution bands of the average Raman spectrum are analyzed, and the differences between the SERS spectra of different samples are examined by deconvolution.

[0048] In the present invention, in step (7), data clustering is used to determine whether the sample is within a specified interval and category. The cluster analysis method is specifically selected according to actual needs, and can be principal component analysis, K-means analysis, partial least squares, and orthogonal partial least squares discriminant analysis. OPLS-DA cluster analysis is preferred, and the specific process is as follows:

[0049] (a) The OPLS-DA algorithm was used to analyze the SERS data of samples of Helicobacter pylori infection negative and positive patients, as well as samples carrying type I and type II antibodies in positive patients. The three evaluation indicators R2X, R2Y, and Q2 were used for quantitative analysis.

[0050] (b) Use the OPLS-DA algorithm to check the discreteness of the sample points after clustering, remove abnormal outliers, and cluster again until no abnormal outliers appear.

[0051] For the present invention, in step (8), data identification is divided into three artificial intelligence classification models, consisting of two machine learning models and one deep learning model. These three methods are computer storage, computer processor, and computer program stored in computer memory and executable on the computer processor; the computer memory stores the model optimization parameter range, and the computer processor implements the following steps when executing the computer program:

[0052] (a) All SERS Raman signals after data preprocessing are divided into different data subsets;

[0053] (b) Pre-set different ensemble learning training parameter ranges, fit the model, and select the optimal parameter combination;

[0054] (c) Using the trained ensemble learning model, the SERS spectra of samples of Helicobacter pylori infection-negative and -positive patients, as well as those carrying type I and type II antibodies, were classified and predicted to obtain different model files;

[0055] (d) Evaluate the performance of different machine learning algorithms using a variety of data evaluations and select the best identification model;

[0056] (e) Randomly recruit subjects to verify the application of the trained model in clinical practice.

[0057] Step (a) involves randomly sampling a pre-constructed serum SERS spectral database into a training set, a validation set, and a test set in a ratio of 6:2:2. The training set and validation set are used to train the fitting model and verify model performance, respectively. The test data does not participate in model training or verification, but is input into the training model file as unfamiliar data to test the model's ability to discriminate between unknown data types.

[0058] Step (b) includes: pre-setting the model parameter range that needs to be fitted for different ensemble learning algorithm models, using grid search plus cross-validation to enumerate the model scores of each parameter combination, and selecting the best parameter combination for each model for final model performance comparison.

[0059] Step (c) includes data classification and prediction, using the best model parameter combination fitted in step (b) to perform model training, comparing the performance of the artificial intelligent model in distinguishing between negative and positive Helicobacter pylori infection tasks, and distinguishing between type I and type II antibodies in positive patients on the SERS dataset, selecting the optimal model, and using the trained optimal classifier to classify and annotate the serum SERS spectral data and store them in a database.

[0060] Step (d) includes: selecting test set samples to test the model's predictive ability, using accuracy, precision, recall, F1 score, 5-fold cross-validation and training time to modify the identification model, evaluating the performance of different machine learning algorithms, and selecting the best identification model.

[0061] Step (e) includes: randomly recruiting subjects as double-blind samples, where the subjects are unaware of whether they are infected or have antibodies, collecting SERS signals from the newly included subjects, and using the model trained in step (c) to detect and verify the application of the model in clinical practice.

[0062] Preferably, in step (9), the artificial intelligence algorithm is a convolutional neural network, a random forest, and a decision tree; the spectral data set is divided into a training set, a validation set, and a test set in a ratio of 6:2:2; the training set and the validation set are used for fitting and validation, wherein the validation adopts 5-fold cross validation with K being 5; the test set data is only used to test the model performance, and the grid parameters are used to search for the best model parameter combination.

[0063] The technical solution of the present invention has the following advantages:

[0064] The use of ensemble learning combined with SERS method to classify and identify Helicobacter pylori serum antibodies provides a rapid detection method. The application of Raman spectrometer provides a fast and effective analysis approach, reflecting the application advantages of Raman spectrometer.

[0065] Through integrated learning, a Raman spectrum library of different serum samples is established, and the newly collected fingerprint spectra are directly imported into the library for judgment. The comparison and combination of different machine learning methods and evaluation indicators can better ensure the accuracy of the identification results.

[0066] High accuracy: In the task of distinguishing different gastric juice samples, the optimal model CNN of the present invention scored more than 85% for each indicator; good stability: The 5-fold cross-validation score of CNN was 88.22%, which was close to the scores of various scoring indicators, indicating that the model did not overfit and had strong universality.

[0067] By including randomized subjects, the value of the model in practical clinical application is revealed, and the potential of the model in identifying Helicobacter pylori serum antibody detection is intuitively reflected. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 are the average Raman spectra and deconvolution spectra of negative and positive samples in the present invention;

[0069] Figure 2 This is the clustering result diagram of negative and positive SERS samples from the OPLS-DA data cluster analysis in the present invention;

[0070] Figure 3 The radar chart of scores of different machine learning algorithm models in serum-negative and serum-positive samples and the confusion matrix chart of the optimal model are shown in the figure.

[0071] Figure 4 This is a performance result diagram of the optimal model in the present invention for verifying the negative and positive infection conditions of clinical samples;

[0072] Figure 5 The average Raman spectra and deconvolution spectra of type I and type II samples in the present invention are shown in FIG.

[0073] Figure 6 This is the clustering result diagram of type I and type II SERS samples from the OPLS-DA data cluster analysis in the present invention;

[0074] Figure 7 The radar chart of scores and the confusion matrix of the optimal model for different machine learning algorithm models in serum type I and type II samples in the present invention are shown;

[0075] Figure 8 This is a graph showing the performance results of the optimal model in the present invention for clinical sample type I and type II. DETAILED DESCRIPTION

[0076] Raman spectroscopy is based on the principle of inelastic scattering. That is, when incident light from a laser light source illuminates a substance, it is scattered by the molecules of the substance. A very small portion of the scattered light has a frequency different from the incident light. The change in the scattered light frequency depends on the structural characteristics of the illuminated substance. Different substances produce scattered light of specific frequencies under the same laser irradiation. Therefore, Raman spectroscopy can be used to achieve fast, simple, repeatable and non-destructive detection of material composition.

[0077] Artificial intelligence technology provides an efficient and accurate implementation for Raman spectroscopy-based material composition detection. Existing Raman spectroscopy machine learning algorithms are oriented toward specific substances to be tested. They transform the material identification problem of Raman spectroscopy into a machine learning classification problem. Machine learning models are trained based on standard Raman spectra of known substances, and the trained models are used to accurately identify test samples.

[0078] This specification and claims do not use differences in names as a way to distinguish components, but use differences in components' functions as the criteria for distinction. As mentioned in the description and claims throughout the text, "including" is an open-ended term and should be interpreted as "including but not limited to". "Approximately" means that within an acceptable error range, those skilled in the art can solve the technical problem within a certain error range and basically achieve the technical effect. In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the embodiments so that those skilled in the art can implement it with reference to the text of the specification. The equipment or raw materials used in the embodiments can all be obtained from the market.

[0079] instrument

[0080] Raman spectrometer (Anton Paar, Shanghai), balance (ME104), water purifier (Arium mini), centrifuge (Centrifuge 5430R, Eppendorf, USA), refrigerator (BCD-406WDPD), magnetic stirrer (DF-101S, Tianjin, China), clean bench (SW-CJ-2F), ultra-low temperature refrigerator (994), pH meter (FE28). The models and manufacturers of the instruments used in the implementation case are detailed in Table 1 below.

[0081] Table 1 Experimental instrument information

[0082]

[0083] Drugs and reagents

[0084] Silver nitrate (Sinopharm Chemical Reagent), sodium citrate (Sinopharm Chemical Reagent), sodium chloride (Sinopharm Chemical Reagent), Helicobacter pylori typing test kit (Chongqing Xinsaiya).

[0085] Specific Example 1: Using the method of the present invention to distinguish between positive and negative Helicobacter pylori serum antibody infection 1. Method

[0086] (1) Sample collection: All serum samples were collected from 4 mL of fasting blood from clinically recruited subjects.

[0087] (2) Sample identification:

[0088] The process of negative and positive Helicobacter pylori infection serum antibody test is as follows:

[0089] (a) Take a coagulated blood sample and centrifuge it at 4000 rpm / min for 5 minutes to obtain a serum sample;

[0090] (b) Take 70 μL of serum sample from step (a) and vertically add it to the sample well. Insert the test card into the instrument, select automatic timing, complete the test in 15 minutes, and record the test results.

[0091] (3) Sample classification: The classification of Helicobacter pylori serum antibody negative and positive is determined by Urease. When Urease ≥ 8, it indicates infection or carriage of Helicobacter pylori, and when Urease < 8, it indicates non-infection.

[0092] (4) Sample preprocessing:

[0093] The preparation process of the Raman enhanced substrate is as follows:

[0094] (a) 33.72 mg of silver nitrate was weighed and dissolved in 200 mL of deionized water to obtain a 1 mol / L silver nitrate solution. The solution was heated to boiling using a magnetic stirrer. 8 mL of a 1 wt % sodium citrate solution was then added at once while stirring. The mixture was stirred at 650 rpm for 40 min to obtain a negatively charged silver nanoparticle reaction solution having a diameter of <11 nm.

[0095] (b) 1 mL of the silver nanoparticle reaction solution obtained in step (a) was centrifuged at 7000 rpm for 7 min. The supernatant was discarded and the resulting precipitate was resuspended in 100 μL of deionized water to obtain a Raman enhancement substrate solution of negatively charged silver nanoparticles. The resulting solution was uniformly milky white and stored in the dark at room temperature until use.

[0096] (c) 20 μL of each of the serum antibody negative and positive samples determined in step (3) were mixed evenly with an equal amount of the Raman enhancement substrate solution with negatively charged nanosilver particles prepared in step (b); then 30 μL of the mixed sample was dropped onto a clean silicon wafer to form a circular spot, and the mixture was naturally dried in a safety cabinet. After repeating three times, a Raman spectral fingerprint was obtained by detection.

[0097] (5) SERS detection: The positive and negative serum samples obtained in step (4) are subjected to multiple, multi-site Raman spectral sampling to obtain the corresponding Raman spectra of the negative and positive serum samples, thereby constructing a SERS database of the negative and positive serum samples;

[0098] The Raman spectrum sampling parameters are as follows: the excitation wavelength of the Raman spectrum is 785 nm, the detector type is a high-sensitivity CCD array, the exposure time is 1000 ms, and the wavelength scanning range is 400-2300 cm -1 , output data type: txt, laser intensity: 2, spectrum acquisition time: 10s; before spectrum acquisition, use silicon wafer at 520cm -1 The Raman peak at is used as the reference peak for wavenumber calibration, and the dark current is subtracted at the same integration time. 120 sites are randomly selected for each sample for Raman spectroscopy detection.

[0099] (6) Quality assessment: Perform average spectrum and spectrum deconvolution analysis on the SERS signal obtained in step (5) to check the data quality and differences.

[0100] (7) Cluster analysis: The serum SERS data were input into the OPLS-DA supervised learning algorithm model. By comparing the distribution differences of different serum spectral sample points, a cluster scatter plot of different serum Raman vectors was formed. The clustering results were evaluated using the three indicators R2X, R2Y, and Q2. The clustering results were observed, abnormal outliers were eliminated, and clustering was re-performed until no outliers appeared.

[0101] (8) Data Identification: Three artificial intelligence algorithms, convolutional neural networks, random forests, and decision trees, were used to analyze the SERS signals of negative and positive samples. The SERS signals were divided into training, validation, and test sets, and a classifier was trained. The sample data were assigned and labeled using the trained classifiers and then stored in a database. Before training, all parameters were optimized using grid search. The parameter settings for each algorithm are shown in Table 2 below.

[0102] Table 2 Parameter settings of artificial intelligence algorithm for distinguishing serum negative and positive samples

[0103]

[0104] (9) Data evaluation: Select test set samples to test model differentiation and prediction ability, use precision, recall rate, ROC curve and confusion matrix to modify the final judgment model, evaluate the performance of different machine learning algorithms, and select the best judgment model, including:

[0105]

[0106]

[0107]

[0108] Where: Precision represents precision, Recall represents recall, and AUC represents the area under the ROC curve, which is considered a performance indicator; TP, FP, TN, and FN represent the number of true positives, false positives, true negatives, and false negatives, respectively. These four cases constitute the confusion matrix.

[0109] (10) Performance verification: 20 subjects were randomly recruited as double-blind samples. The subjects were unaware of whether they were infected or had antibodies. The SERS signals of the newly included subjects were collected to test the best training model selected in step (9) to verify the application of the model in clinical practice.

[0110] 2. Results

[0111] Figure 1 Figures (A) and (B) in the middle are the average SERS spectra of serum-negative and serum-positive samples. The shaded areas in the figures represent the standard error bands of the spectra, and the width of the error bands indicates the repeatability of the Raman spectral data. Figures (C) and (D) represent the deconvolution spectra of serum-positive and serum-negative samples. The deconvolution spectra can perform characteristic peak analysis on the average Raman spectra and further refine the differences between the spectral curves.

[0112] Figure 2 The SERS spectra of the two types of serum samples were clustered and analyzed using the OPLS-DA algorithm. However, most of the sample points overlapped, indicating that OPLS-DA could not distinguish the two types of gastric juice samples well. We need to find a better method.

[0113] Figure 3 (A) shows a radar chart of the scores of different machine learning algorithms on SERS data of serum-negative and serum-positive samples. The results show that the convolutional neural network achieves all scores exceeding 85%, outperforming other algorithms in identification. (B) shows the confusion matrix of the optimal machine learning algorithm, the convolutional neural network. The figure details the judgments made by each algorithm on the test set SERS data. The results show that the convolutional neural network achieves an accuracy of 88% for negative and 90% for positive samples, respectively.

[0114] Figure 4 The results of applying the CNN algorithm to clinical samples with real unknown infection conditions were presented. A total of 20 subjects participated in the test. Through the performance of the detection algorithm in 30 SERS spectra of each subject, it was found that CNN was able to achieve identification with an accuracy rate of 80%, indicating that CNN is an efficient identification method for determining the infection status of serum samples.

[0115] Specific Example 1: Using the method of the present invention to distinguish the serum antibody typing of positive patients

[0116] 1. Method

[0117] (1) Sample collection: All serum samples were collected from 4 mL of blood from clinically recruited subjects under fasting conditions.

[0118] (2) Sample identification:

[0119] The process of serum antibody typing test for Helicobacter pylori positive infection is as follows:

[0120] (a) Take a coagulated blood sample and centrifuge it at 4000 rpm / min for 5 minutes to obtain a serum sample;

[0121] (b) Take 70 μL of serum sample from step (a) and vertically add it to the sample well. Insert the test card into the instrument, select automatic timing, complete the test in 15 minutes, and record the test results.

[0122] (3) Sample classification: The classification of Helicobacter pylori serum antibody negative and positive is determined by cagA and vacA. When cagA ≥ 6 and vacA ≥ 4, it is indicated as serum infection type 1, and when cagA ≥ 6 and vacA < 4, it is indicated as serum infection type 2.

[0123] (4) Sample preprocessing:

[0124] The preparation process of the Raman enhanced substrate is as follows:

[0125] (a) 33.72 mg of silver nitrate was weighed and dissolved in 200 mL of deionized water to obtain a 1 mol / L silver nitrate solution. The solution was heated to boiling using a magnetic stirrer. 8 mL of a 1 wt % sodium citrate solution was then added at once while stirring. The mixture was stirred at 650 rpm for 40 min to obtain a negatively charged silver nanoparticle reaction solution having a diameter of <11 nm.

[0126] (b) 1 mL of the silver nanoparticle reaction solution obtained in step (a) was centrifuged at 7000 rpm for 7 min. The supernatant was discarded and the resulting precipitate was resuspended in 100 μL of deionized water to obtain a Raman enhancement substrate solution of negatively charged silver nanoparticles. The resulting solution was uniformly milky white and stored in the dark at room temperature until use.

[0127] (c) 20 μL of each of the serotype 1 and serotype 2 infected samples determined in step (3) were mixed evenly with an equal amount of the Raman enhancement substrate solution containing negatively charged nanosilver particles prepared in step (b); 30 μL of the mixed sample was then dropped onto a clean silicon wafer to form a circular spot, which was then naturally dried in a safety cabinet. After repeating this three times, a Raman spectral fingerprint was obtained by detection.

[0128] (5) SERS detection: The type I and type II samples obtained in step (4) are subjected to multiple, multi-site Raman spectral sampling to obtain the corresponding SERS spectra of the type I and type II samples, thereby constructing a SERS database of the type I and type II samples;

[0129] The Raman spectrum sampling parameters are as follows: the excitation wavelength of the Raman spectrum is 785 nm, the detector type is a high-sensitivity CCD array, the exposure time is 1000 ms, and the wavelength scanning range is 400-2300 cm -1 , output data type: txt, laser intensity: 2, spectrum acquisition time: 10s; before spectrum acquisition, use silicon wafer at 520cm -1 The Raman peak at is used as the reference peak for wavenumber calibration, and the dark current is subtracted at the same integration time. 120 sites are randomly selected for each sample for Raman spectroscopy detection.

[0130] (6) Quality assessment: Perform average spectrum and spectrum deconvolution analysis on the SERS signal obtained in step (5) to check the data quality and differences.

[0131] (7) Cluster analysis: The serum SERS data were input into the OPLS-DA supervised learning algorithm model. By comparing the distribution differences of different gastric fluid spectral sample points, a cluster three-point diagram of different serum Raman vectors was formed. The clustering results were evaluated using the three indicators R2X, R2Y, and Q2. The clustering results were observed, abnormal outliers were eliminated, and clustering was re-performed until no outliers appeared.

[0132] (8) Data Identification: Three artificial intelligence algorithms, convolutional neural networks, random forests, and decision trees, were used to analyze the SERS signals of serotype I and II samples. The SERS signals were divided into training, validation, and test sets, and classifiers were trained. The sample data were assigned and labeled using the trained classifiers and then stored in a database. Before training, all parameters were optimized using grid search. The parameter settings for each algorithm are shown in Table 3 below.

[0133] Table 3 Parameter settings of artificial intelligence algorithm for distinguishing serotype 1 and serotype 2 samples

[0134]

[0135] (9) Data evaluation: Select test set samples to test model differentiation and prediction ability, use precision, recall rate, ROC curve and confusion matrix to modify the final judgment model, evaluate the performance of different machine learning algorithms, and select the best judgment model, including:

[0136]

[0137]

[0138]

[0139] Where: Precision represents precision, Recall represents recall, and AUC represents the area under the ROC curve, which is considered a performance indicator; TP, FP, TN, and FN represent the number of true positives, false positives, true negatives, and false negatives, respectively. These four cases constitute the confusion matrix.

[0140] (10) Performance verification: 20 subjects were randomly recruited as double-blind samples. The subjects were unaware of whether they were infected with type 1 or type 2. The SERS signals of the newly included subjects were collected to test the best training model selected in step (9) to verify the application of the model in clinical practice.

[0141] 2. Results

[0142] Figure 5 Figures (A) and (B) in the middle are the average SERS spectra of serum type I and type II samples. The shaded area in the figure represents the standard error band of the spectrum. The width of the error band indicates the repeatability of the Raman spectrum data. Figures (C) and (D) represent the deconvolution spectra of serum type I and type II samples. The deconvolution spectrum can perform characteristic peak analysis on the average Raman spectrum and further refine the differences between the spectral curves.

[0143] Figure 6The SERS spectra of the two types of serotyping samples were clustered and analyzed using the OPLS-DA algorithm. However, most of the sample points overlapped, indicating that OPLS-DA could not distinguish the two types of gastric juice samples well. We need to find a better method.

[0144] Figure 7 (A) A radar chart shows the scores of different machine learning algorithms on SERS data of serum type I and II samples. The results show that the convolutional neural network achieves all scores exceeding 85%, outperforming other algorithms in identification. (B) A confusion matrix plot shows the results of each algorithm's judgment on the test set SERS data, showing that the convolutional neural network achieves an accuracy of 88% and 90% for type I and type II samples, respectively.

[0145] Figure 8 The results of applying the CNN algorithm to real clinical samples with unknown typing conditions were presented. A total of 20 subjects participated in the test. Through the performance of the detection algorithm in 30 SERS spectra of each subject, it was found that CNN was able to achieve identification with an accuracy rate of 80%, indicating that CNN is an efficient recognition method for determining the typing status of serum samples.

Claims

1. A method for rapid analysis and identification of Helicobacter pylori infection and antibody typing using serum Raman spectroscopy, characterized in that: It includes the following steps: (1) Sample collection: Serum samples of subjects were collected in the clinic; (2) Sample identification: The serum sample obtained in step (1) is tested for Helicobacter pylori antibodies using a quantum dot immunofluorescence assay; (3) Sample division: The serum samples identified in step (2) were divided according to the results of the three antibodies Urease, cagA, and vacA; (4) Sample pretreatment: adding sodium citrate solution to a silver nitrate solution heated to boiling to react with stirring, centrifuging the resulting reaction solution, and resuspending the resulting precipitate in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles. The serum sample identified in step (2) is mixed with the Raman substrate solution to obtain a high-sensitivity surface-enhanced serum Raman sample; (5) SERS detection: The high-sensitivity surface-enhanced serum Raman sample obtained in step (4) is subjected to multiple, multi-site signal sampling to obtain the SERS spectrum of each serum sample, and antibody positive and negative databases for data analysis, as well as Helicobacter pylori type I and type II infection SERS databases, are constructed respectively; (6) Quality assessment: The serum SERS spectra obtained in step (5) were subjected to repeatability and variability analysis; (7) Cluster analysis: The Raman spectral data obtained after evaluation in step (6) were subjected to orthogonal partial least squares discriminant analysis (OPLS-DA) cluster analysis. The differences between patients infected with Helicobacter pylori and the types of infected patients were compared, and abnormal samples were eliminated according to the discreteness of the sample points. (8) Data identification: The data after evaluation and abnormal samples in steps (6) and (7) are analyzed using a variety of machine learning algorithms. The Raman spectroscopy data set is divided into a training set, a validation set, and a test set by uniform random sampling. The optimal model parameters are found by grid search and tested using K-fold cross validation, where K is an integer between 1 and 10. The sample data are classified and labeled by the trained classifier and stored in the database. (9) Data evaluation: Select test set samples to test model differentiation and prediction ability, use precision, recall rate, ROC curve and confusion matrix to modify the final judgment model, evaluate the performance of different machine learning algorithms, and select the best judgment model, including: Where: Precision represents precision, Recall represents recall, and AUC represents the area under the ROC curve, which is considered a performance indicator; TP, FP, TN, and FN represent the number of true positives, false positives, true negatives, and false negatives, respectively. These four cases constitute the confusion matrix. (10) Performance verification: Randomly recruit a certain number of subjects to test the performance of the optimal model selected in step (9).

2. The method for rapid analysis and identification of Helicobacter pylori infection and antibody typing using serum Raman spectroscopy according to claim 1, characterized in that: In step (1), the serum sample collection process is as follows: (a) Place the subject's arm on a cushion on the table, exposing the puncture site. Disinfect the area 5 cm in diameter in a circular motion from the inside out, with the puncture site as the center. Disinfect 1 to 3 times. The disinfectant must remain in contact with the skin for at least 30 seconds before drying naturally. (b) Hold the subject's arm below the puncture site and pull the skin downward with your thumb 2.5 to 5.0 cm below the puncture point to secure the vein, avoiding contact with the disinfected area. Keep the needle bevel upward and insert the lancet into the vein at an angle of approximately 30° to the arm. Stop sampling when the blood volume reaches near the 4 mL mark.

3. The method for rapid analysis and identification of Helicobacter pylori infection and antibody typing using serum Raman spectroscopy according to claim 1, characterized in that: In step (2), the infection status and typing identification of the serum sample are quantitatively detected using a quantum dot immunofluorescence Helicobacter pylori detection kit, and the specific process is as follows: (a) Take a coagulated blood sample and centrifuge it at a low speed of 3000-5000 rpm / min for 5 minutes to obtain a serum sample; (b) Take 60-90 μL of serum sample from step (a) and vertically add it to the sample well. Insert the test card into the instrument, select automatic timing, complete the test in 15 minutes, and record the test results.

4. The method for rapid analysis and identification of Helicobacter pylori infection and antibody typing using serum Raman spectroscopy according to claim 1, characterized in that: In step (3), the serum sample division steps are as follows: (a) The presence or absence of Helicobacter pylori infection or past infection is determined by whether the urease level is ≥8; (b) The positive samples infected with Helicobacter pylori in step (a) are classified and classified. Samples that meet at least one of the following conditions: cytotoxin-associated gene A protein (cagA) ≥ 6 and vacuolating toxin (vacA) ≥ 4 are determined to be infected with Helicobacter pylori type I, while samples with cagA < 6 and vacA < 4 are determined to be infected with Helicobacter pylori type II.

5. The method for rapid analysis and identification of Helicobacter pylori infection and antibody typing using serum Raman spectroscopy according to claim 1, characterized in that: In step (4), the sample preprocessing method is as follows: (a) A 1 mmol / L silver nitrate solution was heated to boiling, and 8 mL of a 1 wt% sodium citrate solution was added while stirring. The mixture was stirred at 600-800 rpm for 30-60 min to obtain a negatively charged silver nanoparticle reaction solution. (b) centrifuging the silver nanoparticle reaction solution obtained in step (a) at 6000-8000 rpm for 5-10 minutes, discarding the supernatant, and resuspending the resulting precipitate in deionized water to obtain a negatively charged silver nanoparticle solution; (c) 10 to 30 μL of the silver nanoparticle solution obtained in step (b) was mixed with an equal amount of serum sample. Then, 30 μL of the mixed sample was dropped onto a clean silicon wafer to form a circular spot, which was then naturally dried in a safety cabinet for detection.

6. The method for rapid analysis and identification of Helicobacter pylori infection and antibody typing using serum Raman spectroscopy according to claim 1, characterized in that: In step (5), the SERS signal detection method adopts multi-site signal acquisition, and 100 to 150 sites are selected for each sample; the spectrum excitation wavelength is 785 nm, the detector type is a high-sensitivity CCD array, the exposure time is 10 s, and the wavelength scanning range is 400-2300 cm-1.

7. The method for rapid analysis and identification of Helicobacter pylori infection and antibody typing using serum Raman spectroscopy according to claim 1, characterized in that: In step (6), the data quality is assessed by plotting the average Raman spectra of Helicobacter pylori infection negative and positive samples and the average Raman spectra of type 1 and type 2 infection positive samples, and calculating the standard error of the average spectra to assess the data quality; the data quality assessment is also performed by calculating the spectral deconvolution of each average Raman spectrum for differential analysis.

8. The method for rapid analysis and identification of Helicobacter pylori infection and antibody typing using serum Raman spectroscopy according to claim 1, characterized in that: In step (7), the quality-assessed data are input into the OPLS-DA algorithm, and cluster analysis is performed on negative and positive samples as well as type I and type II samples respectively. The clustering results are evaluated using three indicators: R2X, R2Y, and Q2. The larger the indicator, the better the model performance, and the difference between R2X and Q2 does not exceed 0.

3. The clustering results are observed, and abnormal outliers are eliminated. After elimination, clustering is performed again until no outliers appear.

9. The method for rapid analysis and identification of Helicobacter pylori infection and antibody typing using serum Raman spectroscopy according to claim 1, characterized in that: In step (8), the data identification machine learning method is a convolutional neural network (CNN), a random forest (RF) and a decision tree (DT); the spectral data set is divided into a training set, a validation set and a test set in a ratio of 6:2:2; the training set and the validation set are used for fitting and validation, wherein the validation adopts 5-fold cross validation and the K is 5.

10. The method for rapid analysis and identification of Helicobacter pylori infection and antibody typing using serum Raman spectroscopy according to claim 1, characterized in that: In step (10), performance verification is used to verify the performance of the model in real clinical samples. 30 to 60 subjects are randomly recruited for blind testing to test the model. The infection status of each subject is unknown, and 30 to 60 SERS spectra are collected for each blind test sample.