A gastric juice-based non-invasive early warning method and system for gastric cancer and helicobacter pylori

By combining Raman spectroscopy and machine learning, the problems of insufficient standardization in gastric fluid sample processing and procedures in existing technologies have been solved, enabling efficient and accurate detection of non-invasive gastric cancer and early warning of Helicobacter pylori.

CN120853974BActive Publication Date: 2026-04-17PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY)
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY)
Filing Date
2025-07-16
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies lack a systematic approach to combining gastric juice Raman spectroscopy with machine learning, particularly in terms of sample bias handling, process standardization, and the generation of multi-dimensional diagnostic indicators. A non-invasive early warning system integrating sample quality control, spectral analysis, model optimization, and clinical validation has not been established.

Method used

Raman spectroscopy information of gastric fluid samples is obtained non-invasively. The data is processed by machine learning algorithms to generate detection adaptation parameters, perform sample quality control and process optimization, perform oversampling to address sample grouping bias, optimize algorithm parameters, and finally generate detection results that include pathological prediction values ​​and confidence levels.

Benefits of technology

It provides a reliable, non-invasive early warning system for gastric cancer and Helicobacter pylori, improving the accuracy and reliability of detection while avoiding the invasiveness and complexity of traditional detection methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853974B_ABST
    Figure CN120853974B_ABST
Patent Text Reader

Abstract

The application provides a gastric juice-based non-invasive early warning method and system for gastric cancer and Helicobacter pylori, and is applied to the technical field of clinical gastric cancer early screening processing. Based on machine learning algorithm information processing, Raman spectrum detection information and sample information are generated to generate detection adaptation parameter information; based on detection process information, operation process information is processed to generate detection prevention parameter information; based on machine learning algorithm information and technical standard information corresponding to the machine learning algorithm information, sample information is processed to generate sample preparation adjustment factors; based on the sample preparation adjustment factors, the operation process information and the step number information corresponding to the operation process information are processed to generate the step classification information and the execution sequence information of the operation process; the detection adaptation index, the step classification information and the execution sequence information of the operation process, and the detection prevention parameter information are input into a target gastric cancer and Helicobacter pylori detection model for processing to generate detection execution result information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of clinical gastric cancer early screening technology, and in particular to a non-invasive early warning method and system for gastric cancer and Helicobacter pylori based on gastric juice. Background Technology

[0002] Gastric cancer is a highly prevalent and deadly cancer worldwide, making early diagnosis crucial. Helicobacter pylori (HP) infection is a significant preventable risk factor, but existing detection methods such as endoscopic biopsy and serological testing have limitations, including invasiveness, insufficient accuracy, and operational complexity. Gastric juice, as a biological fluid that directly reflects the pathological state of the stomach, contains abundant proteins and metabolites and can reflect changes in the gastric mucosa, yet it is not fully utilized. Raman spectroscopy can capture high-resolution molecular information from gastric juice, and combined with machine learning algorithms, it can process high-dimensional data, improving diagnostic efficiency.

[0003] However, current technologies lack a systematic approach to combining gastric juice Raman spectroscopy with machine learning, particularly in areas such as sample bias handling, process standardization, and the generation of multi-dimensional diagnostic indicators. A comprehensive, non-invasive early warning system integrating sample quality control, spectral analysis, model optimization, and clinical validation has not been established, leaving gaps in sample bias handling and process standardization.

[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The purpose of this application is to provide a non-invasive early warning method and system for gastric cancer and Helicobacter pylori based on gastric juice, which at least to some extent overcomes the problems of existing technologies. It acquires Raman spectroscopy detection information and full life-cycle sample information from gastric juice samples non-invasively, and uses machine learning algorithms to process the data and generate detection adaptation parameters. Characteristic peaks such as adenine and glucose are screened, and high-dimensional spectral data is processed using dimensionality reduction methods. Simultaneously, sample quality control and process optimization are performed, addressing oversampling of the HP+ group due to sample grouping bias, optimizing algorithm parameters, and generating sample preparation adjustment factors to optimize the operation process. Finally, the relevant parameters are input into the target model to generate detection results including pathological prediction values ​​and confidence levels, providing a reliable non-invasive solution for early warning of gastric cancer and Helicobacter pylori.

[0006] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part by practice of the invention.

[0007] According to one aspect of this application, a non-invasive early warning method for gastric cancer and Helicobacter pylori based on gastric juice is provided, comprising: acquiring relevant information for gastric cancer and Helicobacter pylori detection, including Raman spectroscopy detection information and gastric juice sample information throughout the entire life cycle; processing the Raman spectroscopy detection information and sample information based on machine learning algorithm information to generate detection adaptation parameter information; processing the operation process information based on detection process information to generate detection prevention parameter information; processing the sample information based on machine learning algorithm information and technical standard information corresponding to the machine learning algorithm information to generate sample preparation adjustment factors; processing the operation process information and step number information corresponding to the operation process information based on the sample preparation adjustment factors to generate step classification information and execution order information of the operation process; inputting the detection adaptation indicators, the step classification information and execution order information of the operation process, and the detection prevention parameter information into a target gastric cancer and Helicobacter pylori detection model for processing to generate detection execution result information.

[0008] Another aspect of this application discloses a non-invasive early warning device for gastric cancer and Helicobacter pylori based on gastric juice, characterized by comprising: an acquisition module for acquiring relevant information on gastric cancer and Helicobacter pylori detection, including Raman spectroscopy detection information and gastric juice sample information throughout the entire life cycle; a processing module for processing Raman spectroscopy detection information and sample information based on machine learning algorithm information to generate detection adaptation parameter information; processing operation process information based on detection process information to generate detection prevention parameter information; processing sample information based on machine learning algorithm information and corresponding technical standard information to generate sample preparation adjustment factors; processing operation process information and corresponding step number information based on sample preparation adjustment factors to generate step classification information and execution order information of the operation process; and inputting the detection adaptation indicators, operation process step classification information and execution order information, and detection prevention parameter information into a target gastric cancer and Helicobacter pylori detection model for processing to generate detection execution result information.

[0009] This application provides a non-invasive early warning method and system for gastric cancer and Helicobacter pylori based on gastric juice. The system uses a server to non-invasively acquire Raman spectroscopy detection information and full life-cycle sample information from gastric juice samples. Machine learning algorithms are then used to process the data and generate appropriate detection parameters. Characteristic peaks such as adenine and glucose are screened, and high-dimensional spectral data is processed using dimensionality reduction methods such as t-SNE and PCA. Simultaneously, sample quality control and process optimization are performed. Oversampling of the HP+ group is addressed to mitigate sample grouping bias, and algorithm parameters are optimized to generate sample preparation adjustment factors for improved operation. Finally, the relevant parameters are input into the target model to generate detection results including pathological prediction values ​​and confidence levels, providing a reliable non-invasive solution for early warning of gastric cancer and Helicobacter pylori.

[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0011] Figure 1 The flowchart illustrates a non-invasive early warning method for gastric cancer and Helicobacter pylori based on gastric juice according to an embodiment of this application.

[0012] Figure 2 This illustration shows a flowchart of a gastric juice diagnosis process based on Raman spectroscopy and machine learning, provided in one embodiment of this application.

[0013] Figure 3 This illustration shows a schematic diagram of the acquisition and analysis of gastric juice Raman spectra according to an embodiment of this application;

[0014] Figure 4 This illustration shows a schematic diagram of the diagnostic performance and potential biomarkers of gastric juice Raman spectroscopy in distinguishing gastric mucosal lesions, provided in an embodiment of this application.

[0015] Figure 5 This illustration shows a schematic diagram of Raman spectroscopy for spectral feature analysis and quantitative analysis of gastric lesions according to an embodiment of this application;

[0016] Figure 6 This illustration shows a schematic diagram of a diagnostic analysis of Helicobacter pylori based on gastric juice Raman spectroscopy according to an embodiment of this application;

[0017] Figure 7 This application illustrates the diagnostic performance and potential biomarkers of gastric juice Raman spectroscopy in the diagnosis of Helicobacter pylori, as provided in one embodiment of the present application.

[0018] Figure 8 The diagram shows a structural schematic of a non-invasive early warning device for gastric cancer and Helicobacter pylori based on gastric juice, provided in an embodiment of this application. Detailed Implementation

[0019] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0020] The following is combined with Figure 1 This application describes a non-invasive early warning method for gastric cancer and Helicobacter pylori based on gastric juice, according to exemplary embodiments of the present application. In one embodiment, the present application also proposes a non-invasive early warning method and system for gastric cancer and Helicobacter pylori based on gastric juice.

[0021] S101, obtain information related to gastric cancer and Helicobacter pylori testing.

[0022] In one implementation, gastric fluid samples were collected from 133 patients and classified according to pathological type into chronic superficial gastritis (CSG, 38 cases), intestinal metaplasia (IM, 28 cases), dysplasia (DYS, 30 cases), and early gastric cancer (EGC, 35 cases). They were also classified according to Helicobacter pylori infection status into HP+ (27 cases) and HP- (104 cases). Gastric fluid was collected non-invasively using a traction suture collector, avoiding the invasiveness of traditional endoscopic biopsy, as described in Example 1 of Article 2.

[0023] Sample quality control involved excluding two invalid samples due to a lack of detectable spectra, ultimately retaining 131 valid samples to ensure the reliability of subsequent spectral analysis. Samples were observed under a bright-field microscope after collection (e.g., ...). Figure 3 A) Although there were no significant differences in the microscopic morphology of the samples from different groups, it was necessary to ensure that there was no severe contamination or hemolysis. Raman spectroscopy was used to acquire the spectra, with a resolution of 3.3 cm⁻¹. -1 Wavenumber range 300-3000cm -1 A total of 3,887 spectra were collected, including 1,137 in the CSG group, 822 in the IM group, 900 in the DYS group, and 1,028 in the EGC group. Gastric fluid was dropped onto a microscope slide, irradiated with a laser, and the scattered light was collected and converted into molecular vibrational information by a spectrometer.

[0024] The original spectrum was preprocessed with baseline correction and normalization, and characteristic peaks (such as adenine at 717 cm⁻¹) were extracted. -1 1003cm of phenylalanine -1 The differences in intensity and relative intensity of the four groups of samples are shown in the figure. Figure 3 As shown in B, although the spectral shapes are similar, the intensity of characteristic peaks (such as glucose in the EGC group at 1343 cm⁻¹) is different. -1 There are significant differences in the peak values.

[0025] like Figure 2 As shown, the Raman spectroscopy generation process is as follows. The matching point illustrates the process of placing a gastric fluid sample on a glass slide, focusing the laser through a laser aperture, and collecting Raman scattered light. This intuitively presents the physical operation process of spectral acquisition, corresponding to placing gastric fluid on a microscope slide for Raman spectroscopy measurement.

[0026] Figure 3 Matching point A: Observe gastric fluid samples under a bright-field microscope to verify the consistency of sample morphology, corresponding to the sample quality control step. Figure 3 B. Matching point: The average Raman spectral curves and standard deviations of the four groups of samples directly demonstrate the differences in spectral characteristics of different pathological types, such as the EGC group at 717 cm⁻¹. -1 The adenine peak intensity at the site was higher than that in the CSG group.

[0027] Figure 5 A represents the assignment of characteristic Raman peaks. Matching points: These indicate characteristic peaks in different pathological types (e.g., increasing intensity of adenine and glucose peaks from CSG to EGC), corresponding to the extraction process of molecular markers in the spectral data. For example, Figure 5 A clearly lists adenine (717cm). -1 ), glucose (1343cm) -1 Assigning peaks such as α to β provides feature input for subsequent machine learning.

[0028] Through systematic sample collection and spectral detection, a dataset containing pathological groupings and molecular characteristics was constructed. Figure 2 , Figure 3 AB Figure 5 A directly verifies the technical details of information acquisition from three aspects: operational procedures, data format, and feature extraction, ensuring the accuracy and reliability of subsequent machine learning analysis. This process not only guarantees the clinical representativeness of the samples but also captures molecular changes in the gastric microenvironment through Raman spectroscopy, laying a data foundation for non-invasive diagnosis.

[0029] S102 uses machine learning algorithms to process Raman spectroscopy detection information and sample information to generate detection adaptation parameter information.

[0030] In one implementation, a stacked fusion technique is introduced based on multi-algorithm performance evaluation. Combining the dimensional characteristics of spectral features, feature peak intensities and sample grouping labels are extracted from preprocessed Raman spectral data and sample pathological information to generate a multi-dimensional diagnostic feature vector set containing algorithm weights. Various machine learning algorithms are employed, covering supervised learning and deep learning. Specifically, Multilayer Perceptron (MLP): captures the nonlinear relationships of spectral features through multilayer neural networks, suitable for complex pattern recognition of high-dimensional spectral data. Artificial Neural Network (ANN): simulates the structure of biological neurons, optimizing weights through backpropagation to improve classification robustness. Gated Recurrent Unit (GRU): as a variant of Recurrent Neural Network (RNN), it can handle sequence correlations in spectral data and is suitable for dynamic feature extraction. Gradient Boosting (GB): reduces model bias through iterative optimization using integrated weak learners (decision trees). Other algorithms, such as KNN, RF, and LSVM, are used to compare the adaptability of different algorithms to spectral features.

[0031] 80% of the Raman spectral data (910 CSG, 658 IM, 720 DYS, and 822 EGC) was used as the training set, and 20% as the validation set to ensure the model's generalization ability. The core evaluation metrics for the model are as follows: Accuracy: The proportion of correctly classified spectra out of the total samples. For example, MLP's accuracy in pathological staging is 78%, meaning that 78% of the spectra were correctly classified into the CSG, IM, DYS, or EGC groups. Sensitivity: The model's ability to correctly identify a certain type of pathological sample. For example, MLP's sensitivity to EGC is 78%, indicating that 78% of EGC spectra were correctly detected. Specificity: The model's ability to correctly exclude non-target pathological samples. MLP's specificity of 93% indicates that it misclassifies non-EGC samples as EGC by only 7%. The outputs of MLP, GRU, ANN, and GB were used as input features for secondary models, and weighted voting was used to generate the final prediction, improving the model's comprehensive ability to identify spectral features.

[0032] Raman spectral characteristic peak selection, adenine (717 cm⁻¹) -1 As a nucleic acid biomarker, its peak intensity gradually increases from the CSG to the EGC stage. Figure 5 B) The average intensity in the EGC group reached 0.22au, a 120% increase compared to the CSG group (0.10au), which can serve as a key indicator of gastric cancer progression. Glucose (1343cm) -1 Phenylalanine (1003 cm⁻¹): Related to cellular metabolic activity, the peak intensity of the EGC group (0.18 au) was significantly higher than that of the CSG group (0.08 au), reflecting enhanced glycolysis in tumor cells. -1 Phenylalanine is an aromatic amino acid, and its peak intensity in the DYS and EGC groups was 0.16au and 0.18au, respectively, which was about 50% higher than that in CSG (0.12au), indicating abnormal protein metabolism.

[0033] Each spectral sample's feature vector contains 10-15 key peak intensity values, such as [717cm]. -1 0.15, 1003cm -1 0.13,1343cm -1 [0.10,...], with dimensions consistent with the number of feature peaks. Weights are dynamically allocated based on the performance of each individual algorithm: MLP: Due to its high specificity for EGC (93%), it accounts for 35% of the weight, primarily contributing to tumor sample identification. GRU: With strong ability to capture spectral sequence features, it has a 25% weight, optimizing the judgment of staging continuity (e.g., progression from IM to DYS). ANN: With good robustness, it has a 20% weight, balancing classification errors across different pathological types. GB: As a representative of ensemble learning, it has a 20% weight, reducing the risk of overfitting by a single algorithm.

[0034] A certain EGC sample has an adenine peak intensity of 0.20 au. After MLP processing, the output probability is 0.90, GRU output is 0.85, ANN output is 0.88, and GB output is 0.82. Therefore, the final EGC probability in the feature vector is: 0.90 × 35% + 0.85 × 25% + 0.88 × 20% + 0.82 × 20% = 0.8675, meaning the probability of this sample being identified as EGC is 86.75%. Single Raman spectra are in the range of 300-3000 cm⁻¹. -1 The range contains approximately 2700 wavenumber points, forming a 2700-dimensional original feature space. Directly inputting this into the algorithm can easily lead to the "curse of dimensionality," increasing computational complexity and the risk of overfitting. t-SNE ( Figure 3 C): Mapping high-dimensional spectra to 2D space allows visualization of clustering trends among different pathological groups, such as the clear separation of EGC and CSG spectra in t-SNE plots. PCA ( Figure 3 D): Principal components were extracted. The first three principal components could explain 85% of the spectral variation. The first principal component was mainly related to the peak intensity of adenine and glucose.

[0035] Figure 3 E shows a scatter plot of the linear discriminant (LD) contribution of Raman spectroscopy data, used to assess the importance and contribution of different Raman spectral features in distinguishing different gastric histopathological types. The horizontal and vertical axes represent different linear discriminant factors (LD1, LD2, etc.), which are feature vectors extracted from the raw Raman spectral data through linear discriminant analysis (LDA) to maximize the differences between groups. Each scatter point in the figure represents the projection of the Raman spectral data of a gastric fluid sample onto the corresponding linear discriminant factor. Different colors or shapes of scatter points represent different gastric histopathological types, such as chronic superficial gastritis (CSG), intestinal metaplasia (IM), dysplasia (DYS), and early gastric cancer (EGC).

[0036] OPLS-DA ( Figure 3 F): Supervised dimensionality reduction algorithms improve the discriminative power of pathological staging by maximizing between-group differences and minimizing within-group variation. For example, the OPLS-DA model achieves a separation degree of Q for IM and DYS. 2 =0.78. The key peak selection criteria are as follows: priority is given to molecular marker peaks associated with gastric cancer, such as adenine (nucleic acid metabolism) and amide III (1235 cm⁻¹). -1 (Changes in protein structure). ANOVA analysis was used to screen for peaks with significant differences (p<0.05) among different pathological groups, such as 717 cm⁻¹. -1 The p-values ​​of the peaks in the CSG, IM, DYS, and EGC groups were all <0.0001.

[0037] The multi-dimensional feature vector set contains 131 samples × 15 key feature peaks, forming a 131 × 15 matrix. Each row corresponds to the feature vector of one sample, and each column corresponds to the intensity distribution of a feature peak across all samples. This set can capture the molecular trajectory of gastric cancer development, such as the linear increase in adenine and glucose peak intensities and the decrease in amide III peak intensity from CSG to EGC, providing quantitative evidence for pathological staging. Through performance comparison and stacking fusion of 12 algorithms, combined with Raman spectroscopy feature peak extraction and high-dimensional data dimensionality reduction, the constructed multi-dimensional diagnostic feature vector set can accurately capture the molecular characteristics of gastric cancer and precancerous lesions. This process not only improves classification accuracy through algorithm integration (reaching 90% in the stacked model) but also ensures the clinical relevance of features through dual biological and statistical screening, laying a data foundation for the generation of subsequent detection adaptation parameters.

[0038] By comparing the feature vector sets of different pathological types and combining the classification error of the confusion matrix localization algorithm, the variation law of spectral features and pathological stages with sample type is determined, generating a diagnostic association feature mapping relationship with pathological differences. This is achieved through the confusion matrix of the stacked model (…). Figure 4 C) Classification error in the localization algorithm: For example, out of 164 spectra in the IM group, 138 were correctly classified (84.1%), while 26 were misclassified as CSG or DYS. This led to the determination of the correlation between spectral characteristics and pathological stages. Specifically, it was found that glucose in the IM group was 1343 cm⁻¹. -1 The peak intensity is higher than CSG but lower than DYS, based on which a mapping relationship of "spectral characteristics - pathological stage" is generated (e.g., 1343 cm⁻¹). -1 Peak intensity > 0.18au tends to be DYS.

[0039] The diagnostic feature mapping relationship is compared with a pre-set spectral standard feature library and pathological staging standards. Combined with ROC curve area under the curve data, feature discrimination evaluation results are generated. The diagnostic mapping relationship is then compared with a pre-set spectral standard library (e.g., adenine in normal gastric mucosa 717 cm⁻¹). -1 The peak intensity standard value (0.12 au) was compared, and the discrimination ability was evaluated by combining the AUC value of the ROC curve. Specifically, the AUC value of MLP in distinguishing DYS+EGC from the control group was 0.98, indicating that 717 cm⁻¹... -1 Peak intensity has a very high degree of distinguishing between intraepithelial neoplasia and other lesions.

[0040] Combining sample pathological type, spectral data quality, and model error, the feature discrimination evaluation results are weighted using algorithm weight distribution to generate a diagnostic influence feature vector representing the contribution of fused features. Algorithm weights (e.g., EGC samples have higher weights than CSG), spectral quality (e.g., spectra with a signal-to-noise ratio > 30 have higher weights), and model error (e.g., the GB algorithm has a 15% misclassification rate for IM) are assigned based on sample pathological type (e.g., MLP weight 40%, GRU weight 30%), and the feature discrimination results are weighted accordingly. Specifically, for the EGC group, glucose 1343 cm⁻¹ -1 The peak intensity discrimination evaluation value was 0.95, and after multiplying by the MLP weight of 40%, its contribution to the diagnostic influence vector was 0.38.

[0041] The feature discrimination results and diagnostic influence vector are normalized (range 0-1), and combined with diagnostic efficiency indicators (such as accuracy, AUC), a feature correlation coefficient (such as 717cm) is generated. -1 The detection adaptation parameters include the correlation coefficient between the peak and EGC (r = 0.85), the pathological predictive value (e.g., EGC probability 95%), and the diagnostic confidence level (0.97).

[0042] S103, based on the detection process information, processes the operation process information to generate detection prevention parameter information.

[0043] In one implementation, feature extraction processing is performed on the detection process information and operation process information to generate sample collection features, spectral acquisition features, data processing features, and model training features. For sample collection features, a non-invasive gastric fluid collection method using a traction wire collector is employed, avoiding the trauma of traditional endoscopic biopsy and improving patient compliance. A total of 133 patient samples were collected, categorized by pathological type into chronic superficial gastritis (CSG, 38 cases), intestinal metaplasia (IM, 28 cases), dysplasia (DYS, 30 cases), and early gastric cancer (EGC, 35 cases), and further categorized by *Helicobacter pylori* infection status into HP+ (27 cases) and HP- (104 cases). Samples were observed using a bright-field microscope (e.g., ...). Figure 3 A) Two samples with missing spectra were excluded, resulting in an effective sample rate of 98.5%, ensuring no hemolysis or severe contamination.

[0044] Spectral acquisition characteristics: Data was acquired using a Raman spectrometer with a resolution of 3.3 cm⁻¹. -1 Wavenumber range 300-3000cm -1 The spectrometer covers the main vibrational ranges of biomolecules. A total of 3887 spectra were collected, including 1137 in the CSG group, 822 in the IM group, 900 in the DYS group, and 1028 in the EGC group, with the number of spectra in each group matching the number of pathological samples. The spectrometer was calibrated for wavelength and intensity before each test to ensure data repeatability.

[0045] Data processing features included t-SNE visualization of high-dimensional spectral data to differentiate clustering trends among different pathological groups; principal component analysis (PCA) extraction, with the first three principal components explaining 85% of spectral variation; and OPLS-DA enhancement of inter-group differences, such as the separation Q between IM and DYS. 2 =0.78. Characteristic peak extraction: Key peaks were selected based on biological significance and statistical significance, such as adenine (717 cm⁻¹). -1 ), glucose (1343cm) -1 Its intensity was significantly higher in the EGC group than in the CSG group (e.g., Figure 5 (As shown in AB). Model training features employed 12 machine learning algorithms, including MLP (78% accuracy), GRU (72%), and ANN (76%), covering supervised learning and deep learning. Spectral data was divided into an 80% training set and a 20% validation set; for example, 105 out of 131 samples were used for training, and 26 were used to validate the model's generalization ability. Stacked model construction involved weighted fusion of four optimal algorithms: MLP, GRU, ANN, and GB, improving pathological staging accuracy to 90% and HP detection accuracy to 96% (e.g., ...). Figure 4 (A, ROC curve shown in Figure 6A).

[0046] The characteristics of sample collection and spectral acquisition were quantitatively analyzed to generate sample collection prevention factors and spectral acquisition prevention factors. The sample collection prevention factors, with the effective sample rate and aseptic operation qualification rate as core indicators, reflect the quality control level of the sample collection process. Two out of 133 samples were excluded due to missing spectra, resulting in an effective sample rate of (133-2) / 133 ≈ 98.5%. The collection process strictly followed aseptic operation procedures, achieving a 100% qualification rate; therefore, the prevention factor = 98.5% × 100% = 0.985. The spectral acquisition prevention factor, combined with the instrument resolution compliance rate and spectral data quality (signal-to-noise ratio), quantifies the reliability of spectral acquisition. Raman spectral resolution was strictly controlled at 3.3 cm⁻¹. -1 (Compliance rate 100%), 95% of the 3887 spectra have a signal-to-noise ratio > 30, which meets the analytical requirements. Therefore, the prevention factor = 100% × 95% = 0.95.

[0047] Quantitative analysis is performed on data processing characteristics and model training characteristics to generate data processing prevention factors and model training prevention factors. The data processing prevention factor assesses the error risk in the data processing stage by multiplying the accuracy of feature peak identification by the effectiveness of the dimensionality reduction algorithm. Key feature peaks (e.g., 717 cm⁻¹) are used to identify key features. -1 The accuracy rate for identifying adenine peaks was 98%, and the Q value of the OPLS-DA dimensionality reduction model was [missing information]. 2 The value is 0.78 (indicating strong predictive ability of the model), therefore the prevention factor = 98% × 0.78 ≈ 0.76.

[0048] The model training prevention factor focuses on key performance indicators (such as AUC) of the stacked model, while also considering the lowest accuracy of a single algorithm, reflecting the stability of model training. The AUC of the stacked model detected by HP is 0.94 (…). Figure 7 A) is directly used as a preventive factor; if the harmonic mean is used, the harmonic mean of the lowest accuracy of the single algorithm (RSVM 41%) and the accuracy of the stacked model (90%) is approximately 57.8%, but in practice, comprehensive indicators such as AUC are given priority, so 0.94 is taken.

[0049] Based on the technical standards of the detection process, sample collection prevention factors, spectral acquisition prevention factors, data processing prevention factors, and model training prevention factors are processed to generate detection prevention assessment information and corresponding weight calculation results. The technical standards of the detection process are as follows: spectral resolution: ≥3cm -1 To ensure the accuracy of molecular vibration peak identification (e.g., 717 cm⁻¹), -1 (Separation degree between adenine peak and adjacent peaks). Sample effectiveness rate ≥95%, avoiding model bias due to too many invalid samples (e.g., ≤6 invalid cases are allowed out of 133 cases). Model AUC ≥0.90, ensuring the diagnostic model's ability to distinguish pathological types (e.g., AUC = 0.94 is acceptable for HP detection stacking model).

[0050] The evaluation information and weight calculation process are as follows: For sample collection, the prevention factor is 0.985 (≥95%, meeting the standard). Because sample quality directly affects subsequent analysis, its weight is allocated to 30%. For spectral acquisition, the prevention factor is 0.95 (≥95%, meeting the standard). Instrument parameter stability is the foundation of data reliability, and its weight is 25%. For data processing, the prevention factor is 0.76 (<0.8, not meeting the standard). Characteristic peak identification and dimensionality reduction effects need optimization, and its weight is adjusted to 20% (lower than the standard weight of 25%). For model training, the prevention factor is 0.94 (≥0.90, meeting the standard). The performance of the stacked model meets clinical needs, and its weight is 25%. Weights are allocated according to the degree of impact of each stage on diagnostic error. Sample collection and spectral acquisition, as the data source, have a higher weight. Data processing, due to not meeting the standard, has its weight appropriately reduced to highlight the priority of optimization.

[0051] The detection and prevention assessment information and weighted calculation results are processed to generate detection and prevention parameter information. The prevention factors at each stage are weighted according to preset weights to reflect the overall risk level of the detection process. Sample collection: 0.985 × 30% = 0.2955; Spectral collection: 0.95 × 25% = 0.2375; Data processing: 0.76 × 20% = 0.152; Model training: 0.94 × 25% = 0.235; Comprehensive detection and prevention parameters: 0.2955 + 0.2375 + 0.152 + 0.235 = 0.923.

[0052] The prevention factors at each stage (e.g., sample collection 0.985, spectral collection 0.95) reflect the quality control level of a single stage; the comprehensive value of 0.923 (close to 1) indicates a low overall process risk, but there is room for optimization in the data processing stage (0.76). Specifically, the data processing prevention factor of 0.76 does not meet the standard (≥0.8), requiring targeted improvement in the accuracy of feature peak identification (e.g., from 98% to 99%) or optimization of the dimensionality reduction algorithm (e.g., improving the Q of OPLS-DA). 2 The value was increased to 0.85 to improve the overall reliability of the test.

[0053] S104, Based on machine learning algorithm information and the corresponding technical standard information, the sample information is processed to generate sample preparation adjustment factors.

[0054] In one implementation, sample information, machine learning algorithm information, and their corresponding technical standard information are extracted and classified to generate sample quality deviation information, sample grouping deviation information, algorithm parameter setting difference information, and technical standard compliance difference information. Regarding the sample quality deviation information, out of 133 samples, 2 were excluded due to spectral deficiencies, resulting in a sample validity rate of 98.5%, and a deviation value of 1 - 98.5% = 1.5%.

[0055] The sample grouping bias information is as follows: HP+ group 27 cases, HP- group 104 cases, group ratio 1:3.85, the bias value from the balanced group (1:1) = |1-3.85| / 1 = 285% Figure 6 A. Differences in the number of spectra.

[0056] The differences in algorithm parameter settings are as follows: the learning rate of the MLP algorithm is set to 0.01 (standard range 0.001-0.1), which is within a reasonable range; however, the kernel function parameter γ of RSVM is 10 (standard recommendation γ≤1), which shows an excessively large parameter deviation.

[0057] The technical standard compliance information is as follows: spectral resolution 3.3 cm. -1 (≥3cm -1 (meeting the standard); however, the sample size of the HP+ group was 27 cases < 50 cases (the standard recommends ≥ 50 cases per group), and the compliance deviation = 1 - 27 / 50 = 46%.

[0058] Based on machine learning algorithm standards, information on sample quality deviations, sample grouping deviations, differences in algorithm parameter settings, and differences in compliance with technical standards is processed to generate deviation type judgment information and deviation severity assessment information. Clinical sample testing requires a valid sample rate of ≥95%, and invalid samples must be removed due to objective factors (such as contamination or missing spectral signals). Of the 133 samples, 2 were excluded due to missing spectral signals, resulting in a valid sample rate of 98.5% (≥95%). However, data incompleteness still exists, constituting a "data integrity deviation." This deviation affects the completeness of the model's input data and needs to be compensated for through resampling or data augmentation.

[0059] Machine learning training requires that the ratio of sample sizes between groups be ≤2:1 to avoid model bias in identifying the minority class. Specifically, the HP+ group had 27 cases and the HP- group had 104 cases, a ratio of 1:3.85 (>2:1), which falls under the category of "sample balance bias". Figure 6 A shows that the number of spectra in the HP+ group is only 26.1% of that in the HP- group. This bias makes the model prone to favoring the majority class (HP-), and needs to be corrected by oversampling or a weighted loss function.

[0060] The recommended range for the RSVM kernel function parameter γ is [0.1, 1]. Values ​​that are too large can easily lead to overfitting. A parameter γ = 10 exceeds the standard range, resulting in a significant decrease in generalization ability, which is considered "algorithm hyperparameter bias." This bias is caused by manually setting parameters outside the optimal range and requires re-optimization through grid search.

[0061] The quantification of severity assessment information is as follows: when the sample quality deviation is mild, the quantification standard is that the proportion of invalid samples <5% is mild, 5%-10% is moderate, and >10% is severe. The proportion of invalid samples = 2 / 133 ≈ 1.5% < 5%, which has a limited impact on the overall data and can be corrected by adding a small number of samples.

[0062] Group bias is considered severe, quantified by a sample size ratio between groups > 3:1 for severe, 2:1-3:1 for moderate, and ≤ 2:1 for mild. The sample size ratio between the HP+ and HP- groups was 1:3.85 > 3:1. Figure 6 In Group A, there were 806 spectra in the HP+ group and 3081 in the HP- group. The sample size of the minority class (HP+) was severely insufficient, which may have reduced the sensitivity of the model to HP infection (e.g., the sensitivity of HP+ was 89%, which was lower than that of HP- at 98%).

[0063] The RSVM parameter bias is moderate. The quantification criteria are: algorithm accuracy more than 20% below average is considered severe, 10%-20% is moderate, and <10% is mild. The average accuracy of the 12 algorithms is approximately 68%, and RSVM accuracy is 41% (27% below the average). However, because RSVM performs poorly in high-dimensional spectral data, and the stacked model does not use this algorithm (MLP / GRU, etc., are preferred), it is classified as having moderate bias. Parameter adjustment is needed, but it is not the highest priority.

[0064] Based on the deviation type and severity assessment information, target data in the sample information are labeled and filtered to generate deviation data filtering results. 27 samples from the HP+ group are labeled as "requiring oversampling" (…). Figure 6 In group A, there were only 806 HP+ spectra and 3081 HP- spectra; the RSVM algorithm parameter γ=10 was marked as "needs adjustment"; the spectral numbers of 2 invalid samples were screened out and excluded from the analysis.

[0065] The bias data screening results are integrated and quantified to generate a sample preparation adjustment factor. The formula for calculating the bias quantification is: Adjustment Factor = 1 - (ω1 × Quality Bias + ω2 × Group Bias + ω3 × Parameter Bias + ω4 × Standard Deviation), where the weights are: ω1 = 0.2, ω2 = 0.4, ω3 = 0.2, ω4 = 0.2. A specific example is as follows: Quality Bias = 1.5% × 0.2 = 0.003; Group Bias = 285% × 0.4 = 1.14; Parameter Bias = (1 - 41%) × 0.2 = 0.118 (quantified as the difference between RSVM accuracy and average accuracy); Standard Deviation = 46% × 0.2 = 0.092. Adjustment Factor = 1 - (0.003 + 1.14 + 0.118 + 0.092) = 1 - 1.353 = -0.353 (the negative value here indicates that the sample balance problem needs to be addressed first due to excessive group bias).

[0066] S105, based on the sample preparation adjustment factor, process the operation process information and the step number information corresponding to the operation process information to generate the step classification information and execution order information of the operation process.

[0067] In one implementation, the operation process information, sample preparation adjustment factors, and their corresponding step number information are extracted and classified to generate step execution difference information, adjustment factor correlation information, and operation effect deviation information. The step execution difference information includes the following: differences in sample collection steps; clinical testing requires a sample validity rate ≥95%, and invalid samples must be recollected within 24 hours to ensure data integrity. Of the 133 samples, 2 were deemed invalid due to missing spectra, but were not recollected in time. The actual valid sample rate was 98.5% (meeting the standard), but there was a delay in process execution. If invalid samples are not replenished, the sample size of the HP+ group may become further insufficient during model training (the original HP+ group only had 27 cases), exacerbating grouping bias.

[0068] The oversampling operation differed between the HP+ group and the model training process. Before training, the HP+ group needed to be oversampled to 50 cases using the SMOTE algorithm to ensure a group-to-group ratio of ≤2:1. In practice, oversampling was not performed before model training, and the HP+ group remained at 27 cases. Figure 6 A shows 806 spectra in the HP+ group and 3081 spectra in the HP- group (ratio 1:3.85), resulting in the stacked model having a sensitivity of only 89% to HP+.

[0069] The adjustment factor information includes the following: The formula for calculating the sample preparation adjustment factor is: Adjustment Factor = 1 - (Quality Bias × 0.2 + Grouping Bias × 0.4 + Parametric Bias × 0.2 + Standard Deviation × 0.2). For example, Quality Bias: Invalid sample rate 1.5% → 1.5% × 0.2 = 0.003; Grouping Bias: (104 / 27 - 1) = 285% → 285% × 0.4 = 1.14; Parametric Bias: (1 - 41%) = 59% → 59% × 0.2 = 0.118 (RSVM accuracy 41%). Standard Deviation is HP + Group Sample Size Gap 46% → 46% × 0.2 = 0.092. Adjustment Factor = 1 - (0.003 + 1.14 + 0.118 + 0.092) = -0.353 (A negative value indicates that grouping bias is the main problem). Figure 6 In A, the number of spectra in the HP+ group is significantly lower than that in the HP- group, which intuitively supports the necessity of oversampling adjustment.

[0070] The operational performance deviation information includes the following: RSVM algorithm parameter deviation; the RSVM kernel function parameter γ = 10 (standard recommendation γ ≤ 1), leading to model overfitting and an accuracy of only 41%, 27% lower than the average accuracy (68%). If this algorithm is incorporated into a stacked model, it may lower the overall performance (the actual stacked model does not use RSVM; MLP / GRU, etc., are preferred). Other potential deviation examples are as follows: spectral acquisition parameter deviation: if the spectral resolution drops to 2.5cm in a certain detection... -1 (Standard ≥3cm) -1 This could potentially lead to a height of 717cm. -1The adenine peak overlaps with adjacent peaks, affecting feature extraction.

[0071] Based on the rules governing the effect of sample preparation adjustment factors, the information on differences in step execution, correlations between adjustment factors, and deviations in operational effects are processed to generate adjustment type judgment information and adjustment degree assessment information. The adjustment type judgment information includes the following: Sample resampling: Data integrity adjustment. When the sample validity rate is <95% or there are missing spectral signals, resampling is required to ensure data integrity. Of the 133 samples, 2 were invalid due to missing spectra. Although the valid sample rate of 98.5% met the standard, the missing data may affect feature statistics; therefore, it is classified as "data integrity adjustment," and supplementary collection is required to eliminate potential bias.

[0072] HP+ group oversampling: Sample balance adjustment. If the sample size ratio between groups is >2:1, it is necessary to balance the groups through oversampling / undersampling to avoid model bias in identifying the minority class. The HP+ group has 27 cases and the HP- group has 104 cases, with a ratio of 1:3.85 (>3:1), which is severely imbalanced. It is necessary to oversample the HP+ group to 50 cases using the SMOTE algorithm, which is classified as "sample balance adjustment".

[0073] RSVM Parameter Optimization: Algorithm parameter adjustment. When algorithm parameters deviate from the standard range and cause a significant performance drop, the parameters need to be re-optimized. The RSVM kernel function parameter γ = 10 (the standard recommendation is γ ≤ 1), which is outside the optimal range, resulting in an accuracy of 41%, which is 27% lower than the average accuracy of 12 algorithms (68%). This is classified as "algorithm parameter adjustment" and the γ value needs to be reset (e.g., 0.5-1.0) through grid search.

[0074] Furthermore, the quantitative standards for assessing the degree of adjustment are as follows: Sample resampling: Mild, the assessment standard is that the proportion of invalid samples <5% is mild, 5%-10% is moderate, and >10% is severe. The quantitative basis is that the proportion of invalid samples = 2 / 133 ≈ 1.5% < 5%, which has a limited impact on the overall data, and the cost of resampling is low (only 2 cases need to be added), so it is assessed as mild.

[0075] HP+ group oversampling: Severe. The evaluation criteria are: a sample size ratio between groups >3:1 is severe, 2:1-3:1 is moderate, and ≤2:1 is mild. The quantification is based on the sample size ratio of HP+ to HP- groups being 1:3.85 > 3:1 (HP+ group has 806 spectra, HP- group has 3081 spectra). The minority class sample size is severely insufficient, which may lead to a significant decrease in the model's sensitivity to HP+ detection (HP+ sensitivity is 89%, lower than HP-'s 98%), hence the assessment as severe.

[0076] RSVM parameter optimization: Moderate. The evaluation criteria are: algorithm accuracy more than 20% below average is severe, 10%-20% is moderate, and <10% is mild. The quantification is based on RSVM's accuracy of 41%, which is 27% lower than the average accuracy of 68% of 12 algorithms (>20%). However, since RSVM performs poorly in high-dimensional spectral data and the stacked model does not use this algorithm (MLP / GRU, etc. are preferred), its bias has a controllable impact on overall performance, hence the evaluation is moderate.

[0077] Based on the adjustment type judgment information and adjustment degree evaluation information, the target data in the operation process information are marked and filtered to generate adjustment data filtering results. 27 samples in the HP+ group are marked as "need to be oversampled to 50 cases"; the RSVM parameter γ=10 in step 3 is marked as "to be optimized"; the acquisition records of 2 invalid samples (step 1) are filtered out, and the spectral acquisition needs to be performed again.

[0078] The adjusted data filtering results are integrated and quantified to generate step classification information and execution order information for the operation process. The logic and weight allocation of the step classification information are processed as follows:

[0079] Preprocessing (weight 40%) ensures the quality of the raw data and lays the foundation for subsequent analysis.

[0080] Sample recollection (step 1): Recollect the 2 invalid samples to ensure that the sample validity rate reaches 100% and eliminate data integrity bias.

[0081] Spectral calibration (step 2): Perform wavelength and intensity calibration on the Raman spectrometer (perform before each detection) to ensure a resolution of 3.3 cm⁻¹. -1 Meets standards Figure 2 A) To avoid characteristic peak identification errors (e.g., 717cm) caused by instrument drift. -1 (Adenine peak shift). Quality control at the data source directly affects all subsequent analyses, therefore it is assigned the highest weight.

[0082] Analysis category (weight 35%): Optimize data processing and model training processes to improve diagnostic accuracy.

[0083] HP+ group oversampling (step 4): Use the SMOTE algorithm to oversample the HP+ group from 27 cases to 50 cases, balance the group ratio to 1:2, and solve the insufficient sensitivity caused by sample imbalance (original HP+ sensitivity 89%).

[0084] RSVM parameter optimization (step 3): Adjust the kernel function parameter γ from 10 to 0.5-1.0, and retrain using grid search to improve the algorithm's accuracy (original RSVM accuracy 41%). The analysis stage directly affects model performance, but depends on the quality of preprocessed data, so its weight is secondary.

[0085] Evaluation category (weight 25%): Verify the diagnostic efficacy of the model and output the final detection results.

[0086] Model validation (step 6): Using ROC curves ( Figure 4 A) Confusion matrix ( Figure 4 C) Evaluate the performance of the stacked model, calculate AUC, accuracy, and other metrics, and ensure that the diagnostic confidence level is ≥0.95. The evaluation stage is at the end of the process and depends on the accuracy of preceding steps; therefore, it has the lowest weight but is indispensable.

[0087] Preprocessing steps were prioritized, as sample resampling and spectral calibration are fundamental to data reliability; skipping these steps could lead to deviations in the entire subsequent analysis process. Specifically, gastric fluid from two invalid samples was resampled using a non-invasive traction wire collector; the spectrometer was calibrated, and its resolution of 3.3 cm⁻¹ was verified using standard samples. -1 Meet the standard and ensure 717cm -1 Characteristic peaks such as peaks can be accurately identified.

[0088] Next, the analysis was completed, and the model parameters were optimized based on high-quality data to avoid "garbage in, garbage out". SMOTE oversampling was performed on the 27 samples in the HP+ group to generate 23 virtual samples, bringing the total sample size to 50 and balancing the HP+:HP- ratio to 1:2.08. The RSVM algorithm was optimized by setting γ∈[0.1,1.0] for grid search and selecting the parameter combination with the highest accuracy (e.g., accuracy increased to 65% when γ=0.8).

[0089] Finally, an evaluation process is performed to ensure that the performance of the pre-adjusted model can be quantified through clinical validation metrics. The adjusted model is then used to predict on the validation set data, generating ROC curves. Figure 4 A), the HP detection AUC value was improved from 0.94 to 0.97; the output confusion matrix ( Figure 4 C) Verification showed that HP+ sensitivity increased from 89% to 95%, while specificity remained at 98%.

[0090] S106: Input the detection adaptation indicators, operation procedure step classification information and execution sequence information, and detection prevention parameter information into the target gastric cancer and Helicobacter pylori detection model for processing, and generate detection execution result information.

[0091] In one implementation, pathological prediction values ​​from the detection adaptation indicators, weight parameters from the step classification information of the operational procedure, and abnormal operation markers from the detection prevention parameters are extracted and statistically analyzed to generate diagnostic prediction consistency information, step weight distribution characteristic information, and abnormal operation location trend information. Using gastroscopy biopsy pathology results as the gold standard, the prediction results of the stacked model are compared with clinical diagnoses. The stacked model's prediction performance for early gastric cancer (EGC): out of 206 EGC samples, 196 were correctly classified. The calculation method is as follows: Normalizing the consistency rate to the [0,1] interval yields 0.95, which reflects the degree of agreement between the model prediction and the clinical diagnosis.

[0092] Chronic superficial gastritis (CSG): 219 out of 228 cases were correctly classified, with a consistency rate of 96%; intestinal metaplasia (IM): 138 out of 164 cases were correctly classified, with a consistency rate of 84%; dysplasia (DYS): 151 out of 180 cases were correctly classified, with a consistency rate of 84%.

[0093] The weight distribution characteristics of the steps are as follows: Preprocessing (40%): Steps such as sample resampling and spectral calibration directly affect the quality of the original data and have the highest weight. Sample resampling resolves the issue of two invalid samples, and spectral calibration ensures a resolution of 3.3 cm. -1 Achieved the target. Analysis (35%): Oversampling and parameter optimization improve model performance; these steps have the next highest weight. In the HP+ group, oversampling was increased to 50 cases, and the RSVM parameter γ was adjusted from 10 to 0.8. Evaluation (25%): Model validation relies on previous steps; this step has the lowest weight but is indispensable. The HP detection AUC was verified to be 0.94 using ROC curves.

[0094] The high weighting of the preprocessing class reflects the principle of "data quality first" and avoids model bias due to sample defects; the weighting of the analysis class reflects the impact of algorithm optimization on diagnostic accuracy; and the weighting of the evaluation class ensures the verifiability of the results and complies with clinical diagnostic standards.

[0095] The abnormal operation location trend information includes the following: Operations that did not meet the standards, such as incomplete oversampling in the HP+ group: target sample size 50 cases, actual 27 cases, a gap of 46%; RSVM parameter γ=10 exceeding the standard range (γ≤1), resulting in an accuracy of 41%. Trend analysis shows that abnormalities in analysis steps accounted for 60% (oversampling, parameter optimization), preprocessing steps accounted for 20% (delayed sample resampling), and evaluation steps accounted for 20% (model validation not covering all pathological types). Imbalanced samples in the HP+ group resulted in the model's sensitivity to HP+ being only 89%, lower than HP-'s 98%; RSVM parameter bias made this algorithm have the lowest accuracy among the 12 models, and including it in a stacked model might further reduce overall performance.

[0096] The diagnostic prediction consistency information, step weight distribution characteristics, and abnormal operation location trend information are processed to generate detection fusion confidence information and diagnostic optimization direction information. The calculation logic and parameter decomposition of the detection fusion confidence information are as follows: Diagnostic consistency (0.95): The prediction accuracy of EGC based on the stacked model (95%), reflecting the model's basic identification ability of pathological types, with a weight allocation of 0.4 (the proportion of the impact of preprocessing steps). Reliability of analysis steps (weight 0.35): The completion degree of operations such as oversampling and parameter optimization is evaluated as 0.9 (assuming 70% of the optimization task has been completed); Reliability of evaluation steps (weight 0.25): The completeness of the model validation process is evaluated as 0.9 (e.g., 90% of pathological types have been covered). Impact of abnormal operations (-0.1): The negative correction of confidence to abnormalities such as incomplete oversampling in the HP+ group and RSVM parameter deviation, based on the severity of the abnormal operation (0.1 deducted for severe deviation, 0.05 deducted for moderate deviation).

[0097] The formula for calculating the fusion confidence score is: Fusion Confidence Score = 0.95 × 0.4 + 0.35 × 0.9 + 0.25 × 0.9 - 0.1 = 0.915. Specifically, the contribution of preprocessing is 0.95 × 0.4 = 0.38; the contribution of analysis is 0.35 × 0.9 = 0.315; the contribution of evaluation is 0.25 × 0.9 = 0.225; the anomaly correction is -0.1; and the final result is 0.915 (out of 1.0), indicating that the overall detection process has high reliability, but there is still room for optimization.

[0098] The priority order for diagnostic optimization information is as follows: Severe anomalies take precedence: Insufficient sample size in the HP+ group (group ratio 1:3.85) constitutes severe bias, directly impacting model sensitivity (HP+ sensitivity 89%), and requires priority handling. Moderate anomalies are secondary: RSVM parameter bias (γ=10) results in an algorithm accuracy of 41%, which, although not included in the stacked model, may interfere with the feature extraction process. Final evaluation and validation: The model's AUC target is increased to 0.97 (current HP detection AUC = 0.94), and it needs to pass validation set testing after parameter optimization.

[0099] The specific optimization measures are as follows: HP+ group oversampling was implemented by generating 23 virtual samples using the SMOTE algorithm, bringing the total HP+ group sample size to 50, balancing the group ratio to 1:2.08; HP+ detection sensitivity was improved to 95% (referencing the performance of the optimized stacked model). RSVM parameters were optimized, with γ∈[0.1,1.0], and the optimal value was selected through grid search (e.g., accuracy improved to 65% when γ=0.8). The RSVM model was retrained, and the classification performance of IM samples before and after optimization was compared (original misclassification rate 16%, target ≤10%). Model AUC validation was performed as follows: the stacked model was retrained using the optimized sample set and parameters, and ROC curves were generated for the 20% validation set; the target AUC for EGC detection was ≥0.97, and for HP detection, it was ≥0.97. Fusion confidence was quantified to provide data support for optimization direction: priority was given to addressing severe anomalies such as sample imbalance, followed by adjusting algorithm parameters, and finally, validation was conducted to ensure model performance.

[0100] Based on the confidence information and diagnostic optimization direction information from the detection fusion, key data in the detection adaptation indicators, operational procedures, and detection prevention parameters are labeled and filtered to generate key data screening results for fusion processing. Samples with EGC prediction values ​​<90% are labeled, with 90% as the cutoff for pathological prediction values; samples below this value are considered to have low prediction reliability and require close attention. In the stacked model's prediction results for EGC, 10 out of 206 samples had a prediction probability <90% (…). Figure 4 The 10 misclassified cases in the C confusion matrix correspond to spectral numbers such as S_EGC_03 and S_EGC_17. The characteristic peak intensities of these samples may be abnormal, such as adenine at 717 cm⁻¹. -1 The peak intensity (0.15au) was lower than the mean of the EGC group (0.22au), causing the model to misclassify it as DYS or IM. Such samples may be "borderline" cases (e.g., EGC and DYS are difficult to distinguish), requiring further clinical and pathological examination or additional testing (such as immunohistochemistry).

[0101] Twenty-seven samples from the HP+ group that were not oversampled were screened. Of these, 50 samples did not reach the oversampling target because the original HP+ sample size was 27. The lack of oversampling resulted in a model sensitivity of only 89% for HP+, lower than the 98% for HP-, posing a risk of missed diagnoses. The phosphatidylinositol content of these samples was 415 cm⁻¹. -1 Although the mean peak intensity (0.12au) was higher than that of the HP- group (0.08au), the statistical bias was large due to insufficient sample size. SMOTE oversampling was performed on 27 samples to generate 23 virtual samples, bringing the total sample size of the HP+ group to 50, thus balancing the group ratio to 1:2.08.

[0102] The RSVM kernel function parameter γ=10 was selected, but this exceeded the standard range (γ≤1), leading to overfitting and an accuracy of only 41%. Under this parameter configuration, the model's performance on phenylalanine was 1003 cm⁻¹. -1 The peak weight was too high, mistakenly classifying some IM samples (peak intensity 0.14au) as EGC (peak intensity 0.18au). By resetting γ∈[0.1,1.0] through grid search, the accuracy can be improved to 65% when γ=0.8, reducing the oversensitivity to high-dimensional features.

[0103] The key data screening results from the fusion processing are integrated and quantified to generate the final detection execution results. Early gastric cancer (EGC) diagnosis, with a prediction probability of 95%: The pathological prediction value of the stacked model for EGC is based on confusion matrix analysis of 206 samples, correctly classifying 196 cases. The calculation method is as follows: This probability reflects the model's ability to identify EGC samples and shows high consistency with clinical pathological diagnoses. The confidence level is 0.97, taking into account the reliability of preprocessing steps (sample resampling, spectral calibration weight 40%), optimization of analytical steps (oversampling, parameter adjustment weight 35%), and evaluation validation (AUC = 0.96). Figure 4 B) Confidence indices are generated through weighted calculations to quantify the reliability of diagnostic results.

[0104] Diagnosis of Helicobacter pylori infection (HP+), with a predictive probability of 89%, based on the stacking model prediction results of 27 samples from the original HP+ group. Figure 7 B) Correctly classified 144 / 161 cases, with a sensitivity of 89%, meaning the model identified HP+ samples with an 89% probability. Specificity was 98%, with a 98% correct exclusion rate for HP- samples, indicating that the model misclassified HP- samples as HP+ with a probability of only 2%, demonstrating high clinical specificity.

[0105] The optimized metrics are as follows: Oversampling effect of HP+ group: SMOTE oversampling was performed on the original 27 HP+ samples to generate 23 virtual samples, which optimized the ratio between groups from 1:3.85 (HP+:HP-=27:104) to 1:2.08 (50:104), thus solving the sample imbalance problem.

[0106] The AUC increased from 0.94 to 0.97. After oversampling and retraining the stacked model, the area under the ROC curve (AUC) for HP detection increased from 0.94 to 0.97, indicating that the model's ability to distinguish between HP+ and HP- was significantly enhanced, approaching the "excellent" standard for clinical diagnosis (AUC ≥ 0.95).

[0107] RSVM algorithm parameter optimization improved accuracy from 41% to 65%: The RSVM kernel function parameter γ was adjusted from 10 to 0.8 (standard range [0.1, 1.0]), and the model was retrained using grid search. After optimization, the classification accuracy for IM samples improved from 41% to 65%, reducing overfitting caused by parameter bias, especially improving the accuracy for phenylalanine 1003cm². -1 Stability of peak and other features identification.

[0108] This application acquires Raman spectroscopy detection information and gastric fluid sample information throughout the entire life cycle; uses machine learning to process the above information to generate detection adaptation parameters; generates prevention parameters based on the detection process and operation procedures; combines machine learning and technical standards to generate sample preparation adjustment factors, thereby optimizing the step classification and execution order of the operation process; and finally inputs the relevant parameters into the target model to generate detection results.

[0109] The system employs a multi-algorithm stacking and fusion technology, encompassing 12 algorithms including MLP and GRU. 80% of the data is used for training and 20% for validation, achieving a pathological staging accuracy of 90% and an HP detection accuracy of 96%. It processes high-dimensional spectral data by screening for characteristic peaks such as adenine and combining dimensionality reduction methods like t-SNE and PCA. Simultaneously, sample quality control and process optimization are implemented, such as balancing oversampled data from the HP+ group and optimizing algorithm parameters to improve model reliability. This provides a non-invasive solution for early warning of gastric cancer and Helicobacter pylori.

[0110] In one implementation, such as Figure 8 As shown, this application also provides a non-invasive early warning device for gastric cancer and Helicobacter pylori based on gastric juice, comprising:

[0111] The acquisition module 801 is used to acquire information related to gastric cancer and Helicobacter pylori detection, including Raman spectroscopy detection information and gastric fluid sample information throughout the entire life cycle;

[0112] The processing module 802 is used to process Raman spectroscopy detection information and sample information based on machine learning algorithm information to generate detection adaptation parameter information; process operation process information based on detection process information to generate detection prevention parameter information; process sample information based on machine learning algorithm information and corresponding technical standard information to generate sample preparation adjustment factors; process operation process information and corresponding step number information based on sample preparation adjustment factors to generate operation process step classification information and execution order information; and input the detection adaptation indicators, operation process step classification information and execution order information, and detection prevention parameter information into the target gastric cancer and Helicobacter pylori detection model for processing to generate detection execution result information.

[0113] While this application discloses the above information, it is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of this application; therefore, the scope of protection of this application shall be determined by the scope defined in the claims.

Claims

1. A non-invasive early warning method for gastric cancer and Helicobacter pylori based on gastric juice, characterized in that, include: Obtain relevant information on gastric cancer and Helicobacter pylori detection, including Raman spectroscopy detection information and gastric fluid sample information throughout the entire life cycle; Based on machine learning algorithms, Raman spectroscopy detection information and sample information are processed to generate detection adaptation parameter information; Based on the detection process information, the operation process information is processed to generate detection and prevention parameter information; The sample information is processed based on machine learning algorithm information and corresponding technical standard information to generate sample preparation adjustment factors. This process includes: extracting and classifying sample information, machine learning algorithm information, and their corresponding technical standard information to generate sample quality deviation information, sample grouping deviation information, algorithm parameter setting difference information, and technical standard compliance difference information; processing the sample quality deviation information, sample grouping deviation information, algorithm parameter setting difference information, and technical standard compliance difference information based on machine learning algorithm technical standards to generate deviation type judgment information and deviation severity assessment information; labeling and filtering target data in the sample information based on the deviation type judgment information and deviation severity assessment information to generate deviation data filtering results; and integrating and quantifying the deviation data filtering results to generate sample preparation adjustment factors. Based on the sample preparation adjustment factor, the operation process information and the step number information corresponding to the operation process information are processed to generate step classification information and execution order information of the operation process. The detection adaptation indicators, the step classification information and execution sequence information of the operation process, and the detection prevention parameter information are input into the target gastric cancer and Helicobacter pylori detection model for processing, and the detection execution result information is generated.

2. The method as described in claim 1, characterized in that, Based on machine learning algorithms, Raman spectroscopy detection information and sample information are processed to generate detection adaptation parameter information, including: Based on multi-algorithm performance evaluation, stacking fusion technology is introduced. Combining the dimensional characteristics of spectral features, feature peak intensities and sample grouping labels are extracted from preprocessed Raman spectral data and sample pathological information to generate a multi-dimensional diagnostic feature vector set containing algorithm weights. By comparing the feature vector sets of different pathological types and combining the classification error of the confusion matrix localization algorithm, the variation law of spectral features and pathological stages with sample type is determined, and a diagnostic association feature mapping relationship with pathological differences is generated. The diagnostic correlation feature mapping relationship is compared with the preset spectral standard feature library and pathological staging standard, and combined with the area under the ROC curve data to generate feature discrimination evaluation results; By combining the pathological type of the sample, the quality of spectral data and the model error, the feature discrimination evaluation results are weighted using the algorithm weight distribution to generate a diagnostic influence feature vector with fused feature contribution. The feature discrimination evaluation results and diagnostic influence feature vectors are normalized and integrated, and combined with the diagnostic efficiency index, to generate detection adaptation parameters that include feature correlation coefficients, pathological prediction values ​​and diagnostic confidence.

3. The method as described in claim 1, characterized in that, Based on the detection process information, the operation process information is processed to generate detection and prevention parameter information, including: Feature extraction processing is performed on the detection process information and operation process information to generate sample collection features, spectral collection features, data processing features, and model training features; Quantitative analysis and processing of sample collection characteristics and spectral collection characteristics are performed to generate sample collection prevention factors and spectral collection prevention factors. Quantitative analysis and processing of data processing characteristics and model training characteristics are performed to generate data processing prevention factors and model training prevention factors. Based on the technical standards of the detection process, the sample collection prevention factors, spectral acquisition prevention factors, data processing prevention factors, and model training prevention factors are processed to generate detection prevention assessment information and corresponding weight calculation results. The detection and prevention assessment information and weight calculation results are processed to generate detection and prevention parameter information.

4. The method as described in claim 1, characterized in that, Based on the sample preparation adjustment factor, the operation process information and the corresponding step number information are processed to generate step classification information and execution order information of the operation process, including: The operation process information, sample preparation adjustment factors and their corresponding step number information are extracted and classified to generate step execution difference information, adjustment factor correlation information and operation effect deviation information. Based on the rules governing the role of adjustment factors in sample preparation, the information on differences in step execution, correlation information of adjustment factors, and deviation information of operational effects are processed to generate information on adjustment type and assessment of adjustment degree. Based on the adjustment type judgment information and adjustment degree assessment information, the target data in the operation process information is marked and filtered to generate adjustment data filtering results; The results of the adjusted data filtering are integrated and quantified to generate step classification information and execution sequence information for the operation process.

5. The method as described in claim 4, characterized in that, The detection adaptation indicators, operational procedure step classification information and execution sequence information, and detection prevention parameter information are input into the target gastric cancer and Helicobacter pylori detection model for processing, generating detection execution result information, including: The pathological prediction values ​​in the detection adaptation indicators, the weight parameters in the step classification information of the operation process, and the abnormal operation markers in the detection prevention parameters are extracted and statistically analyzed to generate diagnostic prediction consistency information, step weight distribution characteristic information, and abnormal operation location trend information. The diagnostic prediction consistency information, step weight distribution feature information, and abnormal operation location trend information are processed to generate detection fusion confidence information and diagnostic optimization direction information. Based on the detection fusion confidence information and diagnostic optimization direction information, key data in the detection adaptation indicators, operation process steps and detection prevention parameters are marked and filtered to generate key data filtering results for fusion processing. The key data screening results of the fusion processing are integrated and quantified to generate the final detection execution result information.

6. A non-invasive early warning device for gastric cancer and Helicobacter pylori based on gastric juice, characterized in that, The apparatus for implementing the method of claim 1 includes: The acquisition module is used to acquire information related to gastric cancer and Helicobacter pylori detection, including Raman spectroscopy detection information and gastric fluid sample information throughout the entire life cycle; The processing module is used to process Raman spectroscopy detection information and sample information based on machine learning algorithms to generate detection adaptation parameter information; process operation process information based on detection process information to generate detection prevention parameter information; process sample information based on machine learning algorithm information and corresponding technical standard information to generate sample preparation adjustment factors; process operation process information and corresponding step number information based on sample preparation adjustment factors to generate operation process step classification information and execution order information; and input the detection adaptation indicators, operation process step classification information and execution order information, and detection prevention parameter information into the target gastric cancer and Helicobacter pylori detection model for processing to generate detection execution result information.

7. An electronic device, characterized in that, include: First processor; and memory for storing executable instructions of the first processor; The first processor is configured to execute the non-invasive early warning method for gastric cancer and Helicobacter pylori based on gastric juice according to any one of claims 1 to 5 by executing the executable instructions.

Citation Information

Patent Citations

  • Method for rapidly identifying stomach helicobacter pylori infection

    CN117054389A

  • Gout diagnosis and recurrence risk prediction method and system based on Raman spectrum and multi-modal machine learning

    CN120148848A

  • Non-invasive disease diagnosis using light scattering probe

    US20090086202A1