Gout diagnosis and recurrence risk prediction method and system based on Raman spectrum and multi-modal machine learning

By combining Raman spectroscopy and multimodal machine learning technology, the spectral characteristics of serum samples and screening metabolic biomarkers are constructed to predict the risk of gout onset, solving the problem of inaccurate prediction of gout onset risk in the prior art, and achieving more accurate health status assessment and treatment optimization.

CN120148848APending Publication Date: 2025-06-13NANFANG HOSPITAL OF SOUTHERN MEDICAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510210155.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the risk of gout onset, and it is impossible to accurately distinguish individuals with asymptomatic hyperuricemia from high-risk future gout development by relying solely on serum uric acid levels.

Method used

Using Raman spectroscopy and multimodal machine learning methods, by analyzing the spectral characteristics of serum samples of the control group, hyperuricemia group and gout group, metabolic biomarkers that can be used for disease identification, and multimodal models are constructed to predict the risk of gout onset.

Benefits of technology

It improves the accuracy of predicting gout attack risk, helps doctors to more accurately judge patients' health status, optimize treatment strategies, and improve patients' prognosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148848A_ABST
    Figure CN120148848A_ABST
Patent Text Reader

Abstract

The invention provides a gout diagnosis and recurrence risk prediction method and system based on Raman spectrum and multi-modal machine learning, and the method comprises the steps: generating a Raman spectrum sample data set, and carrying out the preprocessing; performing data dimension reduction and feature extraction on the preprocessed Raman spectrum sample data set to obtain key feature vectors; performing feature screening and enhancement on the key feature vectors before modeling; modeling the data before and after dimension reduction by using a machine learning algorithm to obtain a plurality of machine learning models; and training the multiple machine learning models by using a training set, verifying the models by using a ten-fold cross validation method, selecting the machine learning model with the best discrimination effect as a prediction model, and performing risk prediction of medical history development and gout recurrence. The Raman spectrum and the advanced artificial intelligence technology are integrated, a new prospect is provided for identifying unique metabolic characteristics related to the diseases, the clinical management level of gout is improved, and intervention is earlier and more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of Raman spectroscopy processing, and particularly to a method and system for gout diagnosis and recurrence risk prediction based on Raman spectroscopy and multimodal machine learning. Background Art

[0002] Gout and hyperuricemia (HUA) are common metabolic diseases globally. Gout is characterized by the chronic deposition of monosodium urate (MSU) crystals in joints and soft tissues. Hyperuricemia refers to the situation where, under normal purine diet, fasting serum uric acid (SUA) exceeds 420 μmol / L (7 mg / dL) twice, and it is currently recognized as a precursor of gout. However, although elevated SUA levels are closely related to gout, not all hyperuricemia patients will inevitably develop gout. Therefore, relying solely on serum uric acid levels may not always be a reliable predictor of gout attacks, which poses a challenge to clinically accurately distinguishing asymptomatic hyperuricemia and individuals at high risk of future gout development.

[0003] Among those patients who have developed into gout, there are differences between sporadic gout (attack ≤ 1 time per year) and frequent gout (attack ≥ 2 times per year). Frequent gout not only seriously affects the quality of life of patients but is also associated with a higher risk of complications. Therefore, early prediction of its attack risk is of great significance for optimizing treatment strategies and improving patient prognosis. Existing studies have shown that factors such as serum uric acid levels, MSU crystal deposition amounts, and CA72-4 levels are related to the prediction of gout recurrence within one year.

[0004] Metabolomics has shown good prospects in the identification of disease biomarkers. It comprehensively studies metabolites in biological systems and has made significant progress in revealing pathological metabolic changes in recent years. Due to the non-invasive nature of Raman spectroscopy and its ability to provide complex biochemical maps of biological samples, it has become an effective tool for detecting molecular vibrations in metabolomics. Previous studies have confirmed the effectiveness of Raman spectroscopy in diagnosing various diseases by identifying disease-specific metabolic fingerprints. However, Raman spectroscopy technology has not been explored in differentiating hyperuricemia, gout, and people with normal uric acid levels, as well as predicting the attack risk of gout. Given the complex metabolism of gout and hyperuricemia, the integration of Raman spectroscopy and advanced artificial intelligence technology provides new prospects for identifying unique metabolic characteristics related to these diseases.

[0005] Therefore, how to improve the accuracy of predicting the attack risk of gout and help doctors more accurately judge the health status of patients is an important research content for those skilled in the art in this field. Summary of the Invention

[0006] In view of this, it is necessary to provide a gout diagnosis and recurrence risk prediction method and system based on Raman spectroscopy and multimodal machine learning. By using Raman spectroscopy fusion AI technology to analyze the spectral characteristics of serum samples in the control group, hyperuricemia group and gout group, metabolic biomarkers for disease identification are screened, and a multimodal model capable of predicting the risk of gout attack is constructed accordingly. This will help improve the clinical management level of gout and enable earlier and more accurate intervention.

[0007] It can at least overcome one of the above defects.

[0008] In a first aspect, an embodiment of the present application provides a gout diagnosis and recurrence risk prediction method based on Raman spectroscopy and multimodal machine learning. The method includes:

[0009] S1: Select Raman spectroscopy samples with a wavenumber range from 402 cm -1 to 2002 cm -1 Collect about 50 spectral samples from each sample to generate a Raman spectroscopy sample dataset, and preprocess the Raman spectroscopy sample dataset to eliminate the intensity differences between different spectra;

[0010] S2: Use the principal component PCA analysis method to perform data dimensionality reduction and feature extraction on the preprocessed Raman spectroscopy sample dataset to obtain key feature vectors;

[0011] S3: Perform feature screening and enhancement before modeling on the key feature vectors to provide a data source for constructing a machine learning model;

[0012] S4: Use machine learning algorithms to model the data before and after dimensionality reduction to obtain multiple machine learning models;

[0013] S5: Divide the patients into a training set and a test set according to a ratio of 8:2. Use the training set to train the multiple machine learning models. Use the PCA-SVM method to distinguish the average Raman spectra of the serum of the three types of patients through the test set and use the ten-fold cross-validation method to verify the models. List the accuracy and sensitivity index values of the models and draw an ROC curve graph, and draw the conclusion that the PCA-LDA model has the best discrimination effect;

[0014] S6: Select the machine learning model with the best discrimination effect as the prediction model to predict the risk of medical history development and gout recurrence.

[0015] Optionally, in another implementation manner of the first aspect of the present invention, the preprocessing of the Raman spectroscopy sample dataset to eliminate the intensity differences between different spectra includes:

[0016] S1.1: Perform baseline correction on the spectrum through a polynomial fitting method to eliminate background interference;

[0017] S1.2: Smooth the spectral data through Savitzky-Golay filter to reduce the influence of noise on the analysis results;

[0018] S1.3: Align the spectra of all samples and normalize them to the same scale to ensure the consistency and comparability of the data.

[0019] Optionally, in another implementation of the first aspect of the present invention, S2: Use the principal component PCA analysis method to perform data dimensionality reduction and feature extraction on the preprocessed Raman spectral sample dataset to obtain key feature vectors, including:

[0020] S2.1: Project the high-dimensional data into the low-dimensional space by using PCA transformation in the high-dimensional space to reduce the dimensionality of the Raman spectral sample dataset, and finally obtain a reconstructed dataset, including:

[0021] S2.1.1: Represent the Raman spectral sample features as matrix A = [x 1 , x 2 ,..., x N T , where x i represents the feature vector of Raman spectral sample i, and N represents the number of Raman spectral samples;

[0022] S2.1.2: Calculate the deviation matrix B of the Raman spectral samples, and the formula is:

[0023]

[0024] where; f represents the count of the number of samples;

[0025] S2.1.3: Perform eigenvalue decomposition on the deviation matrix through the orthogonal matrix V composed of eigenvectors B = VΛV T , where V = [v 1 , v 2 ,..., v M , Λ = diag(λ 1 , λ 2 ,....λ M ) is a diagonal matrix, λ f is the element on the diagonal, take the eigenvalue, and λ 1 ≥ λ 2 ≥.... ≥ λ M ≥ 0, M represents the dimension;

[0026] S2.1.4: Select the eigenvectors V corresponding to the H largest eigenvalues H = [v 1 , v 2 ,..., v​H , then the projection Y of the original dataset A in the principal component space is Y = AV H is the data after PCA transformation, where the dimension of Y is N×H, realizing the dimensionality reduction from M to H;

[0027] S2.2: Determine the degree of abnormality of each sample by calculating the reconstruction error. The reconstruction error E s (x) The calculation formula is:

[0028]

[0029] where H is the dimension of dimensionality reduction, represents the center point of the i-th cluster after clustering, and i and j represent the number of clusters and the dimension count respectively, C i represents the i-th cluster, and k represents the number of clusters.

[0030] Optionally, in another implementation manner of the first aspect of the present invention, S3: Perform feature screening and enhancement before modeling the key feature vectors, including:

[0031] S3.1: Use LASSO regression to screen the key features of the Raman spectrum after dimensionality reduction to reduce the model complexity;

[0032] S3.2: For the missing data in the clinical features, adopt a targeted imputation strategy: continuous variables are imputed through the KNN model to maintain the continuity of the data; categorical variables are imputed using the mode to ensure that the class distribution of the data is not affected;

[0033] S3.3: Standardize the continuous variables in all data, and perform one-hot encoding conversion on the categorical variables to meet the input requirements of the machine learning model;

[0034] S3.4: Before modeling, remove the features with strong collinearity through multilinear analysis to reduce model redundancy, and identify and remove the outlier samples that may have a negative impact on model training;

[0035] S3.4: Generate enhanced sample data through the adversarial network, where the objective function of the adversarial network is:

[0036]

[0037] where D(x) represents the probability that the real data x is judged to be true, D(G(z)) represents the probability that the forged data z is judged to be true, represents the conditional distribution, represents the noise distribution, G is the generator, D is the discriminator, and λ is the coefficient of distribution penalty;

[0038] The ascending gradient during the training process is: The descending gradient is: where m represents the number of samples,

[0039] ▽ θd and ▽ θg represent the ascending gradient penalty term and the descending gradient penalty term respectively. D(x (i) ) represents the probability that the true data x is judged to be true on the exponent i, and D(G(z (i) ) represents the probability that the forged data z is judged to be true on the exponent i respectively.

[0040] Optionally, in another implementation manner of the first aspect of the present invention, in step S4: using a machine learning algorithm to model the data before and after dimensionality reduction to obtain multiple machine learning models, including:

[0041] According to the comparison of multiple groups of Raman spectrum differences, the peak intensities with statistical significance between groups are used as characteristic spectra to input into the algorithm for learning and modeling. 80% of the data is used as the training set, and 20% of the data is used as the validation set, so as to develop multiple machine learning models;

[0042] where the machine learning algorithm at least includes Random Forest algorithm, Support Vector Machine (SVM) algorithm, Linear Discriminant Analysis (LDA) algorithm, and Extreme Gradient Boosting algorithm.

[0043] Optionally, in another implementation manner of the first aspect of the present invention, when using the Random Forest algorithm to develop a machine learning model, it includes:

[0044] Initializing the random forest, which is composed of several decision trees;

[0045] The samples and labels are input from the root node, and the samples are continuously recursively divided into left and right child nodes according to the sample feature attributes until the sample labels are of the same class or cannot be divided, and this node is the leaf node.

[0046] Optionally, in another implementation manner of the first aspect of the present invention, in step S6: selecting the machine learning model with the best discrimination effect as the prediction model to perform the risk prediction of the medical history development and gout recurrence, including:

[0047] Introducing the accuracy rate as the evaluation criterion of the prediction model, and the accuracy rate is:

[0048] where T p and T n represent the number of correct classifications and incorrect classifications of the model respectively.

[0049] Second aspect, an embodiment of the present application provides a gout diagnosis and recurrence risk prediction system based on Raman spectroscopy and multimodal machine learning, which is applied to a gout diagnosis and recurrence risk prediction method based on Raman spectroscopy and multimodal machine learning as described in the first aspect. The system includes:

[0050] A data acquisition module, configured to select Raman spectroscopy samples with a wavenumber range from 402 cm -1 to 2002 cm -1 , collect about 50 spectral samples from each sample to generate a Raman spectroscopy sample dataset, and preprocess the Raman spectroscopy sample dataset to eliminate the intensity differences between different spectra;

[0051] A feature extraction module, configured to perform data dimensionality reduction and feature extraction on the preprocessed Raman spectroscopy sample dataset by using the principal component PCA analysis method to obtain key feature vectors;

[0052] A feature processing module, configured to perform feature screening and enhancement on the key feature vectors before modeling to provide a data source for constructing a machine learning model;

[0053] A model construction module, configured to use machine learning algorithms to model the data before and after dimensionality reduction to obtain multiple machine learning models;

[0054] A model selection module, configured to divide patients into a training set and a test set according to a ratio of 8:2, use the training set to train the multiple machine learning models, use the PCA-SVM method to distinguish the average Raman spectra of three types of patients' sera through the test set, and use the ten-fold cross-validation method to verify the models, list the accuracy rate and sensitivity index values of the models, and draw an ROC curve graph, and draw the conclusion that the PCA-LDA model has the best discrimination effect;

[0055] A risk prediction module, configured to select the machine learning model with the best discrimination effect as a prediction model to perform risk prediction on the development of medical history and gout recurrence.

[0056] Third aspect, an embodiment of the present application provides an electronic device, including:

[0057] A processor;

[0058] A memory for storing processor-executable instructions;

[0059] Wherein, when the processor is configured to execute the instructions, it implements a gout diagnosis and recurrence risk prediction method based on Raman spectroscopy and multimodal machine learning as described in the first aspect.

[0060] Fourthly, an embodiment of the present application provides a computer-readable storage medium storing a program that instructs a device to execute a gout diagnosis and recurrence risk prediction method based on Raman spectroscopy and multimodal machine learning as described in the first aspect.

[0061] In the technical solution provided by the present invention, by selecting Raman spectroscopy samples with a wavenumber range from 402 cm -1 to 2002 cm -1 , collecting about 50 spectral samples from each sample to generate a Raman spectroscopy sample data set, and preprocessing the Raman spectroscopy sample data set to eliminate the intensity differences between different spectra; using the principal component PCA analysis method to perform data dimensionality reduction and feature extraction on the preprocessed Raman spectroscopy sample data set to obtain key feature vectors; performing feature screening and enhancement before modeling on the key feature vectors to provide a data source for constructing a machine learning model; using machine learning algorithms to model the data before and after dimensionality reduction to obtain multiple machine learning models; dividing the patients into a training set and a test set according to a ratio of 8:2, using the training set to train the multiple machine learning models, using the test set to distinguish the average Raman spectra of three types of patients from the multiple machine learning models by the PCA-SVM method and using the ten-fold cross-validation method to verify the models, listing the accuracy and sensitivity index values of the models and drawing an ROC curve graph, and obtaining the conclusion that the PCA-LDA model has the best discrimination effect; selecting the machine learning model with the best discrimination effect as the prediction model to predict the risk of medical history development and gout recurrence.

[0062] A gout diagnosis and recurrence risk prediction method and system based on Raman spectroscopy and multimodal machine learning provided by an embodiment of the present application integrates Raman spectroscopy with advanced artificial intelligence technology, providing a new prospect for identifying unique metabolic characteristics related to these diseases, facilitating the improvement of the clinical management level of gout, and enabling earlier and more accurate interventions. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is a schematic flowchart of a gout diagnosis and recurrence risk prediction method based on Raman spectroscopy and multimodal machine learning provided by an embodiment of the present application.

[0064] Figure 2 A-C are diagrams comparing the differences of different classifiers provided by an embodiment of the present application.

[0065] Figure 3 A-C are diagrams comparing the differences of different classifiers provided by an embodiment of the present application.

[0066] Figure 4 It is an ROC curve graph of a random forest classifier provided by an embodiment of the present application.

[0067] Figure 5 A gout diagnosis and recurrence risk prediction method based on Raman spectroscopy and multimodal machine learning provided by an embodiment of the present application.

[0068] Figure 6 Schematic diagram of an electronic terminal device provided by an embodiment of the present application. Detailed implementation manners

[0069] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments.

[0070] It should be noted that "at least one" in the embodiments of the present application means one or more, and multiple means two or more. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used in the specification of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application.

[0071] It should be noted that in the embodiments of the present application, terms such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order. Features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.

[0072] Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.

[0073] The present application provides a gout diagnosis and recurrence risk prediction method and system based on Raman spectroscopy and multimodal machine learning. The integration of Raman spectroscopy and advanced artificial intelligence technology provides new prospects for identifying unique metabolic characteristics related to these diseases. Using Raman spectroscopy fusion AI technology to analyze the spectral characteristics of serum samples in the control group, hyperuricemia group and gout group, screening metabolic biomarkers that can be used for disease identification, and constructing a multimodal model that can distinguish these three types of populations and predict the risk of gout attacks based on this. This will be beneficial to improving the clinical management level of gout, making the intervention earlier and more accurate.

[0074] Figure 1 Schematic diagram of the gout diagnosis and recurrence risk prediction method based on Raman spectroscopy and multimodal machine learning provided by an embodiment of the present application.

[0075] As Figure 1 shown, a gout diagnosis and recurrence risk prediction method based on Raman spectroscopy and multimodal machine learning, the method includes:

[0076] S1: Select Raman spectroscopic samples with a wavenumber range from 402 cm -1 to 2002 cm -1 Collect about 50 spectral samples from each sample to generate a Raman spectroscopic sample dataset, and preprocess the Raman spectroscopic sample dataset to eliminate the intensity differences between different spectra.

[0077] It can be understood that in this embodiment, Raman spectra in the range from 402 to 2002 cm -1 were recorded, about 50 spectra were collected for each sample, and a total of 17,125 Raman spectroscopic data were obtained. To ensure the consistency of data quality, the following preprocessing steps were performed on all spectral data in this study: baseline correction of the spectra using the polynomial fitting method to eliminate background interference; application of the Savitzky-Golay filter to smooth the spectral data and reduce the influence of noise on the analysis results; alignment of the spectra of all samples and normalization to the same scale to ensure the consistency and comparability of the data.

[0078] Specifically, in the embodiment of the present application, the preprocessing of the Raman spectroscopic sample dataset to eliminate the intensity differences between different spectra includes:

[0079] S1.1: Perform baseline correction on the spectra through the polynomial fitting method to eliminate background interference;

[0080] S1.2: Smooth the spectral data through the Savitzky-Golay filter to reduce the influence of noise on the analysis results;

[0081] S1.3: Align the spectra of all samples and normalize them to the same scale to ensure the consistency and comparability of the data.

[0082] S2: Use the principal component PCA analysis method to perform data dimensionality reduction and feature extraction on the preprocessed Raman spectroscopic sample dataset to obtain key feature vectors.

[0083] It is understandable that, in order to reduce the computational complexity of high-dimensional data and extract key features, we have adopted a variety of dimensionality reduction techniques, including principal component analysis (PCA), kernel principal component analysis (kPCA), multidimensional scaling (MDS), locally linear embedding (LLE), modified locally linear embedding (MLLE), and spectral embedding, etc.

[0084] Specifically, baseline correction, noise reduction, and normalization were adopted in the spectral data preprocessing to ensure the consistency of data quality. At the same time, the use of dimensionality reduction techniques (such as PCA, kPCA, MDS, LLE, etc.) provides a reliable basis for high-dimensional data analysis.

[0085] Specifically, in the embodiment of the present application, in step S2: the preprocessed Raman spectroscopy sample data set is subjected to data dimensionality reduction and feature extraction by using the principal component PCA analysis method to obtain key feature vectors, including:

[0086] S2.1: In the high-dimensional space, the high-dimensional data is projected into the low-dimensional space by using the PCA transformation to perform dimensionality reduction on the Raman spectroscopy sample data set, and finally a reconstructed data set is obtained, including:

[0087] S2.1.1: Represent the Raman spectroscopy sample features as matrix A = [x 1 , x 2 ,..., x N T , where x i represents the feature vector of Raman spectroscopy sample i, and N represents the number of Raman spectroscopy samples;

[0088] S2.1.2: Calculate the deviation matrix B of the Raman spectroscopy samples, and the formula is:

[0089]

[0090] where; f represents the count of the number of samples;

[0091] S2.1.3: Perform eigenvalue decomposition on the deviation matrix through the orthogonal matrix V composed of eigenvectors B = VΛV T , where V = [v 1 , v 2 ,..., v M , Λ = diag(λ 1 , λ 2 ,....λ M ) is a diagonal matrix, λ f is the element on the diagonal, take the eigenvalue, and λ 1 ≥ λ 2 ≥.... ≥ λ M ≥ 0, and M represents the dimension; ​

[0092] S2.1.4: Select the eigenvectors V corresponding to the H largest eigenvalues H = [v 1 , v 2 ,..., v H . Then the projection Y of the original dataset A in the principal component space is Y = AV H which is the data after PCA transformation. The dimension of Y is N×H, achieving dimensionality reduction from M to H;

[0093] S2.2: Judge the degree of abnormality of each sample by calculating the reconstruction error. The reconstruction error E s (x) is calculated as follows:

[0094]

[0095] where H is the dimension of dimensionality reduction, represents the center point of the i-th cluster after clustering. i and j represent the number of clusters and the dimension count respectively, C i represents the i-th cluster, and k represents the number of clusters.

[0096] Through the extraction of data features by PCA, the main features in the data can be extracted, important biomarkers or clinical information can be captured, the spectral difference features between different types of samples can be distinguished. Subsequently, machine learning algorithms can be used to establish different classification models for classifying and quantitatively analyzing spectral samples, so as to understand the change trends of different types of spectral samples and evaluate and optimize the performance of the models.

[0097] The average Raman spectral morphologies and spectral peaks of the three types of sera are generally similar. By comparing and analyzing the Raman spectral signals of the three groups of sera, the main spectral peaks are located at 469 - 501, 629 - 639, 721 - 731, 1201 - 1218, 1311 - 1342, 1435 - 1454, 1568 - 1594, 1986 - 2002 cm-1 displacements. The HUA group and the GOUT group show lower intensities, and the possible explanations for the corresponding substances and the pathophysiology of the diseases are discussed.

[0098] The peak intensities of the control group and the HUA group are both 636.2 cm-1, while the peak intensity of the GOUT group is 637.3 cm-1. The differences in spectral peaks between the control group and the HUA group and the gout group may reflect the different distributions of metabolic biomarkers, especially the significant differences at 478.8 cm^-1 and 518.7 cm^-1.

[0099] S3: Perform feature screening and enhancement on the key feature vectors before modeling to provide data sources for constructing machine learning models.

[0100] Specifically, S3: Feature screening and enhancement before modeling the key feature vectors, including:

[0101] S3.1: Use LASSO regression to screen the key features of the Raman spectrum after dimensionality reduction, reducing the model complexity;

[0102] S3.2: For the missing data in the clinical features, adopt targeted imputation strategies: continuous variables are imputed through the KNN model to maintain data continuity; categorical variables are imputed using the mode to ensure that the class distribution of the data is not affected;

[0103] S3.3: Standardize the continuous variables in all data, and perform one-hot encoding conversion on the categorical variables to meet the input requirements of the machine learning model;

[0104] S3.4: Before modeling, remove features with strong collinearity through multilinear analysis to reduce model redundancy, and identify and remove outlier samples that may have a negative impact on model training.

[0105] It can be understood that in the feature preprocessing stage, LASSO regression is used to screen the key features of the Raman spectrum after dimensionality reduction, thereby reducing the model complexity. For the missing data in the clinical features, targeted imputation strategies are adopted: continuous variables are imputed through the KNN model to maintain data continuity; categorical variables are imputed using the mode to ensure that the class distribution of the data is not affected. In addition, the continuous variables in all data are standardized, and the categorical variables are subjected to one-hot encoding conversion to meet the input requirements of the machine learning model. Before modeling, features with strong collinearity are also removed through multilinear analysis to reduce model redundancy, and outlier samples that may have a negative impact on model training are identified and removed.

[0106] S3.4: Generate enhanced sample data through the adversarial network, where the objective function of the adversarial network is:

[0107]

[0108] where D(x) represents the probability that the real data x is judged to be true, D(G(z)) represents the probability that the forged data z is judged to be true, represents the conditional distribution, represents the noise distribution, G is the generator, D is the discriminator, and λ is the coefficient of the distribution penalty;

[0109] The ascending gradient during the training process is: The descending gradient is: where m represents the number of samples,

[0110] ▽ θd 、▽θg represent the ascending gradient penalty term and the descending gradient penalty term respectively. D(x (i) ) represents the probability that the real data x is judged to be true at index i. D(G(z (i) ) represents the probability that the forged data z is judged to be true at index i respectively.

[0111] S4: Use machine learning algorithms to model the data before and after dimensionality reduction to obtain multiple machine learning models.

[0112] It can be understood that in order to establish a classification model that can distinguish different types of spectral samples, a supervised machine learning algorithm, random forest, can be selected, or machine learning methods such as linear discriminant models and support vector machine models can be used to analyze Raman spectra.

[0113] Specifically, in the step S4: Use machine learning algorithms to model the data before and after dimensionality reduction to obtain multiple machine learning models, including:

[0114] According to the comparison of the differences in multiple groups of Raman spectra, the peak intensities with statistical significance between groups are used as characteristic spectra to input into the algorithm for learning and modeling. 80% of the data is used as the training set, and 20% of the data is used as the validation set, so as to develop multiple machine learning models;

[0115] Among them, the machine learning algorithms at least include Random Forest algorithm, Support Vector Machine (SVM) algorithm, Linear Discriminant Analysis (LDA) algorithm, and Extreme Gradient Boosting algorithm.

[0116] Specifically, when using the Random Forest algorithm to develop a machine learning model, it includes:

[0117] Initialize the random forest, which consists of several decision trees;

[0118] The samples and labels are input from the root node, and the samples are continuously recursively divided into the left and right child nodes according to the sample feature attributes until the sample labels are of the same class or cannot be divided, and this node is the leaf node.

[0119] S5: Divide the patients into a training set and a test set according to a ratio of 8:2. Use the training set to train the multiple machine learning models, use the PCA-SVM method to distinguish the average Raman spectra of the sera of three types of patients through the test set, and use the ten-fold cross-validation method to verify the models. List the accuracy rate and sensitivity index values of the models and draw an ROC curve graph, and draw the conclusion that the PCA-LDA model has the best discrimination effect.

[0120] It is understandable that the patients are divided into a training set and a test set in a ratio of 8:2. In the training set, all spectral data are used for model training (control group: HUA: GOUT = 2467:5185:6158). In the test set, the average spectrum of each patient is calculated (control group: HUA: GOUT = 12:21:36).

[0121] Specifically, 12 machine learning models such as Random Forest (RF), Logistic Regression (LR), and Extreme Gradient Boosting (XGBoost) are developed to distinguish control group, HUA, and GOUT samples. Ten-fold cross-validation is adopted during the training process, and the models are evaluated using accuracy, area under the curve (AUC), recall rate, kappa, precision, and F1 score.

[0122] S6: Select the machine learning model with the best discrimination effect as the prediction model to predict the risk of medical history development and gout recurrence.

[0123] It is understandable that the prediction model can be screened through the performance of the characteristic curve, or the accuracy can be used as an evaluation index to quantitatively make the judgment basis for screening.

[0124] Specifically, the S6: Select the machine learning model with the best discrimination effect as the prediction model to predict the risk of medical history development and gout recurrence, including:

[0125] Introduce accuracy as the evaluation criterion for the prediction model. The accuracy is:[[]]END]]

[0126] where T p 、T n represent the number of correct classifications and incorrect classifications that the model can make respectively.

[0127] Specifically, all experiments are carried out using Python (v3.11), ensuring the readability and reproducibility of the code. Here, machine learning libraries such as Scikit-learn are used to build and verify the models.

[0128] Example 1

[0129] First, select two groups of general clinical characteristics of RA patients: introduce the situations of the two groups of patients in terms of gender, age, disease course, etc., and compare the differences between the two groups of patients in indicators such as the number of swollen joints and the number of tender joints. Secondly, the total Raman spectra of the two groups of RA patients: elaborate on the spectral excitation wavelength and spectral range selected for the experiment, introduce the measurement of 116 serum samples and the drawn total original spectra. Obtain the differential spectra, explain that the mean values of the sample measurement results are taken, and the differential spectra are obtained through noise reduction and other processes. Point out that there are statistical differences in the spectral peaks at specific displacements between the low group and the high group, and list the corresponding substances. Then perform PCA feature extraction: use the PCA dimensionality reduction algorithm to reduce the original 315-dimensional features to 10 dimensions, draw the two-dimensional scatter plot after dimensionality reduction, and find that some samples of the low group and the high group are mixed, and a machine learning model needs to be used to improve the classification accuracy. Establish an RA disease activity assessment model based on Raman spectroscopy: use the SVM and LDA algorithms to model the data before and after dimensionality reduction to obtain four models. Divide the data set into a training set and a test set according to 8:2, verify the model using the ten-fold cross-validation method, list the index values such as the accuracy rate and sensitivity of the model, and draw the ROC curve graph, and draw the conclusion that the PCA-LDA model has the best discrimination effect.

[0130] Example 2

[0131] The general clinical characteristics of the sera of patients with benign and malignant tumors and healthy patients undergoing physical examinations in the Department of Breast Surgery of Sichuan Cancer Hospital were collected, and the number of cases, gender, age, clinical stage, histological grade, pathological type, etc. of the malignant group, benign group, and normal group were introduced. The average Raman spectra of sera of different types of patients: The average Raman spectra of sera of the malignant tumor group, benign group, and healthy group were compared, and it was found that there were differences in the spectra of the three groups. Compared with the healthy group, the intensities of multiple peaks increased or decreased in the malignant tumor group; compared with the healthy group, the intensities of some peaks also changed in the benign group; the spectral signals of the malignant tumor group and the benign group were similar, but there were still obvious differences in the intensities of some peaks. These differences may be related to factors such as tumor tissue proliferation, cell necrosis, metabolism, nutrient consumption, and gene mutations. Raman spectral PCA-SVM analysis of sera of different types of patients: After processing the original Raman spectra, the PCA-SVM method was used to distinguish the average Raman spectra of sera of the three types of patients and cross-validate. The results showed that the overall accuracy rate of the SVM model reached 98%, but there was a relatively high error rate between the cancer group and the benign group. The sensitivity, specificity, and accuracy of the model for the malignant tumor group, benign group, and healthy group were all relatively high, and the AUC value of the ROC curve was close to 1, indicating that the model had a relatively high discrimination ability for early breast cancer screening.

[0132] Building a classification model: Based on the comparison of Raman spectra differences among three groups, the peak intensities with statistical significance between groups were used as characteristic spectra to input into the algorithm for learning and modeling. 80% of the data was used as the training set, and 20% of the data was used as the validation set. In this way, 12 machine learning models, such as Random Forest (RF), Logistic Regression (LR), and Extreme Gradient Boosting (XGBoost), were developed to distinguish contro, hua, and gout samples. 10-fold cross-validation was used to further optimize the algorithm, and the model was evaluated by indicators such as accuracy, area under the curve (AUC), recall rate, kappa, precision, and F1 score (Table 1) to obtain the best classification algorithm. From the experimental results of 12 machine learning models, the auc of 5 out of 12 machine learning classifiers was higher than 0.85 (Table 1). Among them, the micro-AUC of XGBoost was the highest, at 0.89 (95% CI 0.82 - 0.95), and the macro-AUC was the highest, at 0.88 (95% CI 0.80 - 0.94)( Figure 2 A - B). In addition, the precision, recall rate, accuracy, and F1 score of XGBoost were 0.77, and the kappa was 0.62, which was better than the other four classifiers in these indicators( Figure 2 C), although their auc was comparable, but overall their performance was lower compared to the XGBoost classifier. In a specific subgroup, the AUC of XGBoost, RF, and LGBM was 0.89 when classifying the control group. However, the 95% CI of XGBoost was 0.78 - 0.98. The classification AUC of RF for the GOUT group was 0.89 (95% CI 0.80 - 0.96), while the classification AUC of the LGBM classifier for the HUA group was 0.88 (95% CI 0.77 - 0.96) Figure 3 A - C). In summary, the overall performance of the XGBoost classification model was better than other classifiers, so this classifier was used for further experimental research.

[0133] Table 1 Experimental results of 12 machine learning classifiers

[0134]

[0135] By combining Raman spectral features with clinical features, among 12 machine learning classifiers, the Random Forest model had the best comprehensive performance in predicting gout recurrence within 1 year. As Figure 4As shown in Table 1, the AUC of the Random Forest model is 90.42%, the accuracy is 77.42%, the recall rate is 68.75%, the precision is 84.62%, the F value is 75.86%, the kappa value is 55.07%, and the MMC is 56.12%. As shown in Table 1, this study details the performance indicators of various machine learning classifiers for predicting gout recurrence within one year based on clinical characteristics, including the area under the curve (AUC), accuracy, recall rate, precision, F1 value, kappa coefficient, and Matthews correlation coefficient (MCC). Compared with the performance indicators of other algorithm models, the Random Forest model is the optimal model for predicting gout recurrence within one year constructed by combining Raman spectral features and clinical characteristics.

[0136] Using 12 machine learning models and applying ten-fold cross-validation ensured the robustness and generalization ability of the models. The performance of XGBoost was excellent and it became the core classifier of this study. The comparison and combination of clinical + Raman spectral features improved the efficiency of the model.

[0137] Figure 5 This is a schematic diagram of the functional modules of a gout diagnosis and recurrence risk prediction system based on Raman spectroscopy and multimodal machine learning provided by an embodiment of this application. As Figure 5 shown, a gout diagnosis and recurrence risk prediction system based on Raman spectroscopy and multimodal machine learning, the system includes a data acquisition module 11, a feature extraction module 12, a feature processing module 13, a model construction module 14, a model selection module 15, and a risk prediction module 16 that are connected in sequence, where:

[0138] The data acquisition module 11 is used to select Raman spectral samples with a wavenumber range from 402 cm -1 to 2002 cm -1 , collect about 50 spectral samples from each sample to generate a Raman spectral sample dataset, and preprocess the Raman spectral sample dataset to eliminate the intensity differences between different spectra;

[0139] The feature extraction module 12 is used to perform data dimensionality reduction and feature extraction on the preprocessed Raman spectral sample dataset by using the principal component PCA analysis method to obtain key feature vectors;

[0140] The feature processing module 13 is used to perform feature screening and enhancement on the key feature vectors before modeling to provide a data source for constructing a machine learning model;

[0141] The model construction module 14 is used to use machine learning algorithms to model the data before and after dimensionality reduction to obtain a variety of machine learning models;

[0142] A model selection module 15, which is used to divide patients into a training set and a test set according to a ratio of 8:2, train the multiple machine learning models using the training set, use the test set to distinguish the average Raman spectra of three types of patients' sera for the multiple machine learning models by the PCA-SVM method, verify the models using the ten-fold cross-validation method, list the accuracy rate and sensitivity index values of the models, draw an ROC curve, and draw the conclusion that the PCA-LDA model has the best discrimination effect;

[0143] A risk prediction module 16, which is used to select the machine learning model with the best discrimination effect as a prediction model to perform risk prediction on the development of medical history and gout recurrence.

[0144] It can be understood that a gout diagnosis and recurrence risk prediction method and system based on Raman spectroscopy and multimodal machine learning provided by an embodiment of the present application have the clinical value of disease differentiation: Through Raman spectroscopy combined with AI technology, this study successfully explored the spectral differences in serum samples of the control group, hyperuricemia (HUA) group, and gout (GOUT) group, providing the possibility for disease identification. It has potential applications for risk prediction: Gout recurrence risk prediction is of great significance for optimizing patient management and treatment. An effective multimodal prediction model was developed to predict the risk of gout recurrence within one year, providing a decision-making basis for clinical practice.

[0145] See Figure 6 , Figure 6 is an electronic terminal device provided by an embodiment of the present application. As Figure 6 shown, the electronic device at least includes the following parts: a processor 101, a memory 100, a communication interface 103, and a bus 102.

[0146] In an embodiment of the present application, the memory 100 is used to store executable instructions of the processor 101, and the processor 101 is configured to implement the method as Figure 1 shown when executing the instructions.

[0147] In an embodiment of the present application, a computer-readable storage medium includes instructions that instruct a device to execute the method in the first aspect. For example, the instructions instruct the device to execute the method shown in the process steps in Figure 1 .

[0148] The program operating in the electronic device according to an embodiment of the present application can be a program that controls a Central Processing Unit (CPU) and the like to implement the functions of the above-described embodiments related to a solution of the present invention (a program that causes a computer to function). Then, the information processed by these devices is temporarily stored in a Random Access Memory (RAM) during its processing, and thereafter, it is stored in various ROMs such as a Read Only Memory (Flash ROM), a Hard Disk Drive (HDD), etc., and is read out, corrected, and written by the CPU as needed.

[0149] It should be noted that a part of the electronic device of the above-described embodiment can also be implemented by a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and is implemented by reading the program recorded on the recording medium into the computer and executing it.

[0150] It should be noted that the "computer" mentioned here refers to a computer built into the electronic device, which is a computer including hardware such as an OS and peripheral devices. In addition, the "computer-readable recording medium" refers to a removable medium such as a floppy disk, a magneto-optical disk, a ROM, a CD-ROM, etc., and a storage device such as a hard disk built into the computer.

[0151] Moreover, the "computer-readable recording medium" can include: a medium that dynamically stores a program for a short period of time, such as a communication line in the case of transmitting a program via a network such as the Internet or a communication line such as a telephone line; a medium that stores a program for a fixed period of time, such as a volatile memory inside a computer of a server or a client in this case. In addition, the above program can be a program for implementing a part of the above functions, and can also be a program that can implement the above functions by combining with a program already recorded in the computer.

[0152] In addition, the electronic device in the above-described embodiment can also be implemented as an aggregate (device group) composed of multiple devices. Each device constituting the device group can have all or part of the functions or function blocks of the electronic device of the above-described embodiment. As the device group, it is sufficient to have all the functions or function blocks of the electronic device.

[0153] Those of ordinary skill in the art of this technology should recognize that the above embodiments are only used to illustrate the present application, rather than to limit the present application. As long as it is within the scope of the spirit of the present application, appropriate changes and variations made to the above embodiments fall within the scope of protection required by the present application.

Claims

1. A method for gout diagnosis and recurrence risk prediction based on Raman spectroscopy and multimodal machine learning, characterized in that: The method comprises: S1: Select the wave number range to be 402cm -1 To 2002cm -1 Raman spectrum samples, collecting about 50 spectrum samples from each sample to generate a Raman spectrum sample data set, and preprocessing the Raman spectrum sample data set to eliminate intensity differences between different spectra; S2: Use the principal component analysis method to perform data dimension reduction and feature extraction on the preprocessed Raman spectroscopy sample data set to obtain key feature vectors; S3: Screening and enhancing the key feature vectors before modeling to provide a data source for building a machine learning model; S4: Use machine learning algorithms to model the data before and after dimensionality reduction to obtain multiple machine learning models; S5: The patients were divided into a training set and a test set in a ratio of 8:2, and the training set was used to train the multiple machine learning models. The PCA-S VM method was used to distinguish the average Raman spectra of the serum of three types of patients through the test set, and the model was verified by a ten-fold cross validation method. The accuracy and sensitivity index values ​​of the model were listed and the ROC curve was drawn, and it was concluded that the PCA-LDA model had the best discrimination effect; S6: The machine learning model with the best discriminant effect is selected as the prediction model to predict the risk of medical history development and gout recurrence.

2. The method for gout diagnosis and recurrence risk prediction based on Raman spectroscopy and multimodal machine learning according to claim 1, characterized in that: The Raman spectrum sample data set is preprocessed to eliminate the intensity difference between different spectra, including: S1.1: Baseline correction of the spectrum was performed by polynomial fitting method to eliminate background interference; S1.2: Smoothing the spectral data by using Savitzky-Golay filter to reduce the influence of noise on the analysis results; S1.3: Align the spectra of all samples and normalize them to the same scale to ensure data consistency and comparability.

3. A method for diagnosing gout and predicting recurrence risk based on Raman spectroscopy and multimodal machine learning according to claim 1, characterized in that: S2: using the principal component analysis method (PCA) to perform data dimension reduction and feature extraction on the pre-processed Raman spectrum sample data set to obtain key feature vectors, including: S2.1: In the high-dimensional space, the high-dimensional data is projected into a low-dimensional space by using PCA transformation, and the Raman spectroscopy sample data set is reduced in dimensionality, and finally a reconstructed data set is obtained, including: S2.1.1: Represent the Raman spectrum sample features as a matrix A = [x1, x2, ..., x N ] T , where x i represents the feature vector of Raman spectrum sample i, and N represents the number of Raman spectrum samples; S2.1.2: Calculate the dispersion matrix B of the Raman spectrum sample using the formula: Where; f represents the count of the number of samples; S2.1.3: Decompose the deviation matrix by the orthogonal matrix V composed of eigenvectors B = VΛV T , where V = [v1,v2,...,v M ],Λ=diag(λ1,λ2,....λ M ) is a diagonal matrix, λ f For the elements on the diagonal, take the eigenvalue, and λ1≥λ2≥....≥λ M ≥0, M represents dimension; S2.1.4: Select the eigenvectors V corresponding to the H largest eigenvalues H =[v1,v2,...,v H ], then the projection Y of the original data set A in the principal component space is AV H It is the data after PCA transformation, where the dimension of Y is N×H, achieving dimensionality reduction from M dimension to H; S2.2: Determine the abnormality of each sample by calculating the reconstruction error. Reconstruction error E s (x) is calculated as: Among them, H is the dimension of dimensionality reduction, represents the center point of the i-th cluster after clustering, i and j represent the number of clusters and dimension counts respectively, C i represents the i-th family, and k represents the number of clusters.

4. The method for gout diagnosis and recurrence risk prediction based on Raman spectroscopy and multimodal machine learning according to claim 3, characterized in that: S3: feature screening and enhancement of the key feature vector before modeling, including: S3.1: Use LASSO regression to screen the key features of Raman spectra after dimensionality reduction to reduce the complexity of the model; S3.2: For missing data in clinical characteristics, targeted interpolation strategies were adopted: continuous variables were interpolated using the KNN model to maintain data continuity; categorical variables were interpolated using the mode to ensure that the category distribution of the data was not affected; S3.3: Standardize all continuous variables in the data and perform one-hot encoding on categorical variables to meet the input requirements of the machine learning model; S3.4: Before modeling, use multilinear analysis to remove features with strong collinearity to reduce model redundancy and identify and remove outlier samples that may have a negative impact on model training; S3.4: Generate enhanced sample data through the adversarial network, where the adversarial network objective function is: Among them, D(x) represents the probability that the real data x is judged as true, D(G(z)) represents the probability that the forged data z is judged as true, and E x~pdata(x) represents the conditional distribution, E x~pz(z) represents the noise distribution, G is the generator, D is the discriminator, and λ is the coefficient of the distribution penalty; The ascending gradient during training is: The descent gradient is: Among them, m represents the number of samples, They represent the upward gradient penalty term and the downward gradient penalty term respectively, D(x (i) ) represents the probability that the real data x is judged as true at index i, D(G(z (i) ) represent the probability that the forged data z is judged as true at index i.

5. The method for gout diagnosis and recurrence risk prediction based on Raman spectroscopy and multimodal machine learning according to claim 4, characterized in that: S4: using a machine learning algorithm to model the data before and after dimensionality reduction to obtain a variety of machine learning models, including: Based on the comparison of Raman spectra differences among multiple groups, the peak intensities with statistical significance between groups were used as characteristic spectra to input the algorithm for learning and modeling. 80% of the data was used as the training set and 20% of the data was used as the validation set to develop multiple machine learning models. Among them, the machine learning algorithm includes at least a random forest algorithm, a support vector machine SVM algorithm, a linear discriminant analysis LDA algorithm, and an extreme gradient enhancement algorithm.

6. The method for gout diagnosis and recurrence risk prediction based on Raman spectroscopy and multimodal machine learning according to claim 5, characterized in that: When using the Random Forest algorithm, develop a machine learning model that includes: Initialize the random forest, which consists of several decision trees; Samples and labels are input from the root node, and samples are continuously recursively assigned to left and right child nodes based on sample feature attributes until the sample labels are of the same category or cannot be divided. This node is the leaf node.

7. The method for diagnosing gout and predicting recurrence risk based on Raman spectroscopy and multimodal machine learning according to claim 6, characterized in that: S6: selecting the machine learning model with the best discrimination effect as the prediction model, including: The accuracy is introduced as the evaluation criterion of the prediction model. The accuracy is: Among them, T p , T n They represent the number of correct and incorrect classifications that the model can achieve respectively.

8. A gout diagnosis and recurrence risk prediction system based on Raman spectroscopy and multimodal machine learning, applied to a gout diagnosis and recurrence risk prediction method based on Raman spectroscopy and multimodal machine learning as claimed in any one of claims 1 to 7, characterized in that: The system comprises: Data acquisition module, used to select the wave number range in 402cm -1 To 2002cm -1 Raman spectrum samples, collecting about 50 spectrum samples from each sample to generate a Raman spectrum sample data set, and preprocessing the Raman spectrum sample data set to eliminate intensity differences between different spectra; The feature extraction module is used to perform data dimension reduction and feature extraction on the preprocessed Raman spectrum sample data set using the principal component PCA analysis method to obtain key feature vectors; A feature processing module is used to screen and enhance the features of the key feature vectors before modeling, so as to provide a data source for building a machine learning model; The model building module is used to use machine learning algorithms to model the data before and after dimensionality reduction to obtain a variety of machine learning models; A model selection module is used to divide the patients into a training set and a test set in a ratio of 8:2, use the training set to train the multiple machine learning models, use the test set to use the PCA-SVM method to distinguish the average Raman spectra of the serum of three types of patients for the multiple machine learning models and use the ten-fold cross-validation method to verify the model, list the accuracy and sensitivity index values ​​of the model and draw a ROC curve diagram, and draw a conclusion that the PCA-LDA model has the best discrimination effect; The risk prediction module is used to select the machine learning model with the best discrimination effect as the prediction model to predict the risk of medical history development and gout recurrence.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement a method for diagnosing gout and predicting recurrence risk based on Raman spectroscopy and multimodal machine learning as described in any one of claims 1 to 7 when executing the instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, and the program instructs the device to execute a method for diagnosing gout and predicting recurrence risk based on Raman spectroscopy and multimodal machine learning as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Non-invasive early warning method and system for gastric cancer and helicobacter pylori based on gastric juice

    CN120853974A

  • A gastric juice-based non-invasive early warning method and system for gastric cancer and helicobacter pylori

    CN120853974B