Construction method of ankylosing spondylitis ossification progress prediction model based on metabonomics and artificial intelligence

The predictive model established by metabolomics analysis and artificial intelligence technology solves the problem that the existing technology is difficult to identify the rapid progress of ossification in patients with AS early, and realizes early identification and individualized treatment of the progress of AS ossification.

CN119943417APending Publication Date: 2025-05-06THE SECOND AFFILIATED HOSPITAL OF NAVAL MEDICAL UNIVERSITY PLA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411689874.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing methods for evaluating the progression of ankylosing spondylitis (AS) ossification is difficult to identify patients with rapid progression of ossification early, resulting in a lag in clinical intervention.

Method used

Plasma samples were analyzed by metabolomics, differential compounds related to the progress of AS ossification were screened out, and predictive models were established using artificial intelligence technology to distinguish the rate of ossification progress in AS patients.

Benefits of technology

It has achieved early identification of the speed of ossification progress in AS patients, helping clinicians to conduct targeted interventions as soon as possible, and improving the individualization accuracy of treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005150654060000031
    Figure BDA0005150654060000031
  • Figure BDA0005150654060000041
    Figure BDA0005150654060000041
  • Figure BDA0005150654060000071
    Figure BDA0005150654060000071
Patent Text Reader

Abstract

The invention provides a method for constructing an ankylosing spondylitis ossification progress prediction model based on metabonomics and artificial intelligence. The method comprises the following steps: carrying out metabonomics analysis and peak extraction to obtain a characteristic ion peak table, carrying out compound identification on characteristic ion peaks, and carrying out characteristic screening through a lasso regression method, a multivariate statistical method of unbiased variable selection and a BORUTA method to obtain candidate characteristics related to the ankylosing spondylitis ossification progress; carrying out data standardization processing, principal component analysis and orthogonal partial least-partial-square discriminant analysis on the characteristic ion peak table, and carrying out pathway enrichment analysis to obtain differential metabolites; an LR model, an RF model and an SVM model used for distinguishing the ossification progress speed of the AS patient are established based on known risk factors, candidate features and differential metabolites, the efficiency of the three models is verified and compared, and finally the ankylosing spondylitis ossification progress prediction models with the optimal efficiency are obtained and serve as the LR model and the RF model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to the field of ankylosing spondylitis ossification progression prediction technology based on metabolomics and artificial intelligence. Background Art

[0002] Ankylosing spondylitis (AS) is a chronic, progressive rheumatic disease. Structural damage to the spine can lead to limited spinal mobility, which seriously affects the patient's quality of life. Structural damage to the spine in AS includes bone destruction and abnormal bone formation, which is one of its most significant clinical manifestations. It is significantly correlated with the disease activity and spinal mobility of AS patients. Currently, the Modified Stoke Ankylosing Spondylitis Spinal Score (mSASSS) is considered to be the gold standard for evaluating spinal structural damage in AS. Both the Outcome Measures in Rheumatology Clinical Trials (OMERACT) and the Assessment in Ankylosing Spondylitis International Society (ASAS) recommend mSASSS as the preferred scoring method for AS imaging progression. Establishing a prediction model to distinguish people with different progression rates can enable clinicians to conduct targeted interventions on patients earlier and achieve individualized precision treatment. Summary of the invention

[0003] In order to overcome the above technical defects, the purpose of the present invention is to provide a method for constructing an ankylosing spondylitis ossification progression prediction model based on metabolomics and artificial intelligence, which comprises:

[0004] Step S1: Collecting plasma samples from venous blood of several patients with ankylosing spondylitis who meet the inclusion and exclusion criteria;

[0005] Step S2: performing metabolomics analysis on the plasma samples, performing peak extraction, data normalization and data screening to obtain a characteristic ion peak table, and performing compound identification on the characteristic ion peaks to obtain differential compounds associated with the progression of ankylosing spondylitis ossification;

[0006] Step S3: First, perform receiver operating curve analysis on the differential compounds one by one to screen out metabolites with an area under the curve greater than 0.6, and then perform feature screening using the lasso regression method, the multivariate statistical method for unbiased variable selection, and the BORUTA method, and select metabolites that are selected at least twice in the three methods as candidate features for subsequent modeling;

[0007] Step S4: Import the characteristic ion peak table into the MetaboAnalyst website for data standardization, perform T-test and Fold Change calculation between the rapid progression group and the slow progression group, and perform principal component analysis and orthogonal partial least squares discriminant analysis, and screen differential compounds according to P<0.05 and OPLS-DA VIP>1; then perform pathway enrichment analysis on the MetaboAnalyst website based on the differential compounds to obtain differential metabolites;

[0008] Step S5: Based on the known risk factors, the candidate features and the differential metabolites, an LR model (i.e., a Logistics regression model), an RF model (i.e., a random forest model) and an SVM model (i.e., a support vector machine model) are established to distinguish the ossification progression rate of AS patients, and the effectiveness of these three models are verified and compared. Finally, the most effective ankylosing spondylitis ossification progression prediction models are obtained, which are the LR model and the RF model.

[0009] As an important branch of systems biology, metabolomics is an emerging omics technology that has emerged after genomics and proteomics. Using non-targeted methods, it can provide an overall metabolic picture and identify previously unknown molecules involved in the occurrence and progression of AS.

[0010] Further, in step S1, the inclusion and exclusion criteria are:

[0011] (1) Inclusion criteria:

[0012] 1) Sign the informed consent;

[0013] 2) Patients aged ≥18 years when signing the informed consent form;

[0014] 3) diagnosed with AS according to the New York criteria for ankylosing spondylitis;

[0015] 4) The medical records contain at least one full spine X-ray result.

[0016] (2) Exclusion criteria:

[0017] 1) A history of inflammatory arthritis caused by causes other than AS, or any inflammatory arthritis occurring before the age of 17;

[0018] 2) Participants in other clinical trials;

[0019] 3) Patients with previous or concurrent malignant tumors;

[0020] 4) Patients with or combined immunodeficiency disease, diabetes, gout or other diseases that may affect metabolism;

[0021] 5) Long-term use of drugs that may affect metabolism;

[0022] 6) Those with incomplete medical records;

[0023] 7) Blood samples lacking plasma or blood sediment;

[0024] 8) The patient refused to participate in this study.

[0025] Further, in step S2, the metabolomics analysis includes:

[0026] Step S2.1: Plasma sample pretreatment: Plasma sample pretreatment adopts protein precipitation method, taking plasma samples stored at -80℃, slowly dissolving at 4℃, vortexing until completely liquid, accurately aspirating 100μL plasma after thawing, adding four times volume of cold methanol solution (containing 500ng / ml 2-chlorophenylalanine), vortexing for 1min, standing for 10min, and then centrifuging (4℃, 14000r / min, 15min), taking 200μL of supernatant, freeze-drying and ultrasonically re-dissolving with 200μL acetonitrile-water solution (8:2, v / v), vortexing for 1min, and centrifuging again (4℃, 14000r / min, 15min), taking 150μL of supernatant and placing it in a sample injection vial for use, and operating on ice throughout the process to maintain low temperature;

[0027] Step S2.2: Monitor the sample analysis sequence: Take 10 μL of the supernatant of each sample after reconstitution and centrifugation, vortex mix, and prepare the quality control (QC) sample. Before the start of the sequence, repeat the injection of blank solvent 5 times and QC sample 3 times to balance the system and chromatographic column. After that, inject a QC sample every 10 samples in the sequence to evaluate the instrument stability and perform data correction;

[0028] Step S2.3: UHPLC-Q-TOF / MS detection: (1) Chromatographic conditions: The chromatographic column is WATERS, ACQUITY UPLCHSS T3 chromatographic column (2.1 mm×100 mm, 1.8 μm), the column temperature is 40°C, the injection chamber temperature is 15°C; the flow rate is 0.3 mL / min; the injection volume is 3 μL. The mobile phase A is pure water (0.1% formic acid), the mobile phase B is acetonitrile (0.1% formic acid), and the chromatographic conditions for gradient elution are shown in the following table:

[0029] Time (min) A(%) B(%) 0 95 5 12 80 20 17 60 40 30 5 95 35 5 95 36 95 5 38 95 5

[0030] (2) Mass spectrometry conditions: The acquisition system was Agilent UHPLC-Q-TOF / MS, the drying gas and nebulizing gas were both nitrogen, the collision gas was high-purity nitrogen, and the mass spectrometry parameters were shown in the following table:

[0031]

[0032]

[0033] Step S2.4: Data processing and analysis: (1) Peak extraction and normalization: Agilent QualitativeAnalysis B.10.00 software was used to check the overlap of the total ion current chromatogram superposition, preliminarily determine the stability of the sample sequence and instrument data acquisition, and remove samples with large deviations and missing internal standards; the raw data was converted into .abf format using Abf converter software, and then imported into MS-DIAL software for peak extraction, alignment, removal of batch effects, etc. In positive and negative ion modes, each group of data and blank group data were imported into the software respectively, and data normalization was performed using internal standard peak area correction to obtain a characteristic ion peak table containing characteristic ion peak retention time, mass-to-charge ratio and other information; according to the 80% principle, characteristic ion peaks with a frequency greater than 80% in non-QC samples were retained, and characteristic ion peaks with a frequency greater than half in QC samples and RSD>0.3 were removed; (2) Compound identification: MS FINDER software was used to decompose the .msp format file exported from MS-DIAL, and then based on the Human Metabolome Database, METLIN, the mass bank Web site, the database on the LIPID MAPS website was used for batch comparison, and the comparison and scoring of the secondary fragment ions with the standard library were performed according to the mass-to-charge ratio and relative abundance, and the compounds with the best matching degree were screened according to the comparison scores.

[0034] Further, in step S3, the candidate features are: LysoPC (22:2 (13Z, 16Z)), quinic acid, ascorbic acid-2-sulfate, LysoPC (18:1 (9Z)).

[0035] Further, in step S4 and step S5, the differential metabolites are: 1) cysteine ​​and methionine metabolism; 2) pyrimidine metabolism; 3) biosynthesis of valine, leucine and isoleucine; and 4) arginine and proline metabolism.

[0036] Further, step S5 includes:

[0037] Step S5.1: Divide the original data of the case materials of several patients into a training set and a validation set, and perform data preprocessing on each set;

[0038] Step S5.2: Using the forward-backward method, the known risk factors, the candidate features, and the differential metabolites were used to establish LR models in R4.3.2, and the optimal feature combination was selected after comparing the performance of the models in the validation set; then, using the R package, randomly selecting samples in the training set and adjusting the model parameters by the grid optimization search algorithm, the LR model, the RF model, and the SVM model were established respectively;

[0039] Step S5.3: Then, based on the validation set, a variety of model evaluation indicators are used to comprehensively determine the effectiveness of the three models through the average value of 20 cycles of validation, totaling 100 modeling times, and the models are compared. Finally, the best ankylosing spondylitis ossification progression prediction model is obtained, which is the LR model and / or the RF model.

[0040] Furthermore, in step S5.3, the multiple model evaluation indicators include AUC, accuracy, precision, recall, F1 score, sensitivity, specificity, and Matthews correlation coefficient; the Z test is used to perform pairwise comparisons between the AUC values ​​and the net reclassification improvement index NRI between the models.

[0041] Furthermore, the model parameters of the LR model are default parameters; the optimal parameters of the RF model are: mtry=3; ntree=800; the optimal parameters of the SVM model are: gamma=0.0063, C=355.6559, kernel function='radial'.

[0042] Compared with the prior art, the above technical solution has the following beneficial effects:

[0043] This application uses metabolomics analysis to analyze the metabolic differences between different AS ossification progression groups, and finds that compared with the slow progression group, multiple metabolic pathways in AS patients in the rapid progression group were disturbed, among which the relative levels of methionine and L-cystine related to folate metabolism increased significantly, and for the first time discovered the relationship between folate metabolism disorder and AS ossification progression. Subsequently, through machine learning algorithms, differential metabolites associated with AS ossification progression were screened, and a model for distinguishing the ossification progression rate of AS patients was established based on multi-dimensional information such as patient genotype, metabolome, and clinical characteristics. The overall diagnostic efficiency of the model is relatively good, which helps to identify AS patients with rapid ossification progression in clinical practice and intervene early. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is an overlay of the total ion current of plasma, where A: ESI+; B: ESI-.

[0045] Figure 2This is the PCA graph of plasma samples, where A: ESI+; B: ESI-; C: OPLS-DA graph of plasma samples; D: permutation test graph for model validation; *F: rapid progression group; S: slow progression group; QC: quality control.

[0046] Figure 3 Venn diagrams of the results for the three feature screening methods.

[0047] Figure 4 The ROC curves of four candidate metabolites, including: A: LysoPC (22:2 (13Z, 16Z)); B: Quinic acid; C: Ascorbic acid-2-sulfate; D: LysoPC (18:1 (9Z)).

[0048] Figure 5 The relative contents of the four metabolites in the two groups.

[0049] Figure 6 The following are the results of pathway enrichment analysis, including: 1: biosynthetic pathways of valine, leucine and isoleucine; 2: cysteine ​​and methionine metabolism; 3: pyrimidine metabolism pathway; 4: arginine and proline metabolism pathway; 5: biosynthetic pathways of phenylalanine, tyrosine and tryptophan.

[0050] Figure 7 A in the figure is the ROC curve of different models in the training set; Figure 7 B in the figure is the ROC curve of different models in the validation set; LR: Logistics regression model; RF: Random Forest model; SVM: Support Vector Machine model. DETAILED DESCRIPTION

[0051] The advantages of the present invention are further described below in conjunction with the accompanying drawings and specific embodiments. Those skilled in the art should understand that the following specific description is illustrative rather than restrictive, and should not be used to limit the scope of protection of the present invention.

[0052] This embodiment provides a method for constructing an ankylosing spondylitis ossification progression prediction model based on metabolomics and artificial intelligence, which includes steps S1 to S5:

[0053] Step S1: Collecting plasma samples from venous blood of several ankylosing spondylitis patients who meet the inclusion and exclusion criteria.

[0054] (I) Research subjects and inclusion and exclusion criteria

[0055] 1. Research subjects

[0056] This study was a retrospective, observational study, and the subjects were AS patients who visited the Department of Rheumatology and Immunology and the Department of Clinical Genetics of the Second Affiliated Hospital of Naval Medical University from September 2022 to January 2024. This study followed the relevant provisions and principles of the Declaration of Helsinki, and all patients or their legal guardians signed informed consent.

[0057] 2. Inclusion and Exclusion Criteria

[0058] (1) Inclusion criteria

[0059] 1) Sign the informed consent;

[0060] 2) Patients aged ≥18 years when signing the informed consent form;

[0061] 3) diagnosed with AS according to the New York criteria for ankylosing spondylitis;

[0062] 4) The medical records contain at least one full spine X-ray result.

[0063] (2) Exclusion criteria

[0064] 1) A history of inflammatory arthritis caused by causes other than AS (including but not limited to rheumatoid arthritis, mixed connective tissue disease, systemic lupus erythematosus, dermatomyositis, etc.), or any inflammatory arthritis occurring before the age of 17;

[0065] 2) Participants in other clinical trials;

[0066] 3) Patients with a history of or concurrent malignant tumors, such as breast cancer, gastric cancer, etc.;

[0067] 4) Patients with or combined immunodeficiency disease, diabetes, gout or other diseases that may affect metabolism;

[0068] 5) Those who take anticoagulants, hormones and other drugs that may affect metabolism for a long time;

[0069] 6) Those with incomplete medical records;

[0070] 7) Blood samples lacking plasma or blood sediment;

[0071] 8) The patient refused to participate in this study.

[0072] (II) Sample collection and preservation

[0073] 1) The patient fasts for 12-14 hours and blood is collected on an empty stomach in the morning;

[0074] 2) Avoid strenuous exercise before blood collection, and collect blood after resting for 15 minutes;

[0075] 3) Use ethylenediaminetetraacetic acid (EDTA) anticoagulant tubes to collect 2-3 ml of venous blood from the patient. After collection, gently invert and shake the blood collection tube 4-6 times and place it vertically in a test tube rack. At the same time, record the patient's name, outpatient / inpatient number, age, gender, etc.;

[0076] 4) The collected venous blood was transferred to the sample bank within 2 hours. After centrifugation at 3000 r / min for 5 minutes, the upper plasma and lower blood sediment were taken and frozen in a -80°C refrigerator for testing.

[0077] Step S2: Perform metabolomics analysis on the plasma samples, perform peak extraction, data normalization and data screening to obtain a characteristic ion peak table, and perform compound identification on the characteristic ion peaks to obtain differential compounds related to the progression of ankylosing spondylitis ossification.

[0078] (III) Instruments and reagents

[0079] 1. Instruments and Equipment

[0080] The metabolomics experiment was carried out based on the UHPLC-Q-TOF / MS system. The instruments and equipment used are shown in Table 1.

[0081] Table 1 Main instruments and equipment for metabolomics experiments

[0082]

[0083] 2. Reagents

[0084] The main reagents used in this study are listed in Table 2.

[0085] Table 2 Main reagents for metabolomics experiments

[0086]

[0087]

[0088] (IV) Metabolomics analysis

[0089] 1. Pretreatment of plasma samples

[0090] The plasma samples were pre-treated by protein precipitation. Plasma samples stored at -80°C were slowly dissolved at 4°C and vortexed until completely liquid. After thawing, 100 μL of plasma was accurately aspirated, and a four-fold volume of cold methanol solution (containing 500 ng / ml of 2-chlorophenylalanine) was added. The solution was vortexed for 1 min, allowed to stand for 10 min, and then centrifuged (4°C, 14000 r / min, 15 min). 200 μL of the supernatant was taken, freeze-dried, and ultrasonically reconstituted with 200 μL of acetonitrile-water solution (8:2, v / v), vortexed for 1 min, and centrifuged again (4°C, 14000 r / min, 15 min). 150 μL of the supernatant was taken and placed in a sample injection vial for later use. The whole process was performed on ice to maintain low temperature.

[0091] 2. Monitoring sample analysis sequence

[0092] Take 10 μL of the supernatant of each sample after reconstitution and centrifugation, vortex mix, and prepare the quality control (QC) sample. Before the start of the sequence, repeat the injection of blank solvent 5 times and QC sample 3 times to balance the system and chromatographic column. After that, a QC sample is injected every 10 samples in the sequence to evaluate the stability of the instrument and perform data correction.

[0093] 3. UHPLC-Q-TOF / MS detection conditions

[0094] (1) Chromatographic conditions

[0095] The chromatographic column was a WATERS, ACQUITY UPLC HSS T3 column (2.1 mm × 100 mm, 1.8 μm), the column temperature was 40°C, the injection chamber temperature was 15°C; the flow rate was 0.3 mL / min; the injection volume was 3 μL. The mobile phase A was pure water (0.1% formic acid), and the mobile phase B was acetonitrile (0.1% formic acid). The chromatographic conditions for gradient elution are shown in Table 3.

[0096] Table 3 Gradient elution time series

[0097] Time (min) A(%) B(%) 0 95 5 12 80 20 17 60 40 30 5 95 35 5 95 36 95 5 38 95 5

[0098] (2) Mass spectrometry conditions

[0099] The acquisition system was Agilent UHPLC-Q-TOF / MS, the drying gas and nebulizing gas were both nitrogen, and the collision gas was high-purity nitrogen. Specific mass spectrometry parameters are shown in Table 4.

[0100] Table 4 Mass spectrometry detection parameters

[0101] Test conditions parameter Ion mode ESI+ / ESI- Spray voltage 3500V / 3000V Sheath gas temperature 350℃ Drying gas temperature 325℃ Drying air flow rate 11L / min Atomization pressure 45psi Capture range 100-1700m / z Collision Energy 30V

[0102] 4. Data processing and analysis

[0103] (1) Peak extraction and normalization

[0104] The Agilent Qualitative Analysis B.10.00 software was used to check the overlap of the total ion current chromatogram (TIC) superposition, to preliminarily determine the stability of the sample sequence and instrument data acquisition, and to remove samples with large deviations and missing internal standards. Figure 1 As shown, it can be seen that the TIC graphs of all samples are stable, with good overlap and no outlier samples.

[0105] The data processing process is divided into format conversion, peak extraction, data normalization and data screening. The raw data was converted into .abf format by Abf converter software, and then imported into MS-DIAL software for peak extraction, alignment, removal of batch effects, etc. In positive and negative ion mode, each group of data and blank group data were imported into the software respectively, and the internal standard peak area correction was used for data normalization to obtain a characteristic ion peak table containing characteristic ion peak retention time, mass-to-charge ratio and other information. According to the 80% principle, characteristic ion peaks with a frequency greater than 80% in non-QC samples were retained, and characteristic ion peaks missing half in QC samples and with RSD>0.3 were removed.

[0106] The peak table exported by MSDial was further screened, with the number of data values ​​in QC samples that were not 0 > 50%, the number of data values ​​in each group in non-QC samples that were not 0 > 80%, and the RSD of data values ​​in QC samples < 0.3 as screening criteria. 14,836 and 2,347 ions were eliminated under positive and negative ions, respectively, and the remaining ions were imported into the metaboAnalyst website in .csv format for data standardization and scaling, followed by multivariate statistical analysis and difference analysis.

[0107] PCA plots Figure 2 The QC samples in A and B overlap well, indicating that the detection sequence is stable and the data is reliable. The PCA graph has a poor degree of differentiation, there is obvious overlap between the two groups, the discrete trend is not obvious, and the dispersion within each group is also large. The reason may be the uncontrollability of other factors in clinical samples. At the same time, endogenous substances have not been screened, and there are many interferences, which is a normal phenomenon in clinical sample analysis.

[0108] OPLS-DA score plot ( Figure 2 C) shows that the rapid progression group (F group) and the slow progression group (S group) are well distinguished, and the model validation permutation test diagram ( Figure 2D), R2 represents the ability of the corresponding principal component in the model to explain the overall variation, and Q2 represents the ability of the principal component to predict the variation of the model. Generally speaking, the values ​​of these two parameters must be greater than 0.5 to indicate that the model is successfully constructed. In the figure, both R2 and Q2 are greater than 0.5, and P<0.05, indicating that the model fitting accuracy is good.

[0109] (2) Compound identification

[0110] The .msp format files exported from MS-DIAL were decomposed using MS FINDER software, and then batch comparisons were performed based on databases on websites such as Human Metabolome Database (www.hmdb.ca), METLIN (http: / / metlin.scripps.edu), the massbank Web site (http: / / www.massbank.jp / en / database.html), and LIPID MAPS (http: / / www.lipidmaps.org). The secondary fragment ions were compared and scored with the standard library based on the mass-to-charge ratio and relative abundance, and the compounds with the best match were screened based on the comparison scores.

[0111] The F group and the S group were compared pairwise, and the differential metabolites were screened with the standard of P<0.05 of the single factor T-test and VIP>1 of OPLS-DA. The primary and secondary fragments were identified in MS Finder 3.50 software, and their numbers were searched by HMDB. A total of 117 differential compounds were identified in the positive and negative ion modes, of which 61 were upregulated in the F group and 56 were downregulated.

[0112] 5. Statistical analysis

[0113] 1. Statistical software

[0114] The statistical analysis and graphics software used included SPSS26.0.0, R 4.3.2, and Graphpad Prism 8.

[0115] 2. Data Description and Analysis

[0116] For count data, frequency, rate or composition ratio were used for description, and chi-square test or Fisher's exact test was used for statistical analysis; for measurement data, Shapiro-Wilk test was used for normality test, and those that met normal distribution were expressed as mean ± standard deviation. The data were described, and the median (quartile) M (P0.25, P0.75) was used for those that did not meet the requirements. For two groups of data that met the normal distribution and had equal variance, the independent sample t test was used for analysis, and non-parametric tests were used for those that did not meet the requirements. The statistical results were considered statistically significant when P < 0.05.

[0117] Step S3: First, perform receiver operating curve analysis on the differential compounds one by one to screen out metabolites with an area under the curve greater than 0.6, and then perform feature screening through lasso regression method, multivariate statistical method of unbiased variable selection, and BORUTA method, and select metabolites selected at least twice in the three methods as candidate features for subsequent modeling.

[0118] Metabolomics data often contain thousands of metabolite information and have the characteristics of small sample and high dimensionality. Therefore, it is necessary to filter the metabolite features before modeling to reduce the risk of model overfitting and improve its generalization ability. For the differential compounds obtained by metabolomics, the receiver operating curve (ROC) analysis is performed one by one to screen out metabolites with an area under the curve (AUC) greater than 0.6, and then feature screening is performed through lasso (Lease Absolute Shrinkage and Selection Operator, LASSO) regression, multivariate statistical methods with unbiased variable selection in R (MUVR), BORUTA and other methods, and metabolites selected at least twice in the three methods are selected as candidate features for subsequent modeling.

[0119] Since different variable screening methods look at the problem from different angles and have their own advantages and limitations, this study incorporates three algorithms and compares their screening results to screen out more efficient and concise features.

[0120] Metabolomics feature screening based on multivariate statistical analysis and machine learning algorithms:

[0121] ROC analysis was performed on the 117 differential compounds obtained by multivariate statistical analysis, and a total of 100 compounds with AUC>0.6 were screened as candidate features for feature screening. In order to reduce the collinearity between variables and select variables with good prediction effects, this study used three algorithms to screen the features, and finally selected 4 features that were selected at least twice in the three methods, representing four metabolites of lysophospholipids (22:2(13Z,16Z)), quinic acid, ascorbic acid-2-sulfate, and lysophospholipids (18:1(9Z)) ( Figure 3 The ROC curve and the relative content of metabolites between the two groups are shown in Figure 4 and Figure 5 .

[0122] Step S4: Import the characteristic ion peak table into the MetaboAnalyst website for data standardization, perform T-test and Fold Change calculation between the rapid progression group and the slow progression group, and perform principal component analysis and orthogonal partial least squares discriminant analysis, and screen differential compounds based on P<0.05 and OPLS-DA VIP>1; then, based on the differential compounds, perform pathway enrichment analysis in the MetaboAnalyst website to obtain differential metabolites.

[0123] Univariate and multivariate statistical analysis of metabolomics data:

[0124] The characteristic ion peak table was imported into the MetaboAnalyst website (https: / / www.metaboanalyst.ca / ) for data standardization. T-test and FoldChange (FC) were calculated between the fast progression group (F group) and the slow progression group (S group). Principal component analysis (PCA) and orthogonal partial least squares discrimination analysis (OPLS-DA) were performed. Differential compounds were screened according to P < 0.05 and OPLS-DA VIP > 1 as candidate features for subsequent pathway enrichment analysis and feature screening. PCA is a statistical method that converts a set of variables that may be correlated into a set of linearly uncorrelated variables through orthogonal transformation. PCA can comprehensively reflect the overall metabolic differences between samples in each group and reveal the degree of variability between samples within the group. OPLS-DA is a supervised discriminant analysis statistical method. This method uses partial least squares regression technology to construct a relationship model between metabolite expression and sample category, thereby realizing the prediction of sample category. The stability of the detection sequence was determined by the degree of clustering of the QC samples in the PCA graph. Metabolic pathway enrichment analysis was performed on the MetaboAnalyst website to screen metabolic pathways with P < 0.05 and analyze them in combination with the biological significance of the pathways. The differential compounds between group F and group S were uploaded to MetaboAnalyst for pathway analysis to obtain differential metabolic pathways, which were then analyzed in combination with the biological significance of the pathways. The pathway analysis results are shown in Figure 1. Figure 6 As shown, there are 4 pathways with P values ​​less than 0.05, namely cysteine ​​and methionine metabolism, pyrimidine metabolism, biosynthesis of valine, leucine and isoleucine, and arginine and proline metabolism.

[0125] Step S5: Based on the known risk factors (MTHFR C677>T genotype, hemoglobin (Hb) and alkaline phosphatase (ALP) levels at patient baseline), the candidate features and the differential metabolites, an LR model, a RF model and a SVM model were established to distinguish the ossification progression rate of AS patients, and the effectiveness of the three models were verified and compared. Finally, the LR model and the RF model were obtained as the most effective prediction models for ankylosing spondylitis ossification progression.

[0126] Establishment, validation and comparison of the diagnostic model for AS ossification progression:

[0127] (1) Data division and preprocessing

[0128] Since some data in the patient's medical records are missing, it will bias the result estimation and cause the failure of outcome prediction for some patients. Therefore, when fitting the LR model, for relevant factors with missing values ​​not exceeding 30%, the mean interpolation method is used to fill the missing values. Using the five-fold cross-validation method, the original data set is divided into five subsets, one of which is taken as the validation set, and the other four are combined as the training set, and data preprocessing is performed separately to avoid data leakage. Then the model is established, and the cross-validation is repeated five times, so that a total of five models are established.

[0129] (2) Establishment and verification of diagnostic model

[0130] The risk factors found in the first part and the candidate features of metabolomics, as well as the differential metabolites obtained by pathway enrichment analysis, were used to establish LR models in R 4.3.2 using the forward-backward method. The optimal feature combination was selected after comparing the performance of the models in the validation set. Then, R packages such as stars, randomForest, and e1071 were used to adjust the model parameters in the training set by randomly selecting samples (Bootstrap Sample) and grid optimization search algorithms, and LR, RF, and SVM models were established respectively to estimate the progression grouping of each patient, compare the actual grouping results, draw ROC curves, and calculate the average AUC value of each model. The validation was performed in the validation set data, and the average value of the validation of 20 cycles and a total of 100 modelings was used to comprehensively determine the effectiveness of each model. In this study, the training set and validation set came from the same cohort study, which was an internal validation.

[0131] (3) Model evaluation indicators

[0132] A variety of model indicators, including AUC, accuracy, precision, recall, F1 score, sensitivity, specificity, and Matthews Correlation Coefficient (MCC), are used to comprehensively evaluate the effectiveness of each model.

[0133] (4) Comparison between models

[0134] The Z test was used to compare the AUC values ​​and net reclassification improvement index (NRI) between the models.

[0135] The P value is obtained by the Excel formula: P = (1-NORMSDIST(Z)) × 2.

[0136] Among them, SE1 and SE2 are the standard errors of AUC1 and AUC2 respectively. The calculation formula is:

[0137] SE = (AUC - 95% confidence interval lower limit) / 1.96

[0138] NRI is often used to compare the predictive ability of two models. Its value is between -2 and 2. If it is greater than 0, it is considered that the new model is improved compared to the old model; if it is equal to 0, it is considered that the new model has no improvement; if it is less than 0, it is a negative improvement.

[0139] NRI=(Sensitivity2-Sensitivity1)+(Specificity2-Specificity1)

[0140] Comparison of models incorporating different combinations of dimensional features:

[0141] The performance of the LR model constructed by combining different dimensional features such as clinical indicators only, metabolites, clinical indicators + pathway enriched metabolites, clinical indicators + machine learning screened metabolites, clinical indicators + pathway enriched metabolites + machine learning screened metabolites in the validation set was comprehensively compared. The results showed that the LR model with only clinical indicators had an F1 score of 0.76 (95% CI: 0.742-0.777), MCC of 0.36 (95% CI: 0.314-0.411), and validation set AUC of 0.76 (95% CI: 0.739-0.788). The LR model that included clinical indicators and machine learning screened metabolites had an F1 score of 0.74 (95% CI: 0.722-0.763), MCC of 0.38 (95% CI: 0.336-0.423), and validation set AUC of 0.76 (95% CI: 0.741-0.784). Considering that the purpose of establishing the model is to better distinguish the two groups and identify the rapid progression group, MCC can better take into account the unbalanced data set and measure the classification performance of the model. Therefore, a feature combination with high Recall, MCC and validation set AUC, namely clinical indicators + machine learning screening metabolites, was selected as the feature included in the subsequent model establishment. The evaluation indicators of the model constructed by each included feature combination are shown in Tables 5 and 6.

[0142] Table 5 Evaluation indicators of LR models with different feature combinations (conventional indicators)

[0143]

[0144]

[0145] *a: clinical indicators; b: metabolites obtained by pathway enrichment analysis; c: metabolites screened by machine learning

[0146] Table 6 Evaluation indicators of LR models with different feature combinations (machine learning indicators)

[0147]

[0148] *a: clinical indicators; b: metabolites obtained by pathway enrichment analysis; c: metabolites screened by machine learning

[0149] The parameters of the models (RF, SVM and LR models) with the best performance in this application are as follows:

[0150] LR model: default parameters (family = binomial (link = "logit");

[0151] The best parameters for the SVM model are: gamma = 0.0063, C = 355.6559, kernel function = 'radial';

[0152] The optimal parameters of the RF model are: mtry=3; ntree=800.

[0153] Comparison of evaluation indicators of each model:

[0154] Clinical indicators + machine learning were used to screen metabolite indicators, and LR, RF, and SVM models were constructed respectively. The comparison of the evaluation indicators of the three models in the validation set is shown in Table 7 and Table 8, and the comparison of ROC curves in the training set and the validation set is shown in Figure 7 As can be seen from the chart, compared with other models, in terms of discrimination ability, the RF model has the largest AUC value in the validation set. The ROC curves of the LR model and the RF model are close to the upper left corner, and their AUC values ​​are 0.78 (95% CI: 0.755-0.801) and 0.76 (95% CI: 0.741-0.784), respectively. Although the SVM model has a higher MCC, its AUC in the validation set is 0.63 (95% CI: 0.613-0.643), which is not ideal compared with the other two models. The Recall, F1 score, and MCC of the LR model are better than the other two models. Taking all factors into consideration, it is believed that the LR and RF models have good discrimination abilities.

[0155] Table 7 Comparison of evaluation indicators of different models in the internal validation set (conventional indicators)

[0156]

[0157] LR: Logistic regression model; RF: Random forest model; SVM: Support vector machine model

[0158] Table 8 Comparison of evaluation indicators of different models in the internal validation set (machine learning indicators)

[0159]

[0160] LR: Logistic regression model; RF: Random forest model; SVM: Support vector machine model

[0161] 5. Comparison of model discrimination and accuracy

[0162] The discrimination and accuracy of the three models were compared. As shown in Table 9, the AUC values ​​of the LR and RF models were significantly greater than those of the SVM model, while the AUC of the RF model was greater than that of the LR model, but there was no significant difference; the accuracy of the LR model was improved by 8.5% and 15.6% relative to the RF model and the SVM model, respectively, but there was no significant difference between the LR and RF models in terms of accuracy. Overall, the model performance of the LR and RF models has its own advantages and disadvantages, but the difference is not statistically significant. The sensitivity of the LR model in the internal validation set reached 0.73 (95% CI: 0.709-0.746), and the area under the ROC curve was 0.76 (95% CI: 0.741-0.784). Its overall diagnostic efficacy is relatively good, and it has a certain effect on identifying patients with rapidly progressive AS.

[0163] Table 9 Comparison of the discrimination and accuracy of the models in the internal validation set

[0164] AUC(Z,P) NRI(Z,P) RF vs LR Z = 0.637, P = 0.738 NRI=-0.085, Z=-1.96, P=0.05 RF vs SVM Z=9.72,P<0.0001 NRI=0.07, Z=1.71, P=0.09 LR vs SVM Z=9.99,P<0.001 NRI=0.156, Z=4.39, P<0.001

[0165] LR: Logistic regression model; RF: Random forest model; SVM: Support vector machine model; AUC: Area under the curve; NRI: Net reclassification improvement index

[0166] In summary, metabolomics analysis revealed that compared with the slow progression group, multiple metabolic pathways in AS patients in the rapid progression group were disturbed, among which the relative levels of methionine and L-cystine related to folate metabolism increased significantly, and the relationship between folate metabolism disorder and AS ossification progression was discovered for the first time. Subsequently, the machine learning algorithm was used to screen differential metabolites associated with AS ossification progression, and a model for distinguishing the ossification progression rate of AS patients was established based on multi-dimensional information such as patient genotype, metabolome, and clinical characteristics. The overall diagnostic efficacy of the model is relatively good, which is helpful for clinical identification of AS patients with rapid ossification progression and early intervention.

[0167] It should be noted that the embodiments of the present invention have better practicability and do not impose any form of limitation on the present invention. Any technician familiar with the field may use the technical content disclosed above to change or modify it into an equivalent effective embodiment. However, any modification or equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A method for constructing a prediction model for ossification progression of ankylosing spondylitis based on metabolomics and artificial intelligence, characterized in that: include: Step S1: Collecting plasma samples from venous blood of several patients with ankylosing spondylitis who meet the inclusion and exclusion criteria; Step S2: performing metabolomics analysis on the plasma samples, performing peak extraction, data normalization and data screening to obtain a characteristic ion peak table, and performing compound identification on the characteristic ion peaks to obtain differential compounds associated with the progression of ankylosing spondylitis ossification; Step S3: First, perform receiver operating curve analysis on the differential compounds one by one to screen out metabolites with an area under the curve greater than 0.6, and then perform feature screening using the lasso regression method, the multivariate statistical method for unbiased variable selection, and the BORUTA method, and select metabolites that are selected at least twice in the three methods as candidate features for subsequent modeling; Step S4: Import the characteristic ion peak table into the MetaboAnalyst website for data standardization, perform T-test and Fold Change calculation between the rapid progression group and the slow progression group, and perform principal component analysis and orthogonal partial least squares discriminant analysis, and screen differential compounds according to P<0.05 and OPLS-DA VIP>1; then perform pathway enrichment analysis on the MetaboAnalyst website based on the differential compounds to obtain differential metabolites; Step S5: Based on the known risk factors, the candidate features and the differential metabolites, an LR model, an RF model and an SVM model are established to distinguish the ossification progression rate of AS patients, and the effectiveness of these three models are verified and compared. Finally, the most effective ankylosing spondylitis ossification progression prediction models are obtained, which are the LR model and the RF model.

2. The method for constructing an ankylosing spondylitis ossification progression prediction model based on metabolomics and artificial intelligence according to claim 1, characterized in that: In step S1, the inclusion and exclusion criteria are: (1) Inclusion criteria: 1) Sign the informed consent; 2) Patients aged ≥18 years when signing the informed consent form; 3) diagnosed with AS according to the New York criteria for ankylosing spondylitis; 4) The medical records contain at least one full spine X-ray result. (2) Exclusion criteria: 1) A history of inflammatory arthritis caused by causes other than AS, or any inflammatory arthritis occurring before the age of 17; 2) Participants in other clinical trials; 3) Patients with previous or concurrent malignant tumors; 4) Patients with or combined immunodeficiency disease, diabetes, gout or other diseases that may affect metabolism; 5) Long-term use of drugs that may affect metabolism; 6) Those with incomplete medical records; 7) Blood samples lacking plasma or blood sediment; 8) The patient refused to participate in this study.

3. The method for constructing an ankylosing spondylitis ossification progression prediction model based on metabolomics and artificial intelligence according to claim 1, characterized in that: In step S2, metabolomics analysis includes: Step S2.1: Plasma sample pretreatment: Plasma sample pretreatment adopts protein precipitation method, taking plasma samples stored at -80℃, slowly dissolving at 4℃, vortexing until completely liquid, accurately aspirating 100μL plasma after thawing, adding four times the volume of cold methanol solution containing 500ng / ml of 2-chlorophenylalanine, vortexing for 1min, standing for 10min, then centrifuging at 4℃ and 14000r / min for 15min, taking 200μL of supernatant, freeze-drying, and ultrasonically re-dissolving with 200μL acetonitrile-water solution, vortexing for 1min, and centrifuging again at 4℃ and 14000r / min for 15min, taking 150μL of supernatant and placing it in a sample injection vial for standby use, and operating on ice throughout the process to maintain low temperature; Step S2.2: Monitor the sample analysis sequence: Take 10 μL of the supernatant of each sample after reconstitution and centrifugation, vortex mix, and prepare the quality control QC sample. Before the start of the sequence, repeat the injection of blank solvent 5 times and QC sample 3 times to balance the system and chromatographic column. After that, inject a QC sample every 10 samples in the sequence to evaluate the instrument stability and perform data correction; Step S2.3: UHPLC-Q-TOF / MS detection: (1) Chromatographic conditions: the chromatographic column is WATERS, ACQUITY UPLC HSST3 chromatographic column, 2.1 mm×100 mm, 1.8 μm, the column temperature is 40°C, and the injection chamber temperature is 15°C; flow rate: 0.3 mL / min; the injection volume is 3 μL; the mobile phase A is pure water, the mobile phase B is acetonitrile, and the chromatographic conditions of the gradient elution are shown in the following table: (2) Mass spectrometry conditions: The acquisition system was Agilent UHPLC-Q-TOF / MS, the drying gas and nebulizing gas were both nitrogen, the collision gas was high-purity nitrogen, and the mass spectrometry parameters were shown in the following table: ; Step S2.4: Data processing and analysis: (1) Peak extraction and normalization: Agilent QualitativeAnalysis B.10.00 software was used to check the overlap of the total ion current chromatogram superposition, preliminarily determine the stability of the sample sequence and instrument data acquisition, and remove samples with large deviations and missing internal standards; the raw data was converted into .abf format using Abf converter software, and then imported into MS-DIAL software for peak extraction, alignment, removal of batch effects, etc. In positive and negative ion modes, each group of data and blank group data were imported into the software respectively, and data normalization was performed using internal standard peak area correction to obtain a characteristic ion peak table containing characteristic ion peak retention time, mass-to-charge ratio and other information; according to the 80% principle, characteristic ion peaks with a frequency greater than 80% in non-QC samples were retained, and characteristic ion peaks with a frequency greater than half in QC samples and RSD>0.3 were removed; (2) Compound identification: MS FINDER software was used to decompose the .msp format file exported from MS-DIAL, and then based on the Human Metabolome Database, METLIN, the mass bank Web site, the database on the LIPID MAPS website was used for batch comparison, and the comparison and scoring of the secondary fragment ions with the standard library were performed according to the mass-to-charge ratio and relative abundance, and the compounds with the best matching degree were screened according to the comparison scores.

4. The method for constructing an ankylosing spondylitis ossification progression prediction model based on metabolomics and artificial intelligence according to claim 1, characterized in that: In step S3, the candidate features are: LysoPC (22:2 (13Z, 16Z)), quinic acid, ascorbic acid sulfate, LysoPC (18:1 (9Z)).

5. The method for constructing an ankylosing spondylitis ossification progression prediction model based on metabolomics and artificial intelligence according to claim 1, characterized in that: In step S4 and step S5, the differential metabolites are: 1) cysteine ​​and methionine metabolism; 2) pyrimidine metabolism; 3) biosynthesis of valine, leucine and isoleucine; and 4) arginine and proline metabolism.

6. The method for constructing an ankylosing spondylitis ossification progression prediction model based on metabolomics and artificial intelligence according to claim 1, characterized in that: Step S5 includes: Step S5.1: Divide the original data of the case materials of several patients into a training set and a validation set, and perform data preprocessing on each set; Step S5.2: Using the forward-backward method, the known risk factors, the candidate features, and the differential metabolites were used to establish LR models in R4.3.2, and the optimal feature combination was selected after comparing the performance of the models in the validation set; then, using the R package, randomly selecting samples in the training set and adjusting the model parameters by the grid optimization search algorithm, the LR model, the RF model, and the SVM model were established respectively; Step S5.3: Then, based on the validation set, a variety of model evaluation indicators are used to comprehensively determine the effectiveness of the three models through the average value of 20 cycles of validation, totaling 100 modeling times, and the models are compared. Finally, the best ankylosing spondylitis ossification progression prediction model is obtained, which is the LR model and / or the RF model.

7. The method for constructing an ankylosing spondylitis ossification progression prediction model based on metabolomics and artificial intelligence according to claim 6, characterized in that: In step S5.3, the multiple model evaluation indicators include AUC, accuracy, precision, recall, F1 score, sensitivity, specificity, and Matthews correlation coefficient; The Z test was used to compare the AUC values ​​and net reclassification improvement index (NRI) between models.

8. The method for constructing an ankylosing spondylitis ossification progression prediction model based on metabolomics and artificial intelligence according to any one of claims 1 to 7, characterized in that: The model parameters of the LR model are default parameters; the optimal parameters of the RF model are: mtry=3; ntree=800; the optimal parameters of the SVM model are: gamma=0.0063, C=355.6559, kernel function='radial'.

Citation Information

Cited By

  • Spine sagittal plane form intelligent classification method based on multi-scale feature extraction

    CN121482518A