Artificial intelligence model construction method for auxiliary diagnosis of acute myocardial infarction, storage medium and equipment
By combining Raman spectroscopy and metabolomics data, an artificial intelligence model was constructed using Stacking integrated learning method, which solved the problem of insufficient timeliness and specificity of early diagnosis of acute myocardial infarction, and improved diagnostic efficiency and quality.
Patent Information
- Application Number
- CN202510085031.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has insufficient timeliness and specificity in the early diagnosis of acute myocardial infarction, making it difficult to effectively assist clinical diagnosis.
By combining Raman spectroscopy and metabolomics data, an artificial intelligence model was constructed using Stacking integrated learning method to analyze biological plasma samples and predict the occurrence of acute myocardial infarction.
It improves the diagnostic efficiency and quality of acute myocardial infarction, enhances the ability to identify early lesions, and assists doctors in making more accurate clinical diagnosis.
Smart Images

Figure CN120148812A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medical technology, and particularly to an artificial intelligence model construction method, a storage medium, and a device for auxiliary diagnosis of acute myocardial infarction. Background Art
[0002] With the continuous development of artificial intelligence and machine learning technologies, their applications in the field of medical diagnosis have become increasingly widespread. Some common machine learning algorithms such as Support Vector Machine (SVM), Random Forest (RF), eXtreme Gradient Boosting (XGBoost), and Categorical Boosting (CatBoost) have been widely used. In addition to single machine learning algorithms, ensemble learning methods can reduce overfitting, enhance the generalization ability of the model, and improve the overall performance by integrating multiple machine learning algorithms. Common ensemble learning methods include strategies based on Bagging, Boosting, and Stacking (stacked generalization). The Stacking ensemble learning method first trains multiple base learners using different algorithms, and then uses their prediction results as the input of the meta-learner. The meta-learner obtains the final prediction result by synthesizing these predictions.
[0003] Early diagnosis of acute myocardial infarction is of great significance for reducing mortality and improving prognosis. Traditional diagnostic methods mainly rely on electrocardiogram and imaging examinations, but these methods have certain limitations in terms of timeliness and specificity of diagnosis.
[0004] Raman spectroscopy is a non-destructive detection technology based on molecular vibration scattering. This technology has the advantages of simple sample preparation, non-destructiveness, rapidity, high sensitivity, etc., and can provide rich molecular structure information. It is widely used in molecular characterization and component analysis in the fields of chemistry, materials, biomedicine, etc. Mass spectrometry metabolomics is an analytical technology for systematically studying all small molecule metabolites in an organism. It first separates the metabolites in a complex sample through liquid chromatography or gas chromatography, and then ionizes these separated compounds using a mass spectrometer and detects their mass-to-charge ratio and abundance. By comparing with a metabolite database, qualitative and quantitative analysis of metabolites can be achieved. Mass spectrometry metabolomics has the characteristics of high throughput, high sensitivity, and high resolution, and can simultaneously detect hundreds to thousands of metabolites. It is an important tool for studying the metabolic network of organisms, searching for biomarkers, and revealing the pathogenesis of diseases.
[0005] In recent years, metabolomics analysis has shown good application prospects in the field of disease-assisted diagnosis, because metabolomics data can reflect the dynamic changes of the body's metabolic state. At the same time, Raman spectroscopy can provide molecular structure information and chemical fingerprints of samples. Therefore, by jointly analyzing these two types of data, it is expected to provide new research ideas for the early diagnosis of acute myocardial infarction.
[0006] Therefore, it is a technical problem that needs to be solved urgently to develop an artificial intelligence model for auxiliary diagnosis of acute myocardial infarction based on metabolomics analysis and Raman spectroscopy technology combined with Stacking strategy to assist doctors in clinical diagnosis and improve the diagnostic efficiency and quality of acute myocardial infarction. Summary of the invention
[0007] The purpose of the present invention is to provide an artificial intelligence model construction method, storage medium and device for auxiliary diagnosis of acute myocardial infarction, which can make predictions on acute myocardial infarction based on biological plasma samples by processing their Raman spectral peak data and metabolite data, assist doctors in clinical diagnosis, and improve the diagnostic efficiency and quality of acute myocardial infarction.
[0008] In order to solve the above technical problems, the embodiments of the present invention provide a technical solution as follows:
[0009] A method for constructing an artificial intelligence model for auxiliary diagnosis of acute myocardial infarction comprises the following steps:
[0010] Step S1, converting the Raman spectrum peak diagram of the biological plasma sample into a numerical matrix, including the position of each characteristic peak and the corresponding spectral intensity value, normalizing the spectral intensity value, and screening out the difference characteristic peak data with significant difference as the first data through orthogonal partial least squares discriminant analysis combined with correlation analysis;
[0011] Step S2, performing orthogonal partial least squares discriminant analysis on the metabolite data of the biological plasma samples, and combining with a significance test, identifying and screening differential metabolite data, and using the screened differential metabolite data as the second data;
[0012] Step S3, constructing an original data set, where the original data set is any one of a first data set constructed based on the first data, a second data set constructed based on the second data, and a third data set constructed based on the first data and the second data;
[0013] Step S4: Select an XGBoost model, a CatBoost model, a SVM model, and a RF model as four base learner models, and perform hyperparameter optimization on each base learner model;
[0014] Step S5: Divide the original dataset into a training set and a test set. Apply the K-fold cross-validation strategy to the training set, that is, divide the training set into K parts. In K iterations, each time select 1 part as the sub-test set, and the remaining K - 1 parts as the sub-training set. Train and predict the four base learners respectively to obtain their probability prediction results on the sub-test set. Combine the probability prediction results of the four base learners on the sub-test set by columns to form a new feature training set. For the test set, use the K models generated by each base learner in K iterations to perform probability prediction respectively, and take the average value of their probability prediction results. Combine the average values of the probability prediction results of the four base learners by columns as the new feature test set.
[0015] Step S6: Use logistic regression as the meta-learner model, train it on the new feature training set, and make predictions on the new feature test set. Obtain the prediction results of the Stacking ensemble model as the output of the acute myocardial infarction prediction results.
[0016] Furthermore, it also includes comparative analysis of the prediction results under different original datasets, and uses accuracy, sensitivity, specificity, F1 score, and AUC value as evaluation indicators to evaluate the performance of each model.
[0017] Furthermore, in step S1, the normalization process of the spectral intensity values specifically includes using the maximum-minimum normalization method to adjust the numerical range of the spectral intensity values to the interval [-1, 1]. Its calculation formula is , where is the original spectral intensity value, is the normalized spectral intensity value, and respectively represent the minimum and maximum values of the spectral intensity values; the orthogonal partial least squares discriminant analysis is implemented using SIMCA software, and feature screening is performed by calculating the variable importance in the projection value. Select the significant characteristic peaks with the variable importance in the projection value greater than 1 as the standard.
[0018] Furthermore, in step S1, the correlation analysis includes: calculating the Pearson correlation coefficient between the significant characteristic peaks. In the formula are two different significant characteristic peak variables, is the covariance of the two significant characteristic peak variables, , are the standard deviations of the two significant characteristic peak variables respectively; for the pairs of significant characteristic peaks with a Pearson correlation coefficient greater than 0.8, it indicates a strong correlation. Remove one of the significant characteristic peaks with a larger correlation coefficient with other significant characteristic peaks.
[0019] Further, in the step S2, the orthogonal partial least squares discriminant analysis is implemented using SIMCA software, and the variable importance in the projection (VIP) value is calculated. The significant metabolite data is screened out based on the criterion that the VIP value is greater than 1. The significance test is an independent samples t-test for the significant metabolite data, and the significant metabolite data with is selected as the differential metabolite data with statistical significance, where the p-value represents the probability of observing the current or more extreme difference under the condition that the null hypothesis that there is no significant difference between the metabolite sample data of diseased organisms and the metabolite sample data of non-diseased organisms is true.
[0020] Further, in the step S4, hyperparameter optimization for each base learner model includes using random search to determine an approximate optimal parameter range, and then using grid search to determine the final optimal parameter combination within this approximate optimal parameter range.
[0021] Further, in the step S4, hyperparameter optimization for each base learner model includes: for the XGBoost model, optimizing the learning rate and the maximum number of trees; for the CatBoost model, optimizing the learning rate, the tree depth, and the maximum number of iterations; for the SVM model, optimizing the penalty coefficient, the kernel function type, and the kernel function coefficient; for the RF model, optimizing the maximum number of decision trees.
[0022] Further, in the step S5, the original dataset is divided into a training set and a test set, where the division ratio of the training set to the test set is 9:1.
[0023] To solve the technical problems proposed by the present invention, the present invention also provides a technical solution as follows:
[0024] A readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement any one of the above artificial intelligence model construction methods for the auxiliary diagnosis of acute myocardial infarction.
[0025] To solve the technical problems proposed by the present invention, the present invention also provides a technical solution as follows:
[0026] A computer program device, including a computer program. When the computer program is executed by a processor, it can implement any one of the above artificial intelligence model construction methods for the auxiliary diagnosis of acute myocardial infarction.
[0027] The artificial intelligence model construction method, storage medium, and device for the auxiliary diagnosis of acute myocardial infarction provided by the present invention, compared with the prior art, can, based on biological plasma samples, through the processing of their Raman spectral peak data and metabolite data, make a prediction of acute myocardial infarction based on the Stacking ensemble model, assist doctors in clinical diagnosis, and improve the diagnosis efficiency and quality of acute myocardial infarction. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the drawings in the figures do not constitute a scale limitation.
[0029] Figure 1 is a flowchart of the steps of the artificial intelligence model construction method for the auxiliary diagnosis of acute myocardial infarction in the embodiments of the present invention;
[0030] Figure 2 is a block diagram of the architecture of the artificial intelligence model construction method for the auxiliary diagnosis of acute myocardial infarction in the embodiments of the present invention;
[0031] Figure 3 is a schematic structural diagram of the Stacking ensemble model in the embodiments of the present invention;
[0032] Figure 4 is a bar chart of the classification performance indicators of each model in the embodiments of the present invention under different input data conditions. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] To make the objectives, technical solutions and advantages of the present invention clearer, the following will elaborate on the various embodiments of the present invention in conjunction with the accompanying drawings. However, those of ordinary skill in the art can understand that in the various embodiments of the present invention, many technical details are provided for the readers to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions required to be protected by the claims of the present application can still be achieved.
[0034] The terms "comprise", "include" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0035] In the following detailed description, many specific details are set forth to provide a more thorough understanding of the present invention. However, it will be apparent to those skilled in the art that well-known control algorithms are not shown in detail to avoid obscuring the gist of the present invention; and the orthogonal partial least squares discriminant analysis method, significance test method, XGBoost (Extreme Gradient Boosting) model, CatBoost (Classification Boosting) model, SVM (Support Vector Machine) model, RF (Random Forest) model, LR (Logistic Regression) model, Stacking (Stacked Generalization) ensemble strategy, etc. involved in the following effect embodiments are prior arts that can be retrieved.
[0036] As Figure 1-2 shown, an embodiment of the present invention relates to a method for constructing an artificial intelligence model for the auxiliary diagnosis of acute myocardial infarction, including the following steps:
[0037] Step S1: Convert the Raman spectral peak map of the biological plasma sample into a numerical matrix, including the positions of each characteristic peak and the corresponding spectral intensity values, and then perform normalization processing on the spectral intensity values. Through orthogonal partial least squares discriminant analysis combined with correlation analysis, screen out the differential characteristic peak data with significant differences as the first data.
[0038] Preferably, the biological plasma sample used is a mouse plasma sample. The maximum-minimum normalization method is used for normalizing the spectral intensity values, and the numerical range of the spectral intensity values is adjusted to the interval [-1, 1]. Its calculation formula is, , where is the original spectral intensity value, is the normalized spectral intensity value, and respectively represent the minimum and maximum values of the spectral intensity values; the orthogonal partial least squares discriminant analysis is implemented using SIMCA (Soft Independent Modeling of Class Analogies) software, and feature screening is performed by calculating the variable projection importance value. The significant characteristic peaks that make a significant contribution to the model prediction are screened out with the variable projection importance value greater than 1. Among them, the correlation analysis is performed by calculating the Pearson correlation coefficient between the significant characteristic peak features. In the formula, are two different significant characteristic peak variables, is the covariance of the two features, , They are the standard deviations of the two significant characteristic peak variables respectively; for the pairs of significant characteristic peaks with a Pearson correlation coefficient greater than 0.8, it indicates a strong correlation, and one of the significant characteristic peaks with a relatively large correlation coefficient with other significant characteristic peaks is removed; for the pairs of significant characteristic peaks with a Pearson correlation coefficient less than or equal to 0.8, it indicates no strong correlation, and the original features are retained without dimensionality reduction processing.
[0039] Step S2: Perform orthogonal partial least squares discriminant analysis on the metabolite data of the biological plasma sample, and combine significance testing to identify and screen out the differential metabolite data. These screened differential metabolite data are used as the second data and serve as the input data for subsequent model analysis.
[0040] Preferably, the biological plasma sample is a mouse plasma sample, the metabolite data is implemented by SIMCA software through orthogonal partial least squares discriminant analysis, and the variable importance in the projection value is calculated; the significant metabolite data that makes a significant contribution to the model prediction is screened out with the variable importance in the projection value greater than 1 as the standard; the significance test is an independent sample t-test on the significant metabolite data obtained by orthogonal partial least squares discriminant analysis, and the selected significant metabolite data is used as the differential metabolite data with statistical significance, where the p-value represents the probability of observing the current or more extreme difference under the condition that the null hypothesis that there is no significant difference between the diseased biological metabolite sample data and the non-diseased biological metabolite sample data is true. It means that the probability of this difference occurring is less than 5%, and it can be considered that this difference is not caused by random factors. These differential metabolite data will be used as the second data for the input data of subsequent model analysis.
[0041] Step S3: Construct an original data set, which is any one of the first data set constructed based on the first data, the second data set constructed based on the second data, and the third data set constructed based on the first data and the second data.
[0042] Step S4: Select the XGBoost model, CatBoost model, SVM model, and RF model as the four base learner models; perform hyperparameter optimization on each base learner model to preliminarily predict whether the organism providing the plasma sample has acute myocardial infarction. Preferably, a combination of random search and grid search is used to optimize the parameters of each base learner model, with the AUC value as the evaluation metric. First, use random search to determine the approximate optimal parameter range, and then use grid search within this approximate optimal parameter range to further determine the final optimal parameter combination; in a demonstrative example, the following hyperparameter optimizations are performed on each base learner model: for the XGBoost model, optimize the learning rate learning_rate and the maximum number of trees n_estimators; for the CatBoost model, optimize the learning rate learning_rate, the tree depth max_depth, and the maximum number of iterations iterations; for the SVM model, optimize the penalty coefficient C, the kernel function type kernel, and the kernel function coefficient gamma; for the RF model, optimize the maximum number of decision trees n_estimators. Table 1 shows the optimal parameters of the four base learner models, namely the XGBoost model, CatBoost model, SVM model, and RF model.
[0043] Table 1 Optimal Parameters of the Base Learner Models
[0044]
[0045] Step S5: Divide the original dataset into a training set and a test set. Adopt the K-fold cross-validation strategy for the training set, that is, divide the training set into K parts. In K iterations, each time select 1 part as the sub-test set, and the remaining K - 1 parts as the sub-training set. Train and predict the four base learners respectively to obtain their probability prediction results on the sub-test set. Combine the probability prediction results of the four base learners on the sub-test set column by column to form a new feature training set; for the test set, use the K models generated by each base learner in K iterations to perform probability prediction respectively, take the average value of their probability prediction results, and combine the average values of the probability prediction results of the four base learners column by column as the new feature test set. Preferably, divide the original dataset into a training set and a test set, where the division ratio of the training set to the test set is 9:1.
[0046] As an example, refer to Figure 3, a 5-fold cross-validation strategy is adopted for the training set: the training set is divided into 5 parts. During 5 iterations, each time 1 part of the data is selected as the sub-test set, and the remaining 4 parts of the data are used as the sub-training set, and the loop iteration is carried out until each part of the data is used as the sub-test set once. In this way, four base learner models are trained to obtain their probability prediction results, and these probability prediction results are merged column by column to form a new feature training set; for the test set, probability predictions are respectively made using the models generated in 5 iterations, and the average value of their probability prediction results is taken. The average values of the probability prediction results of the four base learners are merged column by column as the new feature test set.
[0047] Step S6: Use logistic regression as the meta-learner model and train it on the new feature training set composed of the probability prediction results of the base learner models; finally, make predictions on the new feature test set, and the final prediction result of the Stacking ensemble model is used as the acute myocardial infarction prediction result for output. Preferably, under the three original data sets, the parameters of the logistic regression model are all set as regularization strength C = 0.1 and maximum number of iterations max_iter = 200.
[0048] In one embodiment, in order to be able to evaluate the performance of each model, that is, the XGBoost model, the CatBoost model, the SVM model, the RF model and the Stacking ensemble model, and compare and analyze the prediction results of each model under different original data sets, accuracy, sensitivity, specificity, F1 score and AUC value can be used as evaluation indicators to evaluate the performance of each model. The AUC (Area Under Curve) value is defined as the area enclosed by the ROC curve and the coordinate axis.
[0049] Table 2-4 respectively shows the classification performance comparison between each base learner model and the Stacking ensemble model under the condition of inputting three different original data sets.
[0050] Table 2 Classification performance of different models when the input is the first data set
[0051]
[0052] Table 3 Classification performance of different models when the input is the second data set
[0053]
[0054] Table 4 Classification performance of different models when the input is the third data set
[0055]
[0056] As shown in Table 2-4, under different input data conditions, the Stacking ensemble model demonstrates excellent prediction performance: when the input data is the first dataset obtained based on Raman spectroscopy, the AUC values of XGBoost, CatBoost, SVM, and RF are 0.7722, 0.6722, 0.7711, and 0.7611 respectively, while the AUC value of the Stacking ensemble model reaches 0.8111, which is at least 5.04% higher than that of individual base learners; when the input data is the second dataset obtained based on metabolomics, the Stacking ensemble model shows excellent prediction performance in multiple evaluation metrics such as accuracy, specificity, F1-score, and AUC value compared to the four base learners; when the input data is the third dataset based on the dual-modal Raman spectroscopy and metabolomics, the Stacking ensemble model has the highest specificity and AUC value. This proves that the Stacking ensemble model can effectively integrate the prediction results of multiple base learners to obtain better classification performance.
[0057] Figure 4 The bar chart of the classification performance metrics of each model under different input data conditions is shown. By Figure 4 It can more intuitively display the performance metrics of each model when different original datasets are input. The classification performance defects are obvious when using the first dataset or the second dataset alone. When using the third dataset combined with the first data and the second data, the classification performance of the model can be significantly improved, and more reliable prediction results for acute myocardial infarction can be output to assist doctors in clinical diagnosis and provide more reliable technical support for the auxiliary diagnosis of acute myocardial infarction.
[0058] In an embodiment of the present invention, a readable storage medium is further provided, on which a computer program is stored. It is characterized in that when the computer program is executed by a processor, it can implement the above-mentioned artificial intelligence model construction method for the auxiliary diagnosis of acute myocardial infarction.
[0059] In an embodiment of the present invention, a computer program device is further provided, including a computer program, and when the computer program is executed by a processor, it can implement the above-mentioned artificial intelligence model construction method for the auxiliary diagnosis of acute myocardial infarction.
[0060] The artificial intelligence model construction method, storage medium, and device provided by the present invention for the auxiliary diagnosis of acute myocardial infarction can, based on biological plasma samples, through the processing of their Raman spectroscopy peak map data and metabolite data, make predictions for acute myocardial infarction based on the Stacking ensemble model, assist doctors in clinical diagnosis, and improve the diagnosis efficiency and quality of acute myocardial infarction.
[0061] The preferred embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and the devices and structures not described in detail therein should be understood to be implemented in a common manner in the art; any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention, or modify it into an equivalent embodiment with equivalent changes, without departing from the scope of the technical solution of the present invention, which does not affect the essence of the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still be within the scope of protection of the technical solution of the present invention.
Claims
1. A method for constructing an artificial intelligence model for auxiliary diagnosis of acute myocardial infarction, characterized in that: The following steps are involved: Step S1, converting the Raman spectrum peak diagram of the biological plasma sample into a numerical matrix, including the position of each characteristic peak and the corresponding spectral intensity value, normalizing the spectral intensity value, and screening out the difference characteristic peak data with significant difference as the first data through orthogonal partial least squares discriminant analysis combined with correlation analysis; Step S2, performing orthogonal partial least squares discriminant analysis on the metabolite data of the biological plasma samples, and combining with a significance test, identifying and screening differential metabolite data, and using the screened differential metabolite data as the second data; Step S3, constructing an original data set, where the original data set is any one of a first data set constructed based on the first data, a second data set constructed based on the second data, and a third data set constructed based on the first data and the second data; Step S4: Select an XGBoost model, a CatBoost model, a SVM model, and a RF model as four base learner models, and perform hyperparameter optimization on each base learner model; Step S5, dividing the original data set into a training set and a test set, adopting a K-fold cross-validation strategy for the training set, that is, dividing the training set into K parts, selecting 1 part as a sub-test set each time in K iterations, and the remaining K-1 parts as sub-training sets, respectively training and predicting the four base learners, obtaining their probability prediction results on the sub-test sets, and merging the probability prediction results of the four base learners on the sub-test sets by columns to form a new feature training set; For the test set, use the K models generated by each base learner in K iterations to make probability predictions respectively, take the average of their probability prediction results, and merge the average probability prediction results of the four base learners by column as the new feature test set; Step S6: Use logistic regression as the meta-learner model, perform training on the new feature training set, perform prediction on the new feature test set, and obtain the prediction result of the Stacking ensemble model as the output of the acute myocardial infarction prediction result.
2. The method for constructing an artificial intelligence model for auxiliary diagnosis of acute myocardial infarction according to claim 1, characterized in that: It also includes comparative analysis of the prediction results under different original data sets, using accuracy, sensitivity, specificity, F1 score and AUC value as evaluation indicators to evaluate the performance of each model.
3. The method for constructing an artificial intelligence model for auxiliary diagnosis of acute myocardial infarction according to claim 1, characterized in that: In step S1, the normalization process of the spectral intensity value specifically includes: using the maximum and minimum value normalization method to adjust the numerical range of the spectral intensity value to the interval [-1,1], and the calculation formula is: ,in, is the original spectral intensity value, is the normalized spectral intensity value, and Respectively represent the minimum and maximum values of the spectral intensity value; the orthogonal partial least squares discriminant analysis is implemented using SIMCA software, and feature screening is performed by calculating the variable projection importance value, and the significant feature peaks are screened out using the variable projection importance value greater than 1 as the standard.
4. The method for constructing an artificial intelligence model for auxiliary diagnosis of acute myocardial infarction according to claim 3, characterized in that: In step S1, the correlation analysis includes: calculating the Pearson correlation coefficient between significant feature peaks , where are two different significant characteristic peak variables, is the covariance of the two significant characteristic peak variables, , are the standard deviations of the variables of the two significant feature peaks respectively; for the significant feature peak pairs with a Pearson correlation coefficient greater than 0.8, it indicates that there is a strong correlation, and one of the significant feature peaks with a large correlation coefficient with the other significant feature peaks is removed.
5. The method for constructing an artificial intelligence model for auxiliary diagnosis of acute myocardial infarction according to claim 1, characterized in that: In step S2, the orthogonal partial least squares discriminant analysis is implemented using SIMCA software, and the variable projection importance value is calculated. The significant metabolite data are screened out with the variable projection importance value greater than 1 as the standard. The significance test is to perform an independent sample t test on the significant metabolite data, and select The significant metabolite data are taken as statistically significant differential metabolite data, where the p-value represents the probability of observing the current or more extreme difference under the condition that the null hypothesis that there is no significant difference between the metabolite sample data of the diseased organism and the metabolite sample data of the non-diseased organism is true.
6. The method for constructing an artificial intelligence model for auxiliary diagnosis of acute myocardial infarction according to claim 1, characterized in that: In step S4, the hyperparameter optimization of each base learner model includes using random search to determine an approximate optimal parameter range, and then using grid search to determine a final optimal parameter combination within the approximate optimal parameter range.
7. The method for constructing an artificial intelligence model for auxiliary diagnosis of acute myocardial infarction according to claim 6, characterized in that: In step S4, the hyperparameter optimization of each base learner model includes: for the XGBoost model, optimizing the learning rate and the maximum number of trees; for the CatBoost model, optimizing the learning rate, the depth of the tree and the maximum number of iterations; for the SVM model, optimizing the penalty coefficient, the kernel function type and the kernel function coefficient; for the RF model, optimizing the maximum number of decision trees.
8. The method for constructing an artificial intelligence model for auxiliary diagnosis of acute myocardial infarction according to claim 1, characterized in that: In step S5, the original data set is divided into a training set and a test set, wherein the ratio of the training set to the test set is 9:
1.
9. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it can implement the artificial intelligence model construction method for auxiliary diagnosis of acute myocardial infarction as described in any one of claims 1-8.
10. A computer program device comprising a computer program, characterized in that When the computer program is processed and executed, it can implement the artificial intelligence model construction method for auxiliary diagnosis of acute myocardial infarction as described in any one of claims 1-8.