A multi-scale feature enhancement and feature selection metabolite prediction method

By using multi-scale feature enhancement and feature selection methods, combined with information gain and machine learning models, the noise and redundancy problems in mass spectrometry data were solved, and the accuracy and efficiency of metabolite prediction were improved.

CN119380825BActive Publication Date: 2025-10-14HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411647402.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-10-14
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

The high-dimensional features of mass spectrometry data in existing metabolite identification methods carry a large amount of noise and redundant information, resulting in limited ability to process nonlinear relationships, leading to overfitting or information loss, making it difficult to capture metabolite interaction patterns and reducing the accuracy of metabolite prediction.

Method used

A multi-scale feature enhancement and feature selection method is used to extract mass spectrometry data features by sliding different window sizes. Information gain and machine learning models are combined to screen features and generate the final feature set for metabolite prediction.

Benefits of technology

It significantly reduces redundant mass spectrometry data features, improves computational efficiency and model generalization capabilities, and enhances the prediction accuracy of metabolites with similar structures or similar isotope distributions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380825B_ABST
    Figure CN119380825B_ABST
Patent Text Reader

Abstract

The application relates to a metabolite prediction method based on multi-scale feature enhancement and feature selection, and relates to the field of metabolite prediction.The application is used to solve the problem of low prediction accuracy of metabolites.The application comprises the following steps: pre-processing a mass spectrum data feature vector of each biological sample to obtain a multi-scale feature matrix; using the multi-scale feature matrix to obtain information gain of each feature for metabolite classification; screening features in the multi-scale feature matrix according to the information gain of each feature for metabolite prediction to obtain a final feature set; training an ANN model by using the final feature set and corresponding metabolite labels to obtain a metabolite prediction model; obtaining a final feature set of a biological sample to be predicted; and inputting the final feature set of the biological sample to be predicted into the metabolite prediction model to obtain a metabolite class of the biological sample to be predicted.The application is used for predicting a metabolite class of a biological sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of metabolite identification, and in particular to a metabolite prediction method based on multi-scale feature enhancement and feature selection. Background Art

[0002] Untargeted metabolomics is a technique used to comprehensively analyze all metabolites in biological samples, providing in-depth metabolic information for complex biological systems. In mass spectrometry data, metabolites are typically represented by high-dimensional features, including mass, charge ratio, retention time, and signal intensity. Because mass spectrometry data often contain significant noise, redundant information, and complex inter-feature correlations, extracting valid metabolite-related features from them has become a key challenge in metabolomics data analysis.

[0003] The high-dimensional features of mass spectrometry data obtained in current metabolite identification methods often carry a large amount of noise and redundant information, and existing methods have limited ability to process nonlinear relationships between features. Therefore, overfitting or loss of useful information is prone to occur when processing these data, making it difficult to capture complex metabolite interaction patterns. At the same time, due to the overlap of different metabolites in mass, charge ratio and retention time, the accuracy of distinguishing metabolites with similar structures or similar isotope distributions is reduced, resulting in the current low accuracy of metabolite prediction. Summary of the Invention

[0004] The purpose of the present invention is to solve the problem of low accuracy of existing metabolite prediction and propose a metabolite prediction method with multi-scale feature enhancement and feature selection.

[0005] A metabolite prediction method based on multi-scale feature enhancement and feature selection, specifically:

[0006] Obtaining a final feature set of the biological sample to be predicted, inputting the final feature set of the biological sample to be predicted into a metabolite prediction model to obtain a metabolite category of the biological sample to be predicted;

[0007] The final feature set of the biological sample to be predicted is obtained according to the final features included in the final feature set of the biological sample;

[0008] The final feature set of the biological sample is obtained by:

[0009] Step 1: preprocess the mass spectrometry data feature vector of each biological sample to obtain a multi-scale feature matrix;

[0010] Step 2: Use the multi-scale feature matrix to obtain the information gain of each feature for the metabolite category;

[0011] Step 3: Screen the features in the multi-scale feature matrix according to the information gain of each feature for the metabolite category to obtain the final feature set.

[0012] Furthermore, in step 1, the mass spectrometry data feature vector of each biological sample is preprocessed to obtain a multi-scale feature matrix, specifically:

[0013] Step 11: Obtain biological sample S l The mass spectrum data feature vector D l , for vector D l The mid-peak intensity characteristic value is pre-processed;

[0014] First, obtain a biological sample S l The mass spectrum data feature vector D l ;

[0015] The mass spectrometry data feature vector D of the biological sample l The original features included are: peak area, retention time, and peak intensity corresponding to different mass-to-charge ratios;

[0016] Where l is the label of the biological sample, l∈[1,L], and L is the total number of biological samples;

[0017] Then, set the window size A as the small window, the window size B as the large window, and the sliding step size to a;

[0018] Wherein, A and B are positive integers, B is greater than A, A is a positive integer greater than or equal to 3, and a is a positive integer greater than or equal to 1;

[0019] Then, use the small window to step a in vector D l Slide through the sequence of peak intensity characteristic values ​​corresponding to different mass-to-charge ratios in the image to obtain the peak intensity characteristic value in each small window;

[0020] Use a large window to follow a step size in vector D l Slide through the sequence of peak intensity characteristic values ​​corresponding to different mass-to-charge ratios in the image to obtain the peak intensity characteristic value in each large window;

[0021] Finally, the peak intensity eigenvalues ​​in each small window and the peak intensity eigenvalues ​​in each large window are taken as the preprocessed vector D l mid-peak intensity characteristic value;

[0022] Step 1 and 2: Obtain the biological sample S using the pre-processed peak intensity characteristic value l The local eigenvectors and global eigenvectors of ;

[0023] Step 13: Using the biological sample S obtained in step 12 lThe local eigenvector, global eigenvector and biological sample S l The mass spectrum data feature vector D l , forming the biological sample S l The multi-scale feature vector of

[0024] In step 14, the multi-scale feature vectors of different biological samples are padded with zeros to form a multi-scale feature matrix with the same length.

[0025] Furthermore, the peak intensity characteristic value after preprocessing is used to obtain the biological sample S in steps one and two. l The local eigenvectors and global eigenvectors of are:

[0026] First, the mean, maximum, minimum, standard deviation, variance, and median of the peak intensity characteristic values ​​in each window are obtained;

[0027] Then, the mean, maximum, minimum, standard deviation, variance, and median of the peak intensity characteristic values ​​in each small window are combined to form the biological sample S l Finally, the mean, maximum, minimum, standard deviation, variance, and median of the peak intensity feature values ​​in each large window are combined to form the biological sample S l The global eigenvector of .

[0028] Furthermore, in step 2, the multi-scale feature matrix is ​​used to obtain the information gain of each feature for the metabolite category, specifically:

[0029] Step 21: Obtain the overall entropy of the biological sample metabolite category labels;

[0030] Step 22: using the overall entropy of the biological sample metabolite category label to obtain the conditional entropy of the biological sample metabolite category label;

[0031] Step 2 and 3: Use the overall entropy of the biological sample metabolite category label and the conditional entropy of the biological sample metabolite category label to obtain the feature X i' The information gain for metabolite category prediction.

[0032] Furthermore, the overall entropy of the biological sample metabolite category labels obtained in step 21 is specifically:

[0033]

[0034] Among them, i is the category number of the biological sample metabolite, n is the total number of categories of biological sample metabolites, and p i is the probability that the metabolite category label Y of the current biological sample is category i, and H(Y) is the overall entropy of the metabolite category label Y of the biological sample.

[0035] Furthermore, in step 22, the overall entropy of the biological sample metabolite category label is used to obtain the conditional entropy of the biological sample metabolite category label, specifically:

[0036]

[0037] Among them, X i' is the i-th feature in the multi-scale feature matrix, j is the feature X i' The label of the corresponding value, M i',j is feature X i' The corresponding j-th value, k is the feature X i The total number of corresponding values, p(X i' =M i',j ) is X i' Take M i',j The probability of H(Y|X i' =M i',j ) is X i' Take M i',j The total entropy of the label Y under the condition of .

[0038] Furthermore, in steps 2 and 3, the overall entropy of the biological sample metabolite category label and the conditional entropy of the biological sample metabolite category label are used to obtain the feature X i' The information gain for metabolite category prediction is:

[0039] IG(X i' ,Y)=H(Y)-H(Y|X i' )

[0040] Among them, IG(X i' ,Y) is feature X i' Information gain for metabolite classes.

[0041] Furthermore, in step 3, the features in the multi-scale feature matrix are screened according to the information gain of each feature for the metabolite category to obtain the final feature set, specifically:

[0042] Step 3. Preset a feature information gain threshold based on experience, compare the feature information gain threshold with the information gain of each feature for the metabolite category, retain features greater than the feature information gain threshold in the multi-scale feature matrix, and delete features less than or equal to the feature information gain threshold to obtain a multi-scale feature matrix after preliminary screening;

[0043] Step 32: Using each multi-scale feature vector and the corresponding metabolite category label in the preliminarily screened multi-scale feature matrix to form an overall data set, the overall data set is divided into a training set and a test set, and multiple classification models are trained and tested using the training set and the test set, thereby obtaining a feature weight distribution model;

[0044] Various classification models include: ANN model, support vector machine, random forest model;

[0045] Step 3. Use the training set to train the feature weight allocation model multiple times. The feature weight allocation model assigns weights to each feature, and deletes the feature with the lowest weight in the training set each time. At the end of each training, the test set is used to test the performance of the feature weight allocation model. The accuracy of the feature weight allocation model first increases and then decreases during the process of gradual feature deletion. When the accuracy reaches the highest point, stop deleting features and form the final feature set with the remaining features.

[0046] Furthermore, in step 32, the multiple classification models are trained and tested using the training set and the test set respectively, so as to obtain a feature weight allocation model, specifically: the trained model with the highest accuracy is used as the feature weight allocation model.

[0047] Furthermore, the metabolite prediction model is obtained by:

[0048] The final feature set and metabolite category labels are used to form a final training set, and the classification model with the highest accuracy obtained in step 32 is trained using the final training set to obtain a metabolite prediction model.

[0049] The beneficial effects of the present invention are:

[0050] The present invention proposes a feature selection method of multi-scale feature enhancement and information gain (IG)-recursive feature elimination (RFE), and combines it with a machine learning model, which is specifically used for metabolite prediction in non-targeted metabolomics. The present invention performs sliding extraction on the features in the mass spectrometry data feature vector of the biological sample through different window sizes. The small-scale window is used to capture rapidly changing features, and the large-scale window is used to capture the global pattern across multiple feature points, thereby generating features of different scales. Then, the mass spectrometry data features are screened based on the information gain of the metabolite category according to different features. Finally, the mass spectrometry data features are weighted again in combination with the machine learning model to obtain the final feature set, and the final feature set and the machine learning model are combined to predict metabolites. The present invention can significantly reduce redundant mass spectrometry data features, improve computational efficiency, and enhance the generalization ability of the model, and is particularly suitable for processing large-scale mass spectrometry data. The present invention can capture complex metabolite interaction patterns, can improve the accuracy of distinguishing metabolites with similar structures or similar isotope distributions, thereby improving the accuracy of metabolite prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 Flowchart of the present invention;

[0052] Figure 2 It is a schematic diagram of small window sliding;

[0053] Figure 3 This is a schematic diagram of large window sliding. DETAILED DESCRIPTION

[0054] Specific implementation method 1: Figure 1 As shown, the specific process of the metabolite prediction method of multi-scale feature enhancement and feature selection in this embodiment is as follows:

[0055] Obtaining a final feature set of the biological sample to be predicted, inputting the final feature set of the biological sample to be predicted into a metabolite prediction model to obtain a metabolite category of the biological sample to be predicted;

[0056] The final feature set of the biological sample to be predicted is obtained according to the final features included in the final feature set of the biological sample;

[0057] The final characteristics of the biological sample are obtained by:

[0058] Step 1: Preprocess the mass spectrometry data feature vector of each biological sample to obtain a multi-scale feature matrix, specifically:

[0059] Step 11: Obtain biological sample S l The mass spectrum data feature vector D l , for vector D l Preprocess the peak intensity characteristic value:

[0060] First, set the small window size to A and the large window size to B.

[0061] Wherein, A and B are positive integers, B is greater than A, and A is a positive integer greater than or equal to 3; in this step, A is set to 3 and B is set to 9;

[0062] Then, a small window and a large window are used to calculate the vector D. l Slide the sequence of peak intensities corresponding to different mass-to-charge ratios in a step size to obtain the peak intensity characteristic value in each window, such as Figure 2 、 3 As shown;

[0063] The mass spectrometry data feature vector D of the biological sample l The original features in include: peak area, retention time, and peak intensity corresponding to different mass-to-charge ratios;

[0064] Figure 2-3 In the sequence of peak intensities corresponding to different mass-to-charge ratios, the first value is the peak intensity value of 200 corresponding to a mass-to-charge ratio of 100, the second value is the peak intensity value of 500 corresponding to a mass-to-charge ratio of 120, the third value is the peak intensity value of 650 corresponding to a mass-to-charge ratio of 180, and so on until the peak intensity value corresponding to the last mass-to-charge ratio.

[0065] One sliding step of the window represents one unit, which is recorded as a peak value;

[0066] In this step, set a=1, or set it to an integer greater than 1 to control the coverage of the window;

[0067] Where l is the label of the biological sample, l∈[1,L], and L is the total number of biological samples;

[0068] Step 1 and 2: Obtain the biological sample S using the pre-processed peak intensity characteristic value l The local eigenvectors and global eigenvectors of are:

[0069] First, the mean, maximum, minimum, standard deviation, variance, and median of the peak intensity characteristic values ​​in each window are obtained;

[0070] Then, the mean, maximum, minimum, standard deviation, variance, and median of the peak intensity characteristic values ​​in each small window in the mass-to-charge ratio direction of the mass spectrum data characteristic vector are combined to form the biological sample S l The local eigenvector of ;

[0071] The order of data storage in the local feature vector is: first store the average value, maximum value, minimum value, standard deviation, variance, and median of the peak intensity feature value in the first small window, and then store the average value, maximum value, minimum value, standard deviation, variance, and median of the peak intensity feature value in the second small window, until the average value, maximum value, minimum value, standard deviation, variance, and median of the peak intensity feature value in all small windows are stored;

[0072] Then, the mean, maximum, minimum, standard deviation, variance, and median of the peak intensity characteristic values ​​in each large window in the mass-to-charge ratio direction of the mass spectrum data characteristic vector are combined to form the biological sample S l The global eigenvector of ;

[0073] The order of data storage in the global eigenvector is: first store the average value, maximum value, minimum value, standard deviation, variance, and median of the peak intensity eigenvalues ​​in the first large window, and then store the average value, maximum value, minimum value, standard deviation, variance, and median of the peak intensity eigenvalues ​​in the second large window, until the average value, maximum value, minimum value, standard deviation, variance, and median of the peak intensity eigenvalues ​​in all large windows are stored.

[0074] Step 13: Using the biological sample S obtained in step 12 l The local eigenvector, global eigenvector and biological sample S l The mass spectrum data feature vector D l , forming the biological sample S l The multi-scale feature vector of :

[0075] The biological sample S l The original features, local features, and global features constitute multi-scale features;

[0076] Step 14: To ensure that the multi-scale feature vectors of each sample have the same length, all multi-scale feature vectors are padded with zeros to a uniform length, and the multi-scale feature vectors of all samples after the zero padding are combined into a multi-scale feature matrix.

[0077] In this step, features of different scales are generated by different window sizes (small windows are used to capture local changes, and large windows are used to capture global trends). Small-scale windows are used to capture rapidly changing features, and large-scale windows are used to capture global patterns across multiple feature points. The original data is smoothed using methods such as sliding windows and moving averages to generate feature vectors on multiple scales. The average value of all eigenvalues ​​in the window is calculated to smooth local changes. The maximum / minimum value is calculated to capture peak features in the window, which is suitable for predicting mass spectrometry peaks. The variance / standard deviation is calculated, and the variance or standard deviation in the window is used to measure the degree of local change and capture data volatility. The median in the window is calculated, which is suitable for data with outliers. The local feature vectors obtained by the present invention reflect the changes in the biological sample within a local range, and all feature vectors can reflect the overall pattern or change trend of the organism.

[0078] Step 2: Use the multi-scale feature matrix to obtain the information gain of each feature for the metabolite category, specifically:

[0079] Step 2.1: Obtain the overall entropy of the biological sample metabolite category labels:

[0080]

[0081] Among them, i is the category number of the biological sample metabolite, n is the total number of categories of biological sample metabolites, and p i is the probability that the metabolite category label Y of the current biological sample is category i, and H(Y) is the overall entropy of the metabolite category label Y of the biological sample;

[0082] Step 2. Use the overall entropy of the biological sample metabolite category label to obtain the conditional entropy of the biological sample metabolite category label:

[0083]

[0084] Among them, X i' is the i-th feature in the multi-scale feature matrix, j is the feature X i' The label of the corresponding value, M i',j is feature X i' The corresponding j-th value, k is the feature X i The total number of corresponding values, p(X i' =M i',j ) is X i' Take M i',j The probability of H(Y|X i' =M i',j ) is X i' Take M i',j The total entropy of the label Y under the condition of ;

[0085] Step 2 and 3: Use the overall entropy of the biological sample metabolite category label and the conditional entropy of the biological sample metabolite category label to obtain the feature X i' The information gain for metabolite category prediction is:

[0086] IG(X i' ,Y)=H(Y)-H(Y|X i' )

[0087] Among them, IG(X i' ,Y) is feature X i' Information gain for metabolite classes.

[0088] The higher the information gain in this step, the greater the contribution of that feature to metabolite classification. Finally, we screen for key features. We rank the mass spectrometric features based on the calculated information gain value for each feature, selecting those with the highest information gain as key features. These features are more closely associated with the metabolite class labels, helping to improve the accuracy of the classification model.

[0089] Step 3: Screen the features in the multi-scale feature matrix based on the information gain of each feature for metabolite category prediction to obtain the final feature set:

[0090] Step 3. Preset a feature information gain threshold based on experience, compare the feature information gain threshold with the information gain of each feature for the metabolite category, retain features greater than the feature information gain threshold in the multi-scale feature matrix, and delete features less than or equal to the feature information gain threshold to obtain a multi-scale feature matrix after preliminary screening;

[0091] Step 32: Use each multi-scale feature vector and the corresponding metabolite category label in the preliminarily screened multi-scale feature matrix to form an overall data set, divide the overall data set into a training set and a test set, use the training set to train an artificial neural network (ANN), a support vector machine (SVM), and a random forest model, respectively, and use the test set to test the trained artificial neural network (ANN), support vector machine (SVM), and random forest models, respectively, and use the trained model with the highest accuracy and recall rate as the feature weight allocation model;

[0092] Among them, the 10-fold cross-validation method is used to optimize the model parameters during the training process;

[0093] Step 3. Use the training set to train the feature weight allocation model again. The feature weight allocation model assigns weights to each feature, and deletes the feature with the lowest weight in the training set in each round of training. At the end of each training, the test set is used to test the performance of the feature weight allocation model. The accuracy of the feature weight allocation model first increases and then decreases during the process of gradual feature deletion. When the accuracy reaches the highest point, stop deleting features and form the final feature set with the remaining features.

[0094] In this step, a feature selection method for gradually eliminating unimportant features is proposed, which can further optimize the feature set by combining the information gain screening of mass spectrometry data. The information gain of the present invention ensures that features that are strongly correlated with the target variable are screened out; secondly, recursive feature elimination further determines which features contribute most to model performance through actual model evaluation. According to the results of recursive feature elimination, features with higher importance are selected, and an optimized feature subset is finally formed. These features not only have a strong correlation with the target variable, but also contribute significantly to model performance. Peak intensity values ​​corresponding to certain mass-to-charge ratios may remain in the final feature set. When the final feature set of the sample to be predicted is obtained, the peak intensity of the sample to be predicted at the mass-to-charge ratio is obtained according to the mass-to-charge ratio corresponding to the peak intensity value in the final feature set as the feature of the sample to be predicted.

[0095] Specific embodiment 2: The metabolite prediction model is obtained by the following method:

[0096] The final feature set and metabolite category labels are used to train the classification model with the highest accuracy obtained in step 32 to obtain a metabolite prediction model.

Claims

1. A metabolite prediction method based on multi-scale feature enhancement and feature selection, characterized by: The specific process of the method is: Obtaining a final feature set of the biological sample to be predicted, inputting the final feature set of the biological sample to be predicted into a metabolite prediction model to obtain a metabolite category of the biological sample to be predicted; The final feature set of the biological sample to be predicted is obtained according to the final features included in the final feature set of the biological sample; The final feature set of the biological sample is obtained by: Step 1: Preprocess the mass spectrometry data feature vector of each biological sample to obtain a multi-scale feature matrix, specifically: Step 11: Obtain biological sample S l The mass spectrum data feature vector D l , for vector D l The mid-peak intensity characteristic value is pre-processed; First, obtain a biological sample S l The mass spectrum data feature vector D l ; The mass spectrometry data feature vector D of the biological sample l The original features included are: peak area, retention time, and peak intensity corresponding to different mass-to-charge ratios; Where l is the label of the biological sample, l∈[1,L], and L is the total number of biological samples; Then, set the window size A as the small window, the window size B as the large window, and the sliding step size to a; Wherein, A and B are positive integers, B is greater than A, A is a positive integer greater than or equal to 3, and a is a positive integer greater than or equal to 1; Then, use the small window to step a in vector D l Slide through the sequence of peak intensity characteristic values ​​corresponding to different mass-to-charge ratios in the image to obtain the peak intensity characteristic value in each small window; Use a large window to follow a step size in vector D l Slide through the sequence of peak intensity characteristic values ​​corresponding to different mass-to-charge ratios in the image to obtain the peak intensity characteristic value in each large window; Finally, the peak intensity eigenvalues ​​in each small window and the peak intensity eigenvalues ​​in each large window are taken as the preprocessed vector D l mid-peak intensity characteristic value; Step 1 and 2: Obtain the biological sample S using the pre-processed peak intensity characteristic value l The local eigenvectors and global eigenvectors of ; Step 13: Using the biological sample S obtained in step 12 l The local eigenvector, global eigenvector and biological sample S l The mass spectrum data feature vector D l , forming the biological sample S l The multi-scale feature vector of Step 14: padded the multi-scale feature vectors of different biological samples with zeros to make them of the same length to form a multi-scale feature matrix; Step 2: Use the multi-scale feature matrix to obtain the information gain of each feature for the metabolite category; Step 3: Screen the features in the multi-scale feature matrix according to the information gain of each feature for the metabolite category to obtain the final feature set.

2. The metabolite prediction method of claim 1, wherein: In the steps 1 and 2, the peak intensity characteristic value after preprocessing is used to obtain the biological sample S l The local eigenvectors and global eigenvectors of are: First, the mean, maximum, minimum, standard deviation, variance, and median of the peak intensity characteristic values ​​in each window are obtained; Then, the mean, maximum, minimum, standard deviation, variance, and median of the peak intensity characteristic values ​​in each small window are combined to form the biological sample S l The local eigenvector of ; Finally, the mean, maximum, minimum, standard deviation, variance, and median of the peak intensity characteristic values ​​in each large window are combined to form the biological sample S l The global eigenvector of .

3. The metabolite prediction method of multi-scale feature enhancement and feature selection according to claim 2, characterized in that: In step 2, the multi-scale feature matrix is ​​used to obtain the information gain of each feature for the metabolite category, specifically: Step 21: Obtain the overall entropy of the biological sample metabolite category labels; Step 22: using the overall entropy of the biological sample metabolite category label to obtain the conditional entropy of the biological sample metabolite category label; Step 2 and 3: Use the overall entropy of the biological sample metabolite category label and the conditional entropy of the biological sample metabolite category label to obtain the feature X i' The information gain for metabolite category prediction.

4. The metabolite prediction method of claim 3, wherein: The overall entropy of obtaining the biological sample metabolite category label in step 21 is specifically: Among them, i is the category number of the biological sample metabolite, n is the total number of categories of biological sample metabolites, and p i is the probability that the metabolite category label Y of the current biological sample is category i, and H(Y) is the overall entropy of the metabolite category label Y of the biological sample.

5. The metabolite prediction method of multi-scale feature enhancement and feature selection according to claim 4, characterized in that: The step 22 of obtaining the conditional entropy of the biological sample metabolite category label using the overall entropy of the biological sample metabolite category label is specifically as follows: Among them, X i' is the i-th feature in the multi-scale feature matrix, j is the feature X i' The label of the corresponding value, M i',j is feature X i' The corresponding j-th value, k is the feature X i The total number of corresponding values, p(X i' =M i',j ) is X i' Take M i',j The probability of H(Y|X i' =M i',j ) is X i' Take M i',j The total entropy of the label Y under the condition of .

6. The metabolite prediction method of multi-scale feature enhancement and feature selection according to claim 5, characterized in that: In the steps 2 and 3, the feature X is obtained by using the overall entropy of the biological sample metabolite category label and the conditional entropy of the biological sample metabolite category label. i' The information gain for metabolite category prediction is: IG(X i' ,Y)=H(Y)-H(Y|X i' ) Among them, IG(X i' ,Y) is feature X i' Information gain for metabolite classes.

7. The metabolite prediction method of multi-scale feature enhancement and feature selection according to claim 6, characterized in that: In step 3, the features in the multi-scale feature matrix are screened according to the information gain of each feature for the metabolite category to obtain the final feature set, specifically: Step 31: Preset a feature information gain threshold based on experience, compare the feature information gain threshold with the information gain of each feature for metabolite category prediction, retain features greater than the feature information gain threshold in the multi-scale feature matrix, and delete features less than or equal to the feature information gain threshold to obtain a multi-scale feature matrix after preliminary screening; Step 32: Using each multi-scale feature vector and the corresponding metabolite category label in the preliminarily screened multi-scale feature matrix to form an overall data set, the overall data set is divided into a training set and a test set, and multiple classification models are trained and tested using the training set and the test set, thereby obtaining a feature weight distribution model; Various classification models include: ANN model, support vector machine, random forest model; Step 3. Use the training set to train the feature weight allocation model multiple times. The feature weight allocation model assigns weights to each feature, and deletes the feature with the lowest weight in the training set each time. At the end of each training, the test set is used to test the performance of the feature weight allocation model. The accuracy of the feature weight allocation model first increases and then decreases during the process of gradual feature deletion. When the accuracy reaches the highest point, stop deleting features and form the final feature set with the remaining features.

8. The metabolite prediction method of multi-scale feature enhancement and feature selection according to claim 7, characterized in that: In step 32, the multiple classification models are trained and tested using the training set and the test set respectively, so as to obtain a feature weight allocation model, specifically: the trained model with the highest accuracy is used as the feature weight allocation model.

9. The metabolite prediction method of multi-scale feature enhancement and feature selection according to claim 8, characterized in that: The metabolite prediction model is obtained by the following method: The final feature set and metabolite category labels are used to form a final training set, and the classification model with the highest accuracy obtained in step 32 is trained using the final training set to obtain a metabolite prediction model.

Citation Information

Patent Citations

  • Metabolite recognition system based on molecular fingerprint prediction and application method thereof

    CN112735532A

  • Metabolite-based gastric cancer peritoneal metastasis prediction system and method

    CN118448051A