Fungus phenotype bar code screening method and system based on machine learning

Through a machine learning-based method, the importance of analyzing fungal phenotype characteristics using Gini coefficients, arrangement importance and additive interpreted values ​​has solved the shortcomings of the analysis of large fungal populations in the existing technology, and achieved more efficient identification and screening.

CN119993275APending Publication Date: 2025-05-13GUIZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510056177.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing fungal morphology research methods cannot conduct global and comprehensive analysis of all biological traits in all taxa, especially for large fungal populations with overlapping characteristics, and inconsistencies between phenotypes and genotypes are common.

Method used

The fungal phenotype barcode screening method based on machine learning is used to analyze the corresponding feature importance of each category through Gini coefficient, arrangement importance and additive interpretation value, thereby improving the diversity and accuracy of identification screening.

Benefits of technology

The importance analysis and screening of fungal phenotype characteristics has been realized, the diversity and accuracy of identification and screening have been improved, and the shortcomings of large fungal population analysis in the prior art have been solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993275A_ABST
    Figure CN119993275A_ABST
Patent Text Reader

Abstract

The invention provides a fungus phenotype bar code screening method based on machine learning. The method comprises the following steps: acquiring fungus data according to a published specimen and a fresh specimen; performing splitting conversion on phenotypic features in the fungus data to obtain unit features; performing data preprocessing on the unit features to obtain preprocessed data, and dividing the preprocessed data into a test set and a training set; optimal parameters of the classifier are selected through the test set and a supervised learning algorithm, and an optimal model is screened out according to the optimal parameters; based on the optimal model, performing importance analysis on the unit features to obtain an importance sequence; and sorting and screening bar code features according to importance. According to the method, the importance of the features corresponding to each category is analyzed through the Gini coefficient, the arrangement importance and the additive interpretation value, and the diversity and accuracy of identification and screening are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular to a fungal phenotypic barcode screening method and system based on machine learning. Background Art

[0002] Fungal taxonomy is an evolving field, always adapting to new discoveries and technologies. Given the huge diversity of fungal species, the continuous discovery of new taxa, and the intricate relationships between them, the challenges it faces are numerous and complex. Molecular systematics and morphology are two fundamental components of modern fungal taxonomy, which have greatly deepened our understanding of the fungal kingdom. On the one hand, molecular systematics can perform global analysis based on nucleotide sequences of all taxa; on the other hand, morphological analysis and comparison remain an essential part of our understanding of fungal diversity and evolution, and play a key role in the establishment of new taxa as well as production practices.

[0003] However, current fungal morphological research methods cannot perform a global and comprehensive analysis of all biological traits of all taxa, especially for large fungal groups with overlapping characteristics. And the understanding of typical characteristics can only be based on phylogenetic or niche comparisons, and the inconsistency between phenotypes and genotypes is still a common problem. Therefore, it is necessary to design a fungal phenotypic barcode screening method and system based on machine learning. Summary of the invention

[0004] The purpose of the present invention is to provide a method and system for fungal phenotypic barcode screening based on machine learning, so as to analyze the importance of the features corresponding to each category through Gini coefficient, permutation importance and additive explanatory value, so as to improve the diversity and accuracy of identification and screening.

[0005] To achieve the above object, the present invention provides the following solutions:

[0006] A fungal phenotypic barcode screening method based on machine learning comprises the following steps:

[0007] Acquire fungal data based on published specimens and fresh specimens; fungal data include: phenotypic characteristics, DNA sequences and biogeographic information;

[0008] The phenotypic characteristics are split and transformed to obtain unit characteristics; unit characteristics include: continuous characteristics and discrete characteristics;

[0009] Perform data preprocessing on unit characteristics to obtain preprocessed data, and divide the preprocessed data into a test set and a training set;

[0010] Select the best parameters of the classifier through the test set and supervised learning algorithm, and select the best model based on the best parameters;

[0011] Based on the best model, importance analysis of unit characteristics was performed to obtain importance ranking;

[0012] The barcode features were screened out based on the order of importance.

[0013] Optionally, data preprocessing is performed on the unit features to obtain preprocessed data, including:

[0014] Fill missing values ​​for continuous features using mean, median or k-nearest neighbor algorithm;

[0015] Convert discrete features into ordinal integers;

[0016] Both continuous features and ordinal integers are normalized to obtain normalized data, and high correlation features in the normalized data are eliminated through correlation analysis.

[0017] Optionally, based on the best model, importance analysis is performed on unit features to obtain importance rankings, specifically: the importance of unit features is ranked by additive explanatory value, Gini importance, and permutation importance to obtain global feature importance rankings and local feature importance rankings.

[0018] Optionally, the supervised learning algorithm includes: logistic regression algorithm, k-nearest neighbor algorithm, support vector machine algorithm, neural network algorithm, decision tree algorithm, extra tree algorithm, random forest algorithm, extreme random tree algorithm, gradient boosting algorithm, adaptive boosting algorithm and extreme gradient boosting algorithm.

[0019] Optionally, optimal parameters of the classifier are selected through a test set and a supervised learning algorithm, including: obtaining the optimal parameters through grid search, feature selection, and cross validation.

[0020] A fungal phenotypic barcode screening system based on machine learning, comprising:

[0021] The main control module includes: a reading unit and a model task unit; the reading unit is used to read the characteristics, labels and information of the best model of fungal data, and the model task unit is used to perform model training;

[0022] A preprocessing module is used to perform data preprocessing on unit features to obtain preprocessed data;

[0023] The model building module includes: a parameter selection unit and a supervised learning unit; the parameter selection unit is used to select the best parameters of the classifier, and the supervised learning unit has a built-in supervised learning algorithm;

[0024] The model screening module includes: a model selection unit and a graph generation unit; the model selection unit is used to screen out the best model, and the graph generation unit is used to generate a box plot and a confusion matrix;

[0025] Model interpretation module, the model interpretation unit includes: a SHAP value analysis unit and an importance analysis unit; the SHAP value analysis unit is used to rank the global and local importance of unit features by additive interpretation values, and the importance analysis unit is used to rank the Gini importance and permutation importance of unit features.

[0026] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects: the fungal phenotypic barcode screening method based on machine learning provided by the present invention comprises: obtaining fungal data based on published specimens and fresh specimens; splitting and converting the phenotypic features in the fungal data to obtain unit features; performing data preprocessing on the unit features to obtain preprocessed data, and dividing the preprocessed data into a test set and a training set; selecting the best parameters of the classifier through the test set and the supervised learning algorithm, and screening out the best model based on the best parameters; based on the best model, performing importance analysis on the unit features to obtain importance ranking; screening out the barcode features based on the importance ranking. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0028] Figure 1 This is a flow chart of the fungal phenotypic barcode screening method of the present invention;

[0029] Figure 2 This is a specific flow chart of the fungal phenotypic barcode screening method of the present invention;

[0030] Figure 3 This is a structural diagram of the fungal phenotypic barcode screening system of the present invention. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0032] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and understandable, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0033] like Figure 1 As shown, the present invention provides a fungal phenotypic barcode screening method based on machine learning, comprising the following steps:

[0034] Step 100: Acquire fungal data based on published specimens and fresh specimens; the fungal data includes: phenotypic characteristics, DNA sequences, and biogeographic information.

[0035] Step 200: Split and transform the phenotypic features to obtain unit features.

[0036] Step 300: perform data preprocessing on unit features to obtain preprocessed data, and divide the preprocessed data into a test set and a training set.

[0037] Step 400: Select the best parameters of the classifier through the test set and the supervised learning algorithm, and select the best model based on the best parameters.

[0038] Step 500: Based on the optimal model, importance analysis is performed on unit features to obtain importance ranking.

[0039] Step 600: Filter out barcode features according to importance ranking.

[0040] The specific process of the fungal phenotypic barcode screening method is as follows Figure 2 shown.

[0041] Specifically, except for the DNA sequence used for phylogenetic analysis, the rest of the fungal data is used as the original data set for subsequent analysis. The DNA sequence is subjected to phylogenetic analysis to obtain its phylogenetic classification status, which is used as a label for supervised learning.

[0042] Specifically, the specific steps of splitting and converting the phenotypic characteristics to obtain unit characteristics are as follows: decomposing the phenotypic characteristics into unit characteristics that can be represented by a single number or word, including two data types, one is a continuous feature that can be represented by numbers (i.e., numerical data), and the other is a discrete feature that cannot be represented by numbers (i.e., categorical data).

[0043] Specifically, the unit features are preprocessed to obtain preprocessed data as features for supervised learning, and the specific steps of dividing the preprocessed data into a test set and a training set are as follows: for numerical data, missing values ​​are directly filled according to algorithms such as the mean, median, or k-nearest neighbor algorithm (KNN). For categorical data, its missing values ​​are first defined as missing (NA), then converted into ordinal integers, and finally the quantitative features of numerical data and ordinal integers are scaled to [0, 1], and features with high correlation are eliminated using correlation analysis. The preprocessed data is then split into a test set and a training set by stratified sampling.

[0044] Specifically, supervised learning algorithms include:

[0045] Logistic Regression (LR), LR is a special case of a generalized linear model with binomial / Bernoulli conditional distribution and Logit link. It predicts the probability that a sample belongs to a certain category by inputting a linear combination of features into the Sigmoid function, and optimizes the model parameters by minimizing the cross entropy loss. During the training process, logistic regression uses the cross entropy loss function to measure the difference between the predicted probability and the actual label, and uses optimization algorithms (such as gradient descent) to adjust the model parameters to minimize the loss function, thereby improving the prediction accuracy of the model.

[0046] k-Nearest Neighbors (kNN), kNN determines the similarity between data points through distance metrics (including Euclidean distance, Manhattan distance, Chebyshev distance, or cosine similarity, etc.), and learns based on the data points that are closer to each query point in the feature space. When given a new unknown sample, kNN predicts the category of the new sample by finding the categories of the k nearest neighbors of the sample in the training set that are most similar to the sample, where k is a hyperparameter indicating the number of nearest neighbors considered when making decisions. This algorithm does not require an explicit training phase, only the training set needs to be stored. For classification tasks, the kNN algorithm predicts the category of the new sample by majority voting based on the categories of the k nearest neighbors.

[0047] Support Vector Machines (SVM) is a machine that searches for an optimal hyperplane (a straight line in two-dimensional space and a plane in three-dimensional space) in the feature space, through which data points of different categories are separated as much as possible, that is, the interval between two categories is maximized. This interval is called the "support vector", which is the point closest to the hyperplane and determines the position and direction of the hyperplane.

[0048] Neural Networks (NN) simulates the connection and information processing methods of human brain neurons. The input layer receives data, and then extracts features and performs nonlinear transformations through one or more hidden layers, and finally generates prediction results at the output layer. During the training process, the network calculates the predicted value through forward propagation, and uses the loss function to evaluate the prediction error. Then, the gradient is calculated and the weight is updated through the back-propagation algorithm, so as to iteratively learn and optimize the network structure until it can accurately perform specific tasks.

[0049] Decision Trees (DT), DT starts from the root node and recursively applies feature selection algorithms (such as information gain, Gini impurity) to select the best features and thresholds to split the original data set, creating a series of branches based on whether-then rules until the leaf nodes of each branch represent a final classification or regression result. This process involves feature selection, recursive splitting, and post-pruning techniques to build a tree model that can map from input features to output results, ultimately forming a decision tree classifier or regression tree.

[0050] ExtraTrees (ET) algorithm, ET constructs multiple DTs to perform classification or regression tasks. When constructing each decision tree, ET does not randomly sample samples, but uses the entire original data set. At each decision node, a feature and a random split point of the feature are randomly selected for data splitting. This extreme randomness increases the randomness of the model, improves the generalization ability of the model and reduces the risk of overfitting. Finally, ET integrates the prediction results of multiple DTs and obtains the final prediction result by majority voting (classification) or average (regression).

[0051] Random Forests (RF), RF extracts multiple sub-samples from the training set with replacement through bootstrap sampling to construct multiple DTs. Each tree only considers randomly selected sub-samples during training to increase the randomness and generalization ability of the model. In the decision-making process, RF uses a majority voting mechanism to determine the final category for classification problems, and calculates the average of all DT prediction results for regression problems. RF can evaluate the importance of features, provide a basis for feature selection, effectively reduce the risk of overfitting, and improve the stability and prediction accuracy of the model.

[0052] Extremely Randomized Tree (ERT), ERT performs classification or regression tasks by constructing multiple DTs. The entire training set is used when constructing each DT. At each decision node, ERT does not look for the best splitting attribute and threshold, but randomly selects features and random splitting thresholds. This increases the randomness of the model, reduces variance and improves the generalization ability of the model. Specifically, for categorical features, ERT randomly selects samples of certain categories as the left branch and the rest as the right branch; for numerical features, ERT randomly selects any number between the maximum and minimum values ​​of the feature attribute as the splitting threshold, and randomly distributes samples to two branches based on this value.

[0053] Gradient Boosting (GB), GB optimizes the prediction performance of the model by gradually adding weak learners (usually DT). The algorithm starts with an initial prediction model, and then in each iteration, a new weak learner is trained for the residual (i.e., prediction error) of the previous step model, using the gradient descent principle to guide the learning of the new weak learner through the negative gradient direction of the loss function. After training, each weak learner will be weighted and incorporated into the final model with a certain learning rate (shrinkage rate). The learning rate is used to control the contribution of each weak learner to the overall prediction result to avoid overfitting. The iterative process continues until the preset number of iterations is reached or the model performance is no longer significantly improved. Finally, the prediction results of all weak learners are accumulated to form the final prediction of the GB model. This algorithm can effectively process high-dimensional data and improve the stability and accuracy of model performance.

[0054] Adaptive Boosting (AdaBoost, AB), AB is an ensemble learning method for classification tasks, which builds a strong classifier by combining multiple weak classifiers. At the time of initialization, each sample in the training set is given the same weight. In each iteration, the algorithm trains a weak classifier based on the current sample weight and increases the weight of the classifier with a small error rate. At the same time, the sample weights are updated so that the misclassified samples have higher weights in subsequent iterations. Finally, AB combines all weak classifiers into a strong classifier through weighted majority voting, in which the classifier with a higher weight has a greater influence in the final decision. This process continues until the preset number of iterations is reached or the model performance is no longer significantly improved.

[0055] eXtreme Gradient Boosting (XGB), XGB is based on the DT model and uses a variety of optimization techniques to improve the accuracy and computational efficiency of the model. The algorithm gradually improves the model by iteratively building a series of ordered DTs, and each tree tries to correct the prediction error of the previous tree. The core is to use the gradient descent method to optimize the objective function, which includes the prediction error of the model and the regularization term of the model complexity to control overfitting. XGB has multiple built-in loss functions and allows users to customize the loss function according to specific problems. In addition, the complexity of the model is controlled by L1 and L2 regularization terms, a depth-first tree growth strategy is adopted, and the depth of the tree is controlled by early stopping and regularization terms. XGB includes parallel and distributed computing methods, can run on multiple CPU or GPU cores, improves the model training speed, and has an efficient memory usage strategy that can handle large-scale data sets.

[0056] More specifically, grid search is used to test different parameter combinations of machine learning classifiers, and the best parameter combination is selected as the optimal parameter according to the weighted score in the process of predicting the test set data. The optimal parameters are then verified by feature selection and cross-validation. Box plots and confusion matrices are generated during the prediction process. After the optimal model is obtained, the data of fresh specimens can be used as the validation set to manually verify the prediction accuracy of the optimal model.

[0057] Furthermore, grid search searches for the best parameter combination by systematically traversing all possible parameter combinations on a predefined parameter grid. Specifically, a set of discrete values ​​for each hyperparameter is first determined, and then a complete parameter grid is constructed. For each set of parameters in the grid, grid search trains a model and uses a scoring function to evaluate its performance. Ultimately, grid search selects the parameter combination that performs best on the test set.

[0058] Furthermore, feature selection optimizes the performance of machine learning models by identifying and selecting the most relevant feature subsets from the feature set. Specifically: first, generate feature subsets to provide candidate sets for the evaluation function; second, use statistical tests (such as chi-square test, mutual information) or model-driven methods (such as feature importance scoring based on tree models) to evaluate the quality of feature subsets; then, determine the optimal feature subset through model training and verification; finally, select the final feature subset to improve the accuracy of the model, reduce computational costs, and enhance the interpretability of the model. Through feature selection, the most relevant and informative feature subsets are effectively selected from the feature set, which greatly improves the performance and interpretability of the model.

[0059] Furthermore, cross-validation randomly divides the dataset into K subsets and trains the model K times, using K-1 subsets of data for training each time, while the reserved subset is used to verify the model performance. By calculating the average of the performance indicators in all validation processes, a robust and unbiased model performance estimation result is obtained. The use of cross-validation reduces the evaluation bias caused by the data division method and improves the accuracy of model evaluation.

[0060] More specifically, the optimal parameters of the classifier are selected through the test set and the supervised learning algorithm, and the following is included: generating a box plot for calculating the difference in evaluation indicators between different supervised learning algorithms and a confusion matrix for showing the specific performance of the test set data during the training process of different classifiers. The evaluation indicators include: accuracy, precision, recall rate and F1 score.

[0061] Specifically, based on the optimal model, the specific steps of analyzing the importance of unit features and obtaining importance ranking are as follows: the importance of unit features that affect the prediction results is ranked by SHapley additive explanatory value (SHAP value), Gini importance and permutation importance to obtain global feature importance ranking and local feature importance ranking.

[0062] Specifically, the specific steps of selecting barcode features according to importance ranking are as follows: through global feature importance ranking and local feature importance ranking, several unit features with high ranking and high importance can be understood. According to the phenotypic characteristics corresponding to these unit features, the overall characteristic variation of a fungal group and the important representative characteristics of a specific lineage in evolution are determined, and the barcode features that can be used for morphological identification are selected.

[0063] The present invention also provides a fungal phenotypic barcode screening system based on machine learning. Figure 3 As shown, including:

[0064] The main control module includes: a reading unit (LoadData) and a model task unit (ExcelParser); the reading unit includes a first reading subunit (load_excel) and a second reading subunit (load_model); the first reading subunit is used to read the features and label information in the Excel table, and the second reading subunit is used to read the trained model (optimal model) information. The model task unit includes: a first task subunit (train), a second task subunit (predict) and a third task subunit (explain); the first task subunit, the second task subunit and the third task subunit are used to perform model training, model prediction and model explanation respectively.

[0065] The preprocessing module includes: a first processing unit (verification), a second processing unit (normalization) and a third processing unit (correlation); the first processing unit, the second processing unit and the third processing unit are respectively used for preliminary verification of data, normalization processing of data and correlation analysis operations.

[0066] The model building module includes: a parameter selection unit (Classification) and a supervised learning unit (Supervised); the parameter selection unit includes a first parameter selection subunit (test_clf) and a second parameter selection subunit (select_clf); the first parameter selection subunit is used to test the performance of the classifier, and the second parameter selection subunit is used to perform the screening function of the best model; the supervised learning unit has a variety of supervised learning algorithms built in.

[0067] The model screening module includes: a model selection unit (Evaluation) and a graph generation unit (Graphics); the model selection unit includes: a first model selection subunit (feature_selection), a second model selection subunit (grid_search), a third model selection subunit (cross_validation) and a fourth model selection subunit (eva_metrics); the first model selection subunit is used to provide a feature screening function, the second model selection subunit is used to perform grid search, the third model selection subunit is used to perform cross validation, and the fourth model selection subunit is used to calculate evaluation indicators. The graph generation unit includes: a first generation subunit (metric) for generating a box plot and a second generation subunit (cm) for generating a confusion matrix.

[0068] Model explanation module, including: SHAP value analysis unit (Explain) and importance analysis unit (Interpret); SHAP value analysis unit is used to sort the global and local importance of unit features by additive explanation value; SHAP value analysis unit includes: honeycomb subunit (summary_plot) and force field subunit (force_plot); honeycomb subunit is used to generate honeycomb plot, and force field subunit is used to generate force field plot. Importance analysis unit includes: lr feature analysis subunit (lr) and tree model feature subunit (gini_perm); lr feature analysis subunit is used to analyze the feature importance of LR, and tree model feature subunit is used to analyze other Gini importance and permutation importance based on tree model.

[0069] The beneficial effects of the present invention are as follows:

[0070] 1) Be able to identify the category to which a given feature belongs and analyze the importance of the features corresponding to each category;

[0071] 2) Machine learning can process a large number of objective features, thereby screening out accurate and effective features as the basis for classification;

[0072] 3) Through a variety of supervised learning algorithms, the prediction accuracy, stability and generalization ability of the model are improved, and the risk of overfitting is reduced.

[0073] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0074] The present invention uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only used to help understand the method and core ideas of the present invention. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A fungal phenotypic barcode screening method based on machine learning, characterized in that: The steps include: Acquire fungal data based on published specimens and fresh specimens; the fungal data include: phenotypic characteristics, DNA sequences and biogeographic information; Splitting and converting the phenotypic characteristics to obtain unit characteristics; the unit characteristics include: continuous characteristics and discrete characteristics; Performing data preprocessing on the unit features to obtain preprocessed data, and dividing the preprocessed data into a test set and a training set; Selecting optimal parameters of the classifier through the test set and the supervised learning algorithm, and screening out the best model according to the optimal parameters; Based on the optimal model, performing importance analysis on the unit characteristics to obtain importance ranking; The barcode features are selected according to the importance ranking.

2. The fungal phenotypic barcode screening method based on machine learning according to claim 1, characterized in that: Performing data preprocessing on the unit characteristics to obtain preprocessed data includes: Filling missing values ​​of the continuous features by using the mean value, median or k-nearest neighbor algorithm; Converting the discrete features into ordinal integers; The continuous features and the ordinal integers are both normalized to obtain normalized data, and high correlation features in the normalized data are eliminated through correlation analysis.

3. The fungal phenotypic barcode screening method based on machine learning according to claim 1, characterized in that: Based on the optimal model, the importance of the unit features is analyzed to obtain an importance ranking, specifically: the importance of the unit features is ranked by additive explanatory value, Gini importance and permutation importance to obtain a global feature importance ranking and a local feature importance ranking.

4. The fungal phenotypic barcode screening method based on machine learning according to claim 1, characterized in that: The supervised learning algorithms include: logistic regression algorithm, k-nearest neighbor algorithm, support vector machine algorithm, neural network algorithm, decision tree algorithm, extra tree algorithm, random forest algorithm, extreme random tree algorithm, gradient boosting algorithm, adaptive boosting algorithm and extreme gradient boosting algorithm.

5. The fungal phenotypic barcode screening method based on machine learning according to claim 1, characterized in that: The optimal parameters of the classifier are selected through the test set and the supervised learning algorithm, specifically: the optimal parameters are obtained through grid search, feature selection and cross validation.

6. A fungal phenotypic barcode screening system based on machine learning, characterized in that: include: A main control module, the main control module includes: a reading unit and a model task unit; the reading unit is used to read the characteristics, labels and information of the best model of fungal data, and the model task unit is used to perform model training; A preprocessing module is used to perform data preprocessing on unit features to obtain preprocessed data; A model building module, the model building module includes: a parameter selection unit and a supervised learning unit; the parameter selection unit is used to select the best parameters of the classifier, and the supervised learning unit has a built-in supervised learning algorithm; A model screening module, the model screening module comprising: a model selection unit and a graph generation unit; the model selection unit is used to screen out the best model, and the graph generation unit is used to generate a box plot and a confusion matrix; A model interpretation module, wherein the model interpretation unit comprises: a SHAP value analysis unit and an importance analysis unit; the SHAP value analysis unit is used to sort the global and local importance of the unit features by additive interpretation values, and the importance analysis unit is used to sort the Gini importance and permutation importance of the unit features.

Citation Information

Cited By

  • Simulation test key factor screening-oriented method and system and computer program product

    CN121413456A