Model training method and related equipment

By comprehensively using multiple feature selection methods such as partial least squares regression, random forest and breadth-first search, the target features are determined and training samples are constructed, which solves the problem of uneven model performance in existing technologies and improves the overall performance and stability of the model.

CN120632469AActive Publication Date: 2025-09-12SHENZHEN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511123790.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-09-12
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

In the existing technology, the feature selection method is relatively simple, which leads to good performance of the model in some dimensions but poor performance in other dimensions, resulting in poor overall performance of the trained model.

Method used

A variety of feature selection methods are used to extract features from the data, and the target features are determined by the features extracted by different feature selection methods. The performance is good in the dimensions that different feature selection methods focus on, so as to construct training samples to train a model with excellent performance.

Benefits of technology

Through the comprehensive application of multiple feature selection methods, the performance and stability of the model have been significantly improved, especially in high-dimensional data and small sample scenarios. The problems of feature redundancy and overfitting have been solved, and the quality of feature selection and the generalization ability of the model have been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632469A_ABST
    Figure CN120632469A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and related equipment, and belongs to the technical field of model training. The method comprises the following steps: acquiring a plurality of pieces of first data, and determining a plurality of features selected by a feature selection mode in features of each piece of first data by adopting different feature selection modes so as to construct a first feature set corresponding to the feature selection mode; determining target features related to a prediction target according to features in a first feature set corresponding to various feature selection modes; extracting a feature matched with the target feature from each piece of first data to construct a training sample corresponding to each piece of first data, and training a preset model according to each training sample to obtain a prediction model, the prediction model is used for predicting to-be-predicted data to obtain a probability parameter corresponding to the prediction target. According to the invention, the performance of the trained model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of model training technology, and specifically relates to a model training method and related equipment. Background Art

[0002] With the development of large model technology, corresponding models can be built in more and more fields.

[0003] In the exemplary technology, features related to the model prediction target are selected from the data, and training samples are constructed based on the features, so that the model is trained through the training samples.

[0004] However, the above feature selection method is relatively simple, which makes the selected features have better performance in the dimensions focused on by the feature selection method, but poor performance in other dimensions, resulting in poor performance of the trained model. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to provide a model training method and related equipment to solve the problem of poor performance of the model.

[0006] In a first aspect, an embodiment of the present application provides a model training method, comprising: Acquire multiple first data, and adopt different feature selection methods to determine multiple features selected by the feature selection method among the features of each of the first data to construct a first feature set corresponding to the feature selection method; Determining target features related to the prediction target based on the features in the first feature set corresponding to the various feature selection methods; Features matching the target features are extracted from each of the first data to construct training samples corresponding to each of the first data, and a preset model is trained based on each of the training samples to obtain a prediction model, wherein the prediction model is used to predict the data to be predicted to obtain probability parameters corresponding to the prediction target.

[0007] In one embodiment, determining target features related to the prediction target based on the features in the first feature set corresponding to the various feature selection methods includes: Acquire second data in each database, where the second data includes at least one feature in the first feature set; determining a first prediction contribution value of a feature in each data set to the prediction target, wherein one of the data sets is constructed from respective second data corresponding to at least one database; Determining, according to the first prediction contribution value, a plurality of first features in the features of each of the data sets to construct a second feature set corresponding to the data set, wherein a first prediction contribution value corresponding to the first feature is greater than a first preset threshold; An intersection between the second feature sets corresponding to each of the data sets is determined, and the second feature in the intersection is determined as a target feature related to the prediction target.

[0008] In one embodiment, determining a first prediction contribution value of a feature in each data set to the prediction target includes: Determining a type corresponding to the data set, where the type is used to indicate whether the data set is a training set or a test set; Determine, according to the type corresponding to the data set, a first prediction contribution value of the features in the data set to the prediction target.

[0009] In one embodiment, determining target features related to the prediction target based on the features in the first feature set corresponding to the various feature selection methods includes: Determining second prediction contribution values ​​of features in the first feature set corresponding to various feature selection methods to the prediction target; According to the second predicted contribution value, a target feature is determined from the features in the first feature set corresponding to the various feature selection methods, and the second predicted contribution value of the target feature is greater than a second preset threshold.

[0010] In one embodiment, obtaining a plurality of first data includes: Acquire multiple candidate data and determine the sample type to which each candidate data belongs, wherein the sample type includes positive samples and negative samples; According to the sample type of the candidate data, a plurality of first data are determined in each of the candidate data, wherein the sample type of half of the first data in each of the first data is a positive sample.

[0011] In one embodiment, obtaining a plurality of candidate data includes: Acquire multiple initial data, and determine the feature missing rate corresponding to each of the initial data based on the features in the initial data; Determining intermediate data from each of the initial data according to the feature missing rate, wherein the feature missing rate of the intermediate data is less than a preset missing rate; The intermediate data is filled with features to obtain candidate data.

[0012] In a second aspect, an embodiment of the present application provides a model training device, comprising: an acquisition module, configured to acquire a plurality of first data, and adopt different feature selection methods to determine, among the features of each of the first data, a plurality of features selected by the feature selection method, so as to construct a first feature set corresponding to the feature selection method; a determination module, configured to determine target features related to the prediction target based on features in the first feature set corresponding to the various feature selection methods; An extraction module is used to extract features that match the target features in each of the first data to construct training samples corresponding to each of the first data, and to train a preset model based on each of the training samples to obtain a prediction model, wherein the prediction model is used to predict the data to be predicted to obtain probability parameters corresponding to the prediction target.

[0013] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.

[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0016] In an embodiment of the present application, multiple first data are obtained, and different feature selection methods are used to determine the multiple features selected by the feature selection method in the features of each first data to construct a feature set corresponding to the feature selection method, and based on the features in the various feature selection methods, the target features related to the prediction target are determined, and the features matching the target features are extracted from each first data to construct a training sample corresponding to the first data, and then the prediction model is obtained by training each training sample for prediction. In this embodiment, the features in the data are extracted by different feature selection methods, and then the intersection of the features extracted by the different feature selection methods is obtained. The target features related to the prediction target are determined by the intersection in the dimensions focused on by different feature selection methods. The training samples constructed based on the target features are trained. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is one of the flowcharts of a model training method according to an exemplary embodiment; Figure 2This is a second flow chart of a model training method according to an exemplary embodiment; Figure 3 is a schematic diagram of a target feature acquisition process according to an exemplary embodiment; Figure 4 This is a third flow chart of a model training method according to an exemplary embodiment; Figure 5 This is a fourth flowchart of a model training method according to an exemplary embodiment; Figure 6 is a structural block diagram of a model training device according to an exemplary embodiment; Figure 7 The figure is a structural block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0018] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0019] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally refers to before and after.

[0020] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and corresponding operation entrances must be provided for users to choose to authorize or refuse.

[0021] With the development of large model technology, corresponding models can be built in more and more fields.

[0022] In the exemplary technology, features related to the model prediction target are selected from the data, and training samples are constructed based on the features, so that the model is trained through the training samples.

[0023] However, the above feature selection method is relatively simple, which makes the selected features have better performance in the dimensions focused on by the feature selection method, but poor performance in other dimensions, resulting in poor performance of the trained model.

[0024] In response to the problem of poor model performance, the inventors of this application came up with the idea of ​​extracting features from the data through different feature selection methods, and then determining the target features related to the prediction target through the features extracted by different feature selection methods. The target features have good performance in the dimensions that different feature selection methods focus on, so that the performance of the model trained by the training samples constructed based on the target features is excellent.

[0025] The model training method and related equipment provided in the embodiments of the present application are described in detail below through specific embodiments and their application scenarios.

[0026] The model training method provided in the embodiments of the present application can be applied to vehicle driving application scenarios.

[0027] The model training method provided in the embodiments of the present application is described in detail below.

[0028] Reference Figure 1 , Figure 1 One of the flow charts of the model training method provided in this application is as follows: Figure 1 As shown, the model training method includes the following steps: In step S101 , a plurality of first data are acquired, and different feature selection methods are used to determine a plurality of features selected by the feature selection method among the features of each first data, so as to construct a first feature set corresponding to the feature selection method.

[0029] In this embodiment, the execution subject is a model training device. For ease of description, the term "device" is used below to refer to the model training device. The device can be a terminal device with model training capabilities, for example, a server or a computer with large computing power.

[0030] The device obtains multiple first data. The first data may be clinical data of a patient after desensitization. In addition, each first data may be obtained from a different database, for example, the first data may be obtained from a different ICU (Intensive Care Unit) database.

[0031] The first data may include the following types of information: Demographic information: age, sex, race, and weight; Complications: peripheral vascular disease, cerebrovascular disease, dementia, chronic lung disease, etc.; Laboratory indicators: 47 examination indicators including hemoglobin, white blood cells, creatinine, glucose, etc.; Vital signs: heart rate, respiratory rate, systolic blood pressure, diastolic blood pressure, body temperature, etc.

[0032] The device is provided with a target variable, which is the target that the trained model needs to predict. For example, the target variable can be cardiac arrest, that is, the predicted target is cardiac arrest. Of course, the predicted target can also be other types of diseases, not limited to cardiac arrest, and the first data is not limited to clinical data, but can also be any other type of data. The specific type of the first data is determined based on the field of application of the model that the device needs to train. For example, if the device needs to train a model in the financial field, the first data is financial data; if the device needs to train a model in the medical field, the first data is clinical data.

[0033] When the first data is clinical data, the device retrieves the first data from the ICU database using screening criteria. Specifically, the device first filters data from the ICU database based on the screening criteria to obtain data from different ICU databases. The screening criteria include, but are not limited to: patient age greater than 18 years; no history of congestive heart failure, myocardial infarction, or malignant tumors; ICU stay exceeding three hours; and for patients with multiple ICU stays, the first admission record is retrieved.

[0034] The screened data is desensitized, that is, identification labels such as the patient's name are removed from the data, to obtain the first data.

[0035] After obtaining the first data, the device uses different feature selection methods to determine multiple features selected by the feature selection methods from the features of each first data. Different feature selection methods include partial least squares regression, random forest feature selection, and breadth-first search feature selection. The following describes how different feature selection methods are used to determine features.

[0036] 1. Partial least squares regression feature selection method: The input data (first data) is standardized to eliminate the influence of different feature dimensions.

[0037] Build a partial least squares regression model: Use partial least squares regression to build a classification model between the features in the input data and the target variable (predicted target). Partial least squares regression extracts latent components in the first data as features by maximizing the covariance between the input features and the target variable.

[0038] Feature importance assessment: The weight of each feature in the first data is determined by the least squares regression model, so as to evaluate the importance of the feature to the target variable by the weight, and retain the features with higher weights as preliminary candidate features.

[0039] Filter feature subsets: Set a threshold to filter out the features with the strongest correlation with the target variable. The features with weights greater than the threshold are the features selected by this feature selection method.

[0040] 2. Feature selection method of random forest: Generating shadow features: Randomly shuffle the features in the first data to generate a set of "shadow features". These shadow features are used to provide a reference for random noise.

[0041] Train a random forest model: Combine the original features in the first data with the shadow features, train a random forest model, and evaluate the importance of the features.

[0042] Comparing feature importance: Compare the importance scores of each original feature and the shadow feature. If a feature's importance is significantly higher than that of the shadow feature, it is marked as a "confirmed relevant feature." If a feature's importance is close to that of the shadow feature, it is marked as an "uncertain feature." If a feature's importance is lower than that of the shadow feature, it is marked as an "irrelevant feature."

[0043] Iterative optimization: Repeat the above process multiple times to gradually remove irrelevant features until all features are classified as "confirmed relevant features" or "irrelevant features".

[0044] Output feature subset: Output all "confirmed relevant features", that is, the confirmed relevant features are the features selected by the random forest feature selection method.

[0045] 3. Feature selection method of breadth-first search Objective: To evaluate the contribution of candidate clinical variables to the model prediction performance one by one. Candidate clinical variables are features in the first data, such as white blood cells, red blood cells, lymphocytes, anion gap, heart rate, etc.

[0046] Methods: The currently selected optimal clinical variable set is combined with each candidate clinical variable to form a temporary variable set.

[0047] Using a logistic regression model, train the model based on a set of temporary variables and predict the probability of the target variable.

[0048] Calculate model performance metrics, including the receiver operating characteristic (ROC) curve and area under the curve (AUC) to measure overall predictive ability, and sensitivity and specificity to assess model performance across different categories. Save the performance metric for each candidate clinical variable in the results dictionary for subsequent comparison.

[0049] Update the optimal set of clinical variables: Objective: To gradually construct the optimal set of clinical variables based on the evaluation results.

[0050] Methods: In the current round, the clinical variables with the highest performance indicators (such as AUC values) were selected.

[0051] If the addition of a variable significantly improves model performance (e.g., a higher AUC), it is included in the optimal clinical variable set. The selected variable is removed from the candidate variable list to avoid repeated evaluation. If no variable further improves performance, the search is terminated.

[0052] Correlation as a tie resolver: Objective: To break ties when multiple variables behave similarly through correlation analysis.

[0053] Methods: The correlation between each candidate clinical variable and the target variable was calculated. Candidate variables were ranked from highest to lowest correlation. Variables with the highest correlation with the target variable were retained first. These retained variables were then used as features for the breadth-first search feature selection method.

[0054] Each feature selection method selects multiple features from the first data, and the features selected by the feature selection method constitute a feature set corresponding to the feature selection method, that is, each feature selection method has a corresponding feature set, and the feature set is defined as a first feature set.

[0055] In this embodiment, the three feature selection methods mentioned above can efficiently process high-dimensional data from multiple dimensions. Among them, the partial least squares regression feature selection method extracts the most explanatory features by maximizing the covariance between the features and the target variable, which is particularly suitable for high-dimensional data and small sample scenarios; the random forest feature selection method comprehensively identifies relevant features based on the random forest, avoiding the problem of information omission caused by only screening the optimal features; the breadth-first search feature selection method simulates the process of natural selection and genetic variation, and gradually iterates to find the global optimal feature subset, ensuring the efficiency and stability of feature selection. This comprehensive use of multiple feature selection methods helps to overcome the shortcomings of single feature selection technology in comprehensiveness and accuracy, and significantly improves the quality of feature selection.

[0056] Step S102 : determining target features related to the prediction target based on the features in the first feature set corresponding to various feature selection methods.

[0057] After obtaining the first feature set corresponding to each feature selection method, target features related to the prediction target can be determined based on the features in each first feature set.

[0058] Exemplarily, an intersection between each first feature set is determined, and this intersection is defined as a first intersection. The device then determines a target feature associated with the predicted target based on the features in the first intersection. Exemplarily, all features in the first intersection can be used as target features. Exemplarily, the predicted target is cardiac arrest, and the multiple target features associated with cardiac arrest are GCS (Glasgow Coma Scale) score, creatinine, and white blood cell count.

[0059] In step S103, features matching the target features are extracted from each first data to construct training samples corresponding to each first data, and a preset model is trained based on each training sample to obtain a prediction model, wherein the prediction model is used to predict the data to be predicted to obtain probability parameters corresponding to the prediction target.

[0060] After obtaining the target features, features matching the target features are extracted from each first data to construct training samples corresponding to the first data. For example, if the multiple target features are GCS score, creatinine, and white blood cell count, then GCS score, creatinine, and white blood cell count are extracted from the first data and labels are assigned to these three target features to obtain training samples corresponding to the first data. Labels may be, for example, cardiac arrest or non-cardiac arrest.

[0061] After obtaining each training sample in the above manner, the preset model is trained based on each training sample to obtain a prediction model, which is used to predict the data to be tested to obtain the probability parameters corresponding to the prediction target. For example, if the prediction target is cardiac arrest, the prediction model outputs the probability of cardiac arrest in the patient to whom the data to be tested belongs. The prediction model can be deployed on the device or on any other hardware device to predict the prediction target. The prediction model can be applied to fields such as disease prediction and personalized treatment. The exemplary training process of the prediction model is as follows: 1. Data preparation and division: Input features: The screened key features (such as GCS score, creatinine, and white blood cells) were used as input variables of the model.

[0062] Target variable: Based on the research objectives, select the corresponding label (whether the patient suffered cardiac arrest) as the target variable.

[0063] 2. Random Forest Model Construction Model selection: Random forest is used as the basic model because of its strong robustness and ability to capture nonlinear relationships.

[0064] Hyperparameter Optimization: Use the Bayesian optimization algorithm to automatically tune the hyperparameters of the random forest. Tuned hyperparameters include, but are not limited to, the number of decision trees, the maximum depth of a single decision tree, the minimum number of samples required for a node split, the minimum number of samples required for a leaf node, the maximum number of features considered by each tree when splitting, and the goal of Bayesian optimization is to minimize the cross-validation loss function (such as AUC).

[0065] Model training and 5-fold cross-validation: The training set is divided into 5 subsets, 4 of which are used as training data each time, and the remaining 1 subset is used as validation data.

[0066] In each round of validation, the performance indicators of the model are recorded, such as AUC, precision, recall, etc. Finally, the average of the five validation results is taken as the performance evaluation of the model.

[0067] Model training: Retrain the random forest model on the entire training set using the optimized hyperparameters.

[0068] Model performance evaluation: Test set evaluation: Use an independent test set to evaluate the generalization performance of the model and calculate the following metrics: AUC (Area Under Curve, the area under the ROC (receiver operating characteristic curve) curve): measures the model's ability to distinguish between positive and negative samples.

[0069] Accuracy: The proportion of correctly classified samples.

[0070] Recall: The proportion of samples that are actually positive that are correctly predicted to be positive.

[0071] The prediction model constructed above is evaluated. The evaluation steps can be implemented through the following sub-steps: 1.1. Build a prediction model using training samples extracted from the MIMIC database (an ICU database) as the internal training set and training samples extracted from the eICU database (another ICU database) as the external test set. 1.2. Calculate the AUC value of the model to evaluate its predictive performance: The internal test AUC value was 0.697 (95% confidence interval: 0.675-0.719); The AUC value of the external test was 0.756 (95% confidence interval: 0.745-0.767), indicating that the external test performed better than the internal test.

[0072] 2.1. Using the eICU database as the internal training set and the MIMIC database as the external test set, we rebuilt the prediction model and evaluated its performance: The internal test AUC value was 0.768 (95% confidence interval: 0.756-0.780); The AUC value of the external test was 0.661 (95% confidence interval: 0.656-0.686), indicating that the performance of the external test was lower than that of the internal test.

[0073] In this embodiment, a set of efficient solutions is proposed for the common feature redundancy and overfitting problems in high-dimensional data and small sample scenarios, which helps to improve the practical application effect of feature selection. By reducing the data dimension and retaining key information through the partial least squares regression feature selection method, combining the random forest feature selection method to fully explore potential features, and using the breadth-first search feature selection method to find the global optimal solution, this embodiment significantly reduces the feature redundancy phenomenon while improving the explanatory power and practicality of the features. In addition, the extracted key features can be widely used in machine learning modeling, disease prediction, personalized treatment and other fields, especially in complex scenarios such as medical big data. This not only helps to improve the performance and stability of the model, but also provides strong technical support for research in fields such as precision medicine, demonstrating the dual value of this embodiment in theoretical innovation and practical application.

[0074] In this embodiment, multiple first data are obtained, and different feature selection methods are used to determine the multiple features selected by the feature selection method in the features of each first data to construct a feature set corresponding to the feature selection method, and based on the features in the various feature selection methods, target features related to the prediction target are determined, and features matching the target features are extracted from each first data to construct training samples corresponding to the first data, and then a prediction model is obtained through training of each training sample for prediction. In this embodiment, features in the data are extracted by different feature selection methods, and then the intersection of the features extracted by the different feature selection methods is obtained. The target features related to the prediction target are determined by the intersection. The performance on the dimensions focused on by different feature selection methods is good, so that the performance of the model trained based on the training samples constructed based on the target features is excellent.

[0075] Reference Figure 2 , Figure 2 This is the second flow chart of the model training method provided by this application, based on Figure 1 In the embodiment shown, step S102 includes: Step S201: Obtain second data from each database, where the second data includes at least one feature in the first feature set.

[0076] In this embodiment, the device obtains second data from each database based on the features in each first feature set, and each second data includes at least one feature in the first feature set.

[0077] Exemplarily, each first feature set is combined to form a target feature set, and then data containing at least one feature in the target feature set is retrieved from the database as the second data.

[0078] The device may obtain data from multiple databases, for example, obtain data containing features in the target feature set in database A as second data, and obtain data containing features in the target feature set in database B as second data.

[0079] Step S202 : determining a first prediction contribution value of a feature in each data set to a prediction target, wherein a data set is constructed by respective second data corresponding to at least one database.

[0080] After obtaining each second data, at least one data set is constructed through each second database. The data set is composed of the second data corresponding to at least one database. For example, each second data extracted from database A is constructed into a data set, and each second data extracted from database B is constructed into another data set, and the second data extracted from databases A and B are constructed into a data set.

[0081] After obtaining each data set, the contribution value of the features in each data set to the prediction target is determined, that is, the contribution value of the features in each second data item in each data set to the prediction target is determined. The contribution value can be calculated using a random forest model or a contribution value calculation formula. The contribution value calculated by the device is defined as the first prediction contribution value.

[0082] In one example, the first predicted contribution value of the feature in each piece of second data in each data set is calculated using a random forest model or a contribution value calculation formula.

[0083] In another example, the device determines the type corresponding to the data set, and determines the first prediction contribution value of the feature in the data set to the prediction target based on the type corresponding to the data set. The type is used to indicate whether the data set is a training set or a test set. Exemplarily, the data set corresponding to database A is used as the training set, and the data set in database B is used as the test set, and the first prediction contribution value of the feature of each second data in the data sets corresponding to AB is calculated; the data set corresponding to database A is used as the test set, and the data set in database B is used as the training set, and the first prediction contribution value of the feature of each second data in the data sets corresponding to AB is calculated; the first prediction contribution value of the feature of the second data in the data sets corresponding to databases A+B is calculated, and the data sets corresponding to databases A+B are the training sets. It should be noted that when the data set is a training set, the product of the contribution value of the feature of the second data and a larger weight is calculated as the first prediction contribution value; when the data set is a test set, the product of the contribution value of the feature of the second data and a smaller weight is calculated as the first prediction contribution value.

[0084] Step S203 , based on the first prediction contribution value, multiple first features are determined in the features of each data set to construct a second feature set corresponding to the data set, and the first prediction contribution value corresponding to the first feature is greater than a first preset threshold.

[0085] After obtaining the first predicted contribution value corresponding to the feature in each data set, multiple first features are determined in the features of each data set based on the first predicted contribution value, and the first predicted contribution value of the first feature is greater than the first preset threshold. Exemplarily, each feature of the data set has a first predicted contribution value, and a first predicted contribution value greater than the first preset threshold is determined in each feature in the data set, and the feature corresponding to the first predicted contribution value greater than the first preset threshold is used as the first feature of the data set. If the number of first predicted contribution values ​​greater than the first preset threshold is greater than the set number, these first predicted contribution values ​​are sorted, and the features corresponding to the set number of first predicted contribution values ​​in the first sorting are obtained as the first features. For example, if the set number is 10, the features corresponding to the first ten first predicted contribution values ​​are obtained as the first features. The first features corresponding to each data set constitute the second feature set.

[0086] Step S204: determining the intersection between the second feature sets corresponding to each data set, and determining the second feature in the intersection as the target feature related to the prediction target.

[0087] Each data set corresponds to a second feature set, and the intersection between the second feature sets is determined, and the intersection is defined as a second intersection. Each feature in the second intersection is used as a target feature related to the prediction target.

[0088] Reference Figure 3 , Figure 3 This is a simplified flowchart of the target feature screening process in this embodiment. For example, the device obtains three datasets. Each different dataset is subjected to feature selection using a different feature selection method, resulting in multiple first feature sets (unlabeled). The union of the first feature sets for each dataset is determined, i.e., each dataset has a union, such as Union 1 for the first dataset, Union 2 for the second dataset, and Union 3 for the third dataset. Each union is then input into an RF model (random forest model), which performs SHAP calculations on each feature in each of the five datasets obtained from the previous three datasets. SHAP is a game-theoretic method used to explain the contribution of each feature to the prediction results in a machine learning model. Its core concept is to assess the importance of each feature by calculating its marginal contribution under all possible feature combinations. SHAP can be considered a feature contribution value. The five datasets are: MIMIC (test set), MIMIC (training set), eICU (test set), eICU (training set), and MIMIC+eICU (training set). The features in each data set are calculated and screened using SHAP to obtain a second feature set. For example, the top ten SHAP features in the data set are constructed as the second feature set of the data set.

[0089] Multiple generalizable features can be obtained by intersecting the second feature sets, and each generalizable feature is a target feature.

[0090] In this embodiment, by introducing contribution value analysis and cross-database screening processes, key features with high stability and strong generalization capabilities can be effectively extracted. Specifically, in three different databases, MIMIC (for example, database A), eICU (for example, database B), and MIMIC + eICU, features are extracted based on three feature selection methods, and important features that rank high in both the internal test set and the external test set are screened out through intersection operations, thereby ensuring the stability and consistency of the features. In addition, by comprehensively analyzing the results of multiple databases, key features applicable across databases are further screened out, significantly enhancing the versatility of features in different data sets. This approach helps to solve the problem that the exemplary technology is unstable in feature selection results and difficult to adapt to cross-dataset application scenarios, provides reliable guarantees for feature optimization in complex data environments, and improves the robustness and practical application value of the model.

[0091] Reference Figure 4 , Figure 4 This is the third flow chart of the model training method provided in this application, based on Figure 1 or Figure 2In the embodiment shown, step S102 includes: Step S401 : determining second prediction contribution values ​​of features in a first feature set corresponding to various feature selection methods to a prediction target.

[0092] In this embodiment, the first feature sets are combined to obtain a target feature set, and the device calculates the contribution value of the features in the target feature set as the second predicted contribution value based on the contribution value or the random forest model.

[0093] Step S402 : determining a target feature from among the features in the first feature set corresponding to various feature selection methods according to the second predicted contribution value, wherein the second predicted contribution value of the target feature is greater than a second preset threshold.

[0094] After obtaining the second predicted contribution value corresponding to each feature in the target feature set, the device determines the target feature based on the second predicted contribution value and the features in the first feature set corresponding to each feature selection method, wherein the second predicted contribution value of the target feature is greater than a second preset threshold. It should be noted that the first preset threshold and the second preset threshold are both set values, and the first preset threshold and the second preset threshold can be any composite value.

[0095] Exemplarily, the features in the target feature set are sorted in descending order of the second prediction contribution value, and the features with second prediction contribution values ​​greater than a preset threshold are taken as preliminary features. If the number of preliminary features is greater than the set number, the preliminary features ranked first are taken as target features, and the number of target features is the set number.

[0096] In this embodiment, the device determines the second prediction contribution value of the features in the first feature set corresponding to various feature selection methods to the predicted target, and thus accurately determines the target feature among the features in the first feature set corresponding to various feature selection methods based on the second prediction value.

[0097] Reference Figure 5 , Figure 5 This is the fourth flow chart of the model training method provided in this application, based on Figures 1 to 4 In any of the embodiments shown in FIG, step S101 includes: Step S501: Acquire multiple candidate data and determine the sample type to which each candidate data belongs. The sample type includes positive samples and negative samples.

[0098] In this example, the number of positive and negative samples in the database varies. For example, positive samples are patients with cardiac arrest, while negative samples are healthy individuals without cardiac arrest. Since the number of patients with cardiac arrest is far less than the number of healthy individuals without cardiac arrest, the number of negative samples far exceeds the number of positive samples. To improve the generalization and predictive performance of the model, the data needs to be downsampled to make the number of positive and negative samples uniform.

[0099] In one example, the device obtains multiple data from multiple databases as candidate data, and determines the sample type to which the candidate data belongs based on the information in the candidate data. The sample type includes positive samples and negative samples. The positive sample is a patient with disease A, and the negative sample is a normal person who does not have disease A.

[0100] In another example, the device obtains multiple initial data from a database, and determines the feature missing rate of each initial data based on the features in each initial data, and determines the intermediate data in each initial data based on the feature missing rate. The feature missing rate of the intermediate data is less than the preset missing rate, and then the features of the intermediate data are filled to obtain candidate data.

[0101] For example, the feature missing rate in the initial data is first counted, that is, all features in the initial data are traversed and the missing rate of each feature is calculated. The missing rate is the ratio of the number of missing values ​​to the total number of samples.

[0102] Features are divided into two categories: low missing rate features (missing rate <30%): retained and imputed; high missing rate features (missing rate ≥ 30%): directly eliminated, because a higher missing rate may lead to unreliable imputation results.

[0103] The device separates low missing rate features: extracts features with a missing rate lower than 30% and forms a new feature matrix for subsequent interpolation operations.

[0104] Standardized feature data: Since feature filling is sensitive to the dimension of the feature, the data needs to be standardized before performing feature difference filling.

[0105] To this end, the features are normalized to ensure that each feature has the same scale. The normalization formula is: Among them, x represents the original value, μ represents the feature mean, and σ represents the feature standard deviation.

[0106] The device applies the missing value interpolation method and first initializes the missing value interpolation parameters, that is, sets the parameters for missing value interpolation, including: n_neighbors: The number of neighbors K to be selected, usually set to an integer between 3 and 10.

[0107] weights: Neighbor weight type, you can choose "uniform" (uniform weight) or "distance" (distance weighted).

[0108] Metric: distance metric, usually Euclidean distance ("euclidean").

[0109] Specifically, the device calculates the distance between neighboring samples. For each sample with missing values, it calculates the characteristic distance between it and other samples (ignoring the dimension where the missing value is located). The distance calculation formula is: in, and Represent two samples respectively. represents the number of features, and represents two adjacent samples.

[0110] The device selects K nearest neighbors: based on the calculated distance, finds the K neighbor samples that are most similar to the current sample.

[0111] The device calculates the estimated value of the missing value: for each missing value, calculate the average value of its K neighbor samples on the corresponding feature as the filling value: if the "uniform" weight is used, the simple mean is taken; if the "distance" weight is used, the mean is calculated weighted by distance.

[0112] Device fills missing values: The calculated filling values ​​are inserted into the corresponding missing positions in the original data set.

[0113] Device denormalization: After interpolation is completed, if the data has been normalized before, the data needs to be denormalized back to the original scale. The denormalization formula is: .

[0114] Step S502 : determining a plurality of first data in each candidate data according to the sample type of the candidate data, wherein half of the first data in each first data are of positive sample type.

[0115] After determining the sample type of the candidate data, multiple first data are determined in each candidate data based on the sample type, and half of the samples of the first data in each first data are positive samples, that is, half of the first data in each first data are data of patients with diseases, and half of the samples are data of normal people without diseases.

[0116] In this embodiment, by selecting half of the positive sample data and half of the negative sample data, the positive and negative samples of the training samples used for training are balanced, thereby improving the generalization ability and prediction performance of the training samples.

[0117] Based on the same inventive concept, this application also provides a model training device. Figure 6 The model training device provided in the embodiment of the present application is described in detail.

[0118] Figure 6 The figure is a structural block diagram of a model training device according to an exemplary embodiment.

[0119] like Figure 6 As shown, the model training device 600 may include: An acquisition module 610 is configured to acquire a plurality of first data and, using different feature selection methods, determine a plurality of features selected by the feature selection method from among the features of each first data to construct a first feature set corresponding to the feature selection method; A determination module 620 is configured to determine target features related to the prediction target based on features in the first feature set corresponding to various feature selection methods; The extraction module 630 is used to extract features that match the target features in each first data to construct a training sample corresponding to each first data, and train the preset model based on each training sample to obtain a prediction model, wherein the prediction model is used to predict the data to be predicted to obtain probability parameters corresponding to the prediction target.

[0120] In one embodiment, the model training apparatus 600 further includes: Acquire second data in each database, where the second data includes at least one feature in the first feature set; Determining a first prediction contribution value of a feature in each data set to a prediction target, wherein a data set is constructed from respective second data corresponding to at least one database; According to the first prediction contribution value, a plurality of first features are determined in the features of each data set to construct a second feature set corresponding to the data set, wherein the first prediction contribution value corresponding to the first feature is greater than a first preset threshold; An intersection between the second feature sets corresponding to each data set is determined, and the second feature in the intersection is determined as a target feature related to the prediction target.

[0121] In one embodiment, the model training apparatus 600 further includes: Determine the type of data set. The type is used to indicate whether the data set is a training set or a test set. According to the type corresponding to the data set, a first prediction contribution value of the feature in the data set to the prediction target is determined.

[0122] In one embodiment, the model training apparatus 600 further includes: Determine a second prediction contribution value of features in the first feature set corresponding to various feature selection methods to the prediction target; According to the second predicted contribution value, a target feature is determined from the features in the first feature set corresponding to various feature selection methods, and the second predicted contribution value of the target feature is greater than a second preset threshold.

[0123] In one embodiment, the model training apparatus 600 further includes: Obtain multiple candidate data and determine the sample type of each candidate data, which includes positive samples and negative samples; According to the sample type of the candidate data, a plurality of first data are determined in each candidate data, wherein the sample type of half of the first data in each first data is a positive sample.

[0124] In one embodiment, the model training apparatus 600 further includes: Obtain multiple initial data, and determine the feature missing rate corresponding to each initial data based on the features in each initial data; According to the feature missing rate, intermediate data is determined in each initial data, and the feature missing rate of the intermediate data is less than the preset missing rate; Fill in the features of the intermediate data to obtain candidate data.

[0125] The model training device provided in the embodiment of the present application can achieve Figure 1-Figure 5 The various processes implemented in the method embodiment achieve the same technical effect and will not be described again here to avoid repetition.

[0126] In some embodiments, as Figure 7 As shown, an embodiment of the present application also provides an electronic device 700, including a processor 701 and a memory 702, wherein the memory 702 stores a program or instruction that can be run on the processor 701, and when the program or instruction is executed by the processor 701, the various steps of the above-mentioned model training method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0127] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0128] An embodiment of the present application also provides a computer-readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned model training method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0129] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory, a random access memory, a magnetic disk, or an optical disk.

[0130] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned model training method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0131] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0132] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.

[0133] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A model training method, characterized in that: include: Acquire multiple first data, and adopt different feature selection methods to determine multiple features selected by the feature selection method among the features of each of the first data to construct a first feature set corresponding to the feature selection method; Determining target features related to the prediction target based on the features in the first feature set corresponding to the various feature selection methods; Features matching the target features are extracted from each of the first data to construct training samples corresponding to each of the first data, and a preset model is trained based on each of the training samples to obtain a prediction model, wherein the prediction model is used to predict the data to be predicted to obtain probability parameters corresponding to the prediction target.

2. The model training method according to claim 1, characterized in that Determining target features related to the prediction target based on the features in the first feature set corresponding to the various feature selection methods includes: Acquire second data in each database, where the second data includes at least one feature in the first feature set; determining a first prediction contribution value of a feature in each data set to the prediction target, wherein one of the data sets is constructed from respective second data corresponding to at least one database; Determining, according to the first prediction contribution value, a plurality of first features in the features of each of the data sets to construct a second feature set corresponding to the data set, wherein a first prediction contribution value corresponding to the first feature is greater than a first preset threshold; An intersection between the second feature sets corresponding to each of the data sets is determined, and the second feature in the intersection is determined as a target feature related to the prediction target.

3. The model training method according to claim 2, characterized in that Determining a first prediction contribution value of a feature in each data set to the prediction target includes: Determining a type corresponding to the data set, where the type is used to indicate whether the data set is a training set or a test set; Determine, according to the type corresponding to the data set, a first prediction contribution value of the features in the data set to the prediction target.

4. The model training method according to claim 1, characterized in that Determining target features related to the prediction target based on the features in the first feature set corresponding to the various feature selection methods includes: Determining second prediction contribution values ​​of features in the first feature set corresponding to various feature selection methods to the prediction target; According to the second predicted contribution value, a target feature is determined from the features in the first feature set corresponding to the various feature selection methods, and the second predicted contribution value of the target feature is greater than a second preset threshold.

5. The model training method according to claim 1, characterized in that The acquiring of a plurality of first data includes: Acquire multiple candidate data and determine the sample type to which each candidate data belongs, wherein the sample type includes positive samples and negative samples; According to the sample type of the candidate data, a plurality of first data are determined in each of the candidate data, wherein the sample type of half of the first data in each of the first data is a positive sample.

6. The model training method according to claim 5, characterized in that The obtaining of multiple candidate data includes: Acquire multiple initial data, and determine the feature missing rate corresponding to each of the initial data based on the features in the initial data; Determining intermediate data from each of the initial data according to the feature missing rate, wherein the feature missing rate of the intermediate data is less than a preset missing rate; The intermediate data is filled with features to obtain candidate data.

7. A model training device, characterized in that: include: an acquisition module, configured to acquire a plurality of first data, and adopt different feature selection methods to determine, among the features of each of the first data, a plurality of features selected by the feature selection method, so as to construct a first feature set corresponding to the feature selection method; a determination module, configured to determine target features related to the prediction target based on features in the first feature set corresponding to the various feature selection methods; An extraction module is used to extract features that match the target features in each of the first data to construct training samples corresponding to each of the first data, and to train a preset model based on each of the training samples to obtain a prediction model, wherein the prediction model is used to predict the data to be predicted to obtain probability parameters corresponding to the prediction target.

8. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the model training method as described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the model training method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product is stored in a storage medium, and the computer program product is executed by at least one processor to implement the model training method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Feature selection method and device, and network equipment identification method and device

    CN118332452A

  • Feature-based data screening and predicting method and system

    CN120296373A

  • Quantum annealing-based device and method for searching for novel material

    WO2025095499A1