High-frequency hearing loss prediction method
Through the combination of RFECV and XGBoost models, key features are automatically screened and hyperparameters are optimized, which solves the problem of feature selection subjectivity and model performance limitations of high-frequency hearing loss prediction, and is effective and accurate high-frequency hearing loss prediction, suitable for large-scale health data analysis and early screening.
Patent Information
- Application Number
- CN202510623771.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-19
AI Technical Summary
The existing high-frequency hearing loss prediction methods have problems such as strong subjectivity in feature selection, large-scale limitations in model prediction performance, inability to apply on a large scale, and low detection efficiency, making it difficult to achieve efficient and accurate high-frequency hearing loss prediction.
The RFECV algorithm is used for feature selection, and the high-frequency hearing loss prediction model is trained in combination with the XGBoost model. Hyperparameters are optimized through grid search and 10-fold cross-validation, and key features related to high-frequency hearing loss are automatically screened out to build a high-performance prediction model.
It realizes automated and accurate prediction of high-frequency hearing loss, improves the generalization performance and prediction accuracy of the model, is suitable for large-scale health data analysis, and supports early high-frequency hearing loss screening and personalized intervention.
Smart Images

Figure CN120511073A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical health technology, and in particular to a method for predicting high-frequency hearing loss, which is a method for predicting high-frequency hearing loss based on feature selection and machine learning. Background Art
[0002] High-frequency hearing loss is a common hearing problem, usually caused by aging, noise exposure, genetic factors or other health problems. If high-frequency hearing loss can be detected early, measures can be taken to prevent further hearing damage and improve quality of life. However, traditional hearing detection methods (such as audiometry) have the following shortcomings: First, they rely on professional equipment and professional operation, which is costly and not convenient enough; second, they cannot be applied on a large scale and it is difficult to quickly screen a large number of people; third, the detection efficiency is low, the detection process is time-consuming, and it is not suitable for daily health management.
[0003] In recent years, with the development of artificial intelligence and big data technologies, researchers have begun to try to use health data to predict high-frequency hearing loss. However, existing prediction methods have the following problems:
[0004] 1. Too many and too complex features: Health data contains a large number of indicators, but not all indicators are useful for predicting high-frequency hearing loss. Too many useless indicators will make the model complex and inaccurate.
[0005] 2. The model is not reliable enough: Some models perform well in the laboratory but do not work well in practical applications.
[0006] 3. Lack of ability to automatically screen features: Many methods rely on manual experience to select features, which is time-consuming and cannot guarantee the scientificity and accuracy of feature screening.
[0007] Therefore, given the above problems and shortcomings of existing high-frequency hearing loss prediction methods, there is an urgent need for a method that can automatically screen key features and efficiently and accurately predict high-frequency hearing loss. Summary of the Invention
[0008] The technical problem to be solved by the present invention is to provide a method for predicting high-frequency hearing loss, which can simply and efficiently predict high-frequency hearing loss.
[0009] In order to solve the above technical problems, the invention adopts the following technical solutions:
[0010] A method for predicting high-frequency hearing loss is provided, the method comprising the following steps:
[0011] S1: Data preparation: collect multi-dimensional data related to human hearing health as raw data, pre-process the raw data to obtain basic data, and divide the basic data into two parts: a first-level training set and a first-level test set;
[0012] S2: Construct an initial feature set and use the RFECV algorithm for feature selection to screen out several key features that are most relevant to predicting high-frequency hearing loss from the initial feature set;
[0013] S3: constructing a high-frequency hearing loss prediction model, using the key features most relevant to the prediction of high-frequency hearing loss screened in step S2 to construct the high-frequency hearing loss prediction model; obtaining a secondary training set and a secondary test set;
[0014] S4: training a high-frequency hearing loss prediction model, using the key features and secondary training set most relevant to the prediction of high-frequency hearing loss screened out in step S2, and training the high-frequency hearing loss prediction model based on an XGBoost model;
[0015] S5: Using the secondary test set, the high-frequency hearing loss prediction model trained in step S4 is evaluated.
[0016] Preferably, in step S1, the preprocessing of the original data includes filling missing values, data transformation and / or standardization; the specific method of dividing the basic data into two parts, a first-level training set and a first-level test set, is: taking each tested person as the basic unit of division, each tested person corresponds to a piece of data, and the piece of data corresponding to each tested person is determined to enter the first-level training set or the first-level test set by random selection, and the ratio of the data volume of the first-level training set to the first-level test set is 8:2; the piece of data corresponding to each tested person includes all indicator data associated with the tested person.
[0017] Preferably, step S2 further includes the following sub-steps: S2.1: artificially constructing an initial feature set based on past experience, the initial feature set including several features related to predicting high-frequency hearing loss; S2.2: the RFECV algorithm uses a gradient boosting classifier algorithm as its evaluator; S2.3: using all features of the initial feature set to train the gradient boosting classifier algorithm, the gradient boosting classifier algorithm screens out the least important feature from the initial feature set and removes it, thereby completing the zeroth round of iteration and obtaining the first feature set; S2.4: using all features of the first feature set to train the gradient boosting classifier algorithm, the RFECV algorithm performs 10-fold cross validation, and calculates the F1 value of the 10-fold cross validation. The average of the scores is recorded as the F1 score of the first round of iteration; then the gradient boosting classifier algorithm filters out the feature with the lowest importance from the first feature set and eliminates it, thereby completing the first round of iteration and obtaining the second feature set; S2.5: repeating step S2.4 until the last round of iteration is completed and only one feature remains in the current feature set; S2.6: the RFECV algorithm sorts the F1 scores from the first round of iteration to the last round of iteration in descending order, and obtains the maximum F1 score, then obtains the round of iteration that produces the maximum F1 score and the number of features in the corresponding feature set, and uses the features contained in the feature set that produces the maximum F1 score as the key features most relevant to predicting high-frequency hearing loss.
[0018] Preferably, in the step S2.3, "the gradient boosting classifier algorithm filters out the feature with the lowest importance from the initial feature set and eliminates it" is specifically as follows: the gradient boosting classifier algorithm scores the importance of each feature of the initial feature set according to its own feature importance scorer, and then filters out the feature with the lowest importance score from the initial feature set and eliminates it; in the step S2.4, "the gradient boosting classifier algorithm filters out the feature with the lowest importance from the first feature set and eliminates it" is specifically as follows: the gradient boosting classifier algorithm scores the importance of each feature of the first feature set according to its own feature importance scorer, and then filters out the feature with the lowest importance score from the first feature set and eliminates it; in the step S2.4, "the RFECV algorithm performs 10-fold cross validation, calculates the average F1 score of the 10-fold cross validation and records it as the F1 score of the first round of iteration" is specifically as follows: all the data in the first-level training set are evenly divided into 10 mutually exclusive data blocks, and the data amount of the 10 data blocks is equal or nearly equal. Specifically, the data in the first-level training set are evenly divided into A1, A2, A3, A4, A5, A6, A7, A8, A9, and A10 data blocks. In the first cycle, the RFECV algorithm uses the A2, A3, A4, A5, A6, A7, A8, A9, and A10 data blocks to train the gradient boosting classifier algorithm, and uses the A1 data block to calculate the F1 score of the first cycle and record the F1 score of the first cycle; in the second cycle, the RFECV algorithm uses the A1, A3, A4, A5, A6, A7, A8, A9, and A10 data blocks to train the gradient boosting classifier algorithm. , A5, A6, A7, A8, A9, and A10 data blocks are used to train the gradient boosting classifier algorithm, and the A2 data block is used to calculate the F1 score of the second cycle and record the F1 score of the second cycle; and so on, until the RFECV algorithm completes the 10th cycle, and the A10 data block is used to calculate the F1 score of the 10th cycle and record the F1 score of the 10th cycle; calculate the arithmetic mean of the F1 scores of the 10 cycles of 10-fold cross validation, and record the arithmetic mean as the F1 score of the first round of iteration.
[0019] Preferably, in step S2.6, 18 key features most relevant to predicting high-frequency hearing loss are screened out from the initial feature set, and the 18 key features are specifically: province, race, marital status, gender, age group, education level, ear examination record, smoking status, daily activities, dairy product intake, bean intake, fruit intake, platelet count, alkaline phosphatase, total cholesterol, triglycerides, tinnitus and hearing level.
[0020] Preferably, in step S3, "obtaining a secondary training set and a secondary test set" specifically includes: modifying the primary training set according to the several key features most relevant to the prediction of high-frequency hearing loss obtained in step S2, retaining the data related to the primary training set, and deleting the remaining irrelevant data, to obtain a secondary training set; modifying the primary test set according to the several key features most relevant to the prediction of high-frequency hearing loss obtained in step S2, retaining the data related to the primary test set, and deleting the remaining irrelevant data, to obtain a secondary test set.
[0021] Preferably, the step S4 further includes the following sub-steps: S4.1: Select the XGBoost model as the basic machine learning model; S4.2: Construct a hyperparameter grid of the XGBoost model; S4.3: Initialize the grid search cross validator and configure the key parameters of the grid search cross validator; S4.4: Train the XGBoost model by traversing all hyperparameter combinations in the hyperparameter grid of the XGBoost model through the grid search cross validator; S4.5: Select the best hyperparameter, the grid search cross validator compares the F1 scores corresponding to all hyperparameter combinations in the hyperparameter grid of the XGBoost model, and selects the hyperparameter combination with the highest F1 score as the best hyperparameter combination and stores it; S4.6: Best model training, the grid search cross validator uses the best hyperparameter combination obtained in step S4.5 to reconfigure a new XGBoost model, and then uses the secondary training set to train the new XGBoost model to obtain the best XGBoost model and store it.
[0022] Preferably, the step S4.2 further includes the following sub-steps: S4.2.1: defining the hyperparameter grid of the XGBoost model to include four hyperparameters: the number of trees, the learning rate, the maximum depth of the tree, and the minimum number of samples required to split internal nodes, and setting two candidate values for each hyperparameter; S4.2.2: listing all combinations of the four hyperparameters of the hyperparameter grid of the XGBoost model based on the candidate values; in the step S4.2.1, the candidate values of the number of trees are set to 100 and 200, the candidate values of the learning rate are set to 0.01 and 0.1, the candidate values of the maximum depth of the tree are set to 3 and 5, and the candidate values of the minimum number of samples required to split internal nodes are set to 2 and 5.
[0023] Preferably, step S4.3 further includes the following sub-steps: S4.3.1: specifying the evaluator used by the high-frequency hearing loss prediction model as the XGBoost model of step S4.1; S4.3.2: the grid search cross-validator adopts and obtains the hyperparameter grid of the XGBoost model constructed in step S4.2; S4.3.3: setting the number of cross-validation folds to 10 folds; S4.3.4: specifying the evaluation indicator of cross-validation as the F1 score; S4.3.5: setting parallel computing to use all available CPU cores to speed up the grid search process.
[0024] Preferably, the step S4.4 further includes the following sub-steps: S4.4.1: for any hyperparameter combination in the hyperparameter grid of the XGBoost model, the grid search cross validator performs 10-fold cross validation, obtains the F1 score corresponding to the current hyperparameter combination and records it; S4.4.2: repeats the step S4.4.1 until the grid search cross validator obtains the F1 score corresponding to each hyperparameter combination in the hyperparameter grid of the XGBoost model and records it; the step S4.4.1 further includes the following content: for any hyperparameter combination in the hyperparameter grid of the XGBoost model, divide all the data in the secondary training set into 10 mutually exclusive data blocks on average, and the data amounts of the 10 data blocks are equal or approximately equal. Specifically, divide the data in the secondary training set into B1, B2, B3, B4, B5, B6, B7, B8, B9, and B10 data blocks. In the first cycle, the grid search cross validator uses the B2, B3, B4, B5, B6, B7, B8, B9, and B10 data blocks to train the XGBoost model configured with the current hyperparameter combination, and uses the B1 data block to calculate the F1 score of the first cycle and record the F1 score of the first cycle; in the second cycle, the B1, B3, B4, B5, B6, B7, B8, B9, and B10 data blocks are used to train the XGBoost model configured with the current hyperparameter combination, and uses the B2 data block to calculate the F1 score of the second cycle and record the F1 score of the second cycle; and so on, until the grid search cross validator completes the 10th cycle, and uses the B10 data block to calculate the F1 score of the 10th cycle and record the F1 score of the 10th cycle; calculate the arithmetic mean of the F1 scores of the 10 cycles of 10-fold cross validation, and take the arithmetic mean as the F1 score corresponding to the current hyperparameter combination and record it.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] (1) The high-frequency hearing loss prediction method provided by the present invention is a high-frequency hearing loss prediction method based on feature selection and machine learning. This prediction method effectively overcomes the subjectivity of feature selection and the limitations of model prediction performance of traditional methods. It can not only automatically screen key features, but also efficiently and accurately predict whether a person has high-frequency hearing loss, thereby significantly improving the automation, precision and reliability of high-frequency hearing loss prediction, and providing technical support for early accurate screening and personalized intervention of high-frequency hearing loss.
[0027] (2) The high-frequency hearing loss prediction method provided by the present invention uses a grid search and 10-fold cross-validation (Grid Search CV) to systematically search for the optimal hyperparameter combination of the XGBoost model. The optimal parameters are selected based on the cross-validation performance, and then the optimal XGBoost prediction model (best_model) is trained on the complete training set using the optimal hyperparameters. The entire model training process of this prediction method is designed to maximize the model's generalization performance and prediction accuracy, ensuring that the model performs well even on unknown data.
[0028] (3) The high-frequency hearing loss prediction method provided by the present invention can realize automated feature selection and automatically screen key features through the RFECV algorithm, avoiding the subjectivity of relying on manual experience judgment and saving a lot of time and energy.
[0029] (4) The high-frequency hearing loss prediction method provided by the present invention can construct a high-performance prediction model. The model constructed based on the XGBoost algorithm performs well in multiple evaluation indicators, especially the recall rate and AUC value, indicating that the model can well identify patients with high-frequency hearing loss.
[0030] (5) The high-frequency hearing loss prediction method provided by the present invention can screen out 18 key features that are most relevant to predicting high-frequency hearing loss from the initial feature set. These 18 key features are health data that are easy for ordinary people to understand and collect, such as age, gender, eating habits, etc., which are convenient for promotion and application.
[0031] (6) The high-frequency hearing loss prediction method provided by the present invention is highly practical and has a wide range of applications. This prediction method can be applied to large-scale health data analysis, providing technical support for medical institutions or health management platforms, assisting in early high-frequency hearing loss screening, quickly identifying high-risk groups, and providing strong support for subsequent refined diagnosis and personalized intervention. The prediction results of this prediction method are presented in the form of classification results of whether high-frequency hearing loss has occurred, providing users and medical professionals with intuitive and actionable decision-making reference information. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0033] Figure 1 This is a flow chart of a method for predicting high-frequency hearing loss provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0034] In order to explain the present invention more clearly, the present invention is further described below in conjunction with preferred embodiments. Those skilled in the art should understand that the following specific description is illustrative rather than restrictive and should not be used to limit the scope of protection of the present invention.
[0035] like Figure 1 As shown, this embodiment provides a high-frequency hearing loss prediction method, which is applicable to humans. The prediction method is a high-frequency hearing loss prediction method based on feature selection and machine learning, and the prediction method includes the following steps:
[0036] S1: Data preparation: collect multi-dimensional data related to human hearing health as raw data, pre-process the raw data to obtain basic data, and divide the basic data into two parts: a first-level training set and a first-level test set;
[0037] S2: Construct an initial feature set and use the Recursive Feature Elimination with Cross-Validation (RFECV) algorithm for feature selection to screen out several key features most relevant to predicting high-frequency hearing loss from the initial feature set;
[0038] S3: Constructing a high-frequency hearing loss prediction model, using the key features most relevant to the prediction of high-frequency hearing loss screened in step S2 above to construct a high-frequency hearing loss prediction model; obtaining a secondary training set and a secondary test set;
[0039] S4: Training a high-frequency hearing loss prediction model, using the key features and secondary training set most relevant to the prediction of high-frequency hearing loss screened in step S2 above, and training the high-frequency hearing loss prediction model based on the XGBoost model;
[0040] S5: Use the secondary test set to evaluate the high-frequency hearing loss prediction model trained in step S4 above.
[0041] The specific contents of each step of the high-frequency hearing loss prediction method of this embodiment are described below one by one:
[0042] S1: Data preparation, collecting multi-dimensional data related to human hearing health as raw data;
[0043] In step S1 above, multi-dimensional data related to human hearing health is collected as raw data, such as demographic information (e.g., gender, age), lifestyle habits (e.g., smoking status, dietary habits), and physical examination indicators (e.g., platelet count, cholesterol level). Data related to human hearing health can be collected through, for example, epidemiological surveys, such as physical examination data of people in a certain region and lifestyle survey data.
[0044] In the above step S1, in order to ensure the data quality and meet the standardization of the machine learning model input, the original data needs to be preprocessed. The preprocessing includes filling missing values, data transformation and / or standardization processing, etc., to obtain basic data. Specifically, in response to the problem of missing data, this embodiment adopts multiple imputation to fill missing values to reduce the potential impact of missing data on model performance; in response to the problem of data type diversity, this embodiment adopts data encoding technology, such as weight of evidence (WOE) encoding and ordinal encoding, to convert categorical features into numerical features for easy processing by the machine learning model; in response to the problem of inconsistent dimensions of different features, this embodiment adopts Z-score normalization method to scale different features to a unified scale and eliminate dimensional differences to improve the convergence speed and stability of the machine learning model; Yeo-Johnson transformation is used for processing continuous variables. Yeo-Johnson transformation is a powerful data transformation technology, which has the advantage of being able to process continuous data containing zero and negative values. By performing a parameterized power transformation on the data, the transformed data distribution can be made closer to a normal distribution, thereby helping to stabilize variance, reduce data skewness, and potentially improve the performance and stability of subsequent machine learning models. This data preprocessing effectively reduces data noise, eliminates dimensionality effects, and significantly improves the training effectiveness and prediction accuracy of machine learning models.
[0045] In the above step S1, the basic data is divided into two parts: a first-level training set and a first-level test set, and the ratio of the data volume of the first-level training set to the first-level test set is 8:2, wherein the first-level training set is used to train the model, and the first-level test set is used to test the effect of the model. The specific method of dividing the basic data into two parts: a first-level training set and a first-level test set is as follows: take each person being tested as the basic unit of division, each person being tested corresponds to a piece of data, and the piece of data corresponding to each person being tested is determined by random selection to enter the first-level training set or the first-level test set, and the ratio of the data volume of the first-level training set to the first-level test set is 8:2. In this embodiment, for example, there are 8725 people being tested, and each person being tested corresponds to a piece of data, totaling 8725 pieces of data. 6980 pieces of data are randomly selected to enter the first-level training set, and the remaining 1745 pieces of data enter the first-level test set. The ratio of the data volume of the first-level training set to the first-level test set is 6980:1975, that is, 8:2. It should be noted that the data corresponding to each person being tested includes all indicator data associated with the person being tested, such as the demographic information (such as gender, age), living habits (such as whether smoking, eating habits), and physical examination indicators (such as platelet count, cholesterol level) associated with the person being tested.
[0046] S2: Construct an initial feature set and use the Recursive Feature Elimination with Cross-Validation (RFECV) algorithm for feature selection to screen out several key features most relevant to predicting high-frequency hearing loss from the initial feature set;
[0047] The RFECV algorithm is an efficient and robust feature selection method. Its core idea is to iteratively train the model and evaluate feature importance, gradually eliminating redundant or unimportant features. It then combines cross-validation to select the optimal feature subset, thereby reducing feature dimensionality while ensuring model performance. The RFECV algorithm is an intelligent tool that can gradually remove indicators that are not very helpful for prediction, retaining only the most important features. Its advantage is that there is no need to manually guess which indicators are useful; the RFECV algorithm can automatically find the most relevant indicators.
[0048] The above step S2 further includes the following sub-steps:
[0049] S2.1: An initial feature set is artificially constructed based on previous experience. The initial feature set includes several features related to the prediction of high-frequency hearing loss.
[0050] For example, the initial feature set includes height, weight, skin color, province, race, marital status, sex, age group, education level, work status, headphone wearing status, ear check record, smoking status, drinking status, daily activities, dairy product intake, bean intake, fruit intake, meat intake, dessert intake, platelet count, alkaline phosphatase, total cholesterol, triglycerides, blood pressure, heart rate, tinnitus, etc.
[0051] S2.2: The RFECV algorithm uses the Gradient Boosting Classifier (GBC) algorithm as its estimator.
[0052] S2.3: Using all features of the initial feature set (assuming the initial feature set contains, for example, 30 features), the gradient boosting classifier algorithm is trained. The gradient boosting classifier algorithm selects the least important feature from the initial feature set and removes it, thereby completing the zeroth iteration and obtaining the first feature set (the first feature set contains, for example, 29 features).
[0053] In step S2.3 above, "the gradient boosting classifier algorithm selects the least important feature from the initial feature set and removes it" specifically involves: the gradient boosting classifier algorithm scores the importance of each feature in the initial feature set according to its own feature importance scorer (feature_importances_), and then selects the feature with the lowest importance score (i.e., the least important feature) from the initial feature set and removes it;
[0054] S2.4: The gradient boosting classifier algorithm is trained using all features of the first feature set (the first feature set contains, for example, 29 features). The RFECV algorithm performs 10-fold cross-validation, and the average F1 score (F1-score) of the 10-fold cross-validation is calculated and recorded as the F1 score for the first iteration. The gradient boosting classifier algorithm then selects the least important feature from the first feature set and removes it, thereby completing the first iteration and obtaining a second feature set (the second feature set contains, for example, 28 features).
[0055] In step S2.4 above, "the gradient boosting classifier algorithm selects the feature with the lowest importance from the first feature set and removes it" specifically includes: the gradient boosting classifier algorithm scores the importance of each feature in the first feature set according to its own feature importance scorer, and then selects the feature with the lowest importance score (i.e., the feature with the lowest importance) from the first feature set and removes it;
[0056] In the above step S2.4, "RFECV algorithm performs 10-fold cross validation, calculates the average F1 score (F1-score) of the 10-fold cross validation and records it as the F1 score of the first round of iteration" is specifically as follows: all the data in the above-mentioned first-level training set are evenly divided into 10 mutually exclusive data blocks (here, "mutually exclusive data blocks" means that any data in the above-mentioned first-level training set can and can only be divided into one data block, and cannot be divided into two or more data blocks), and the data amounts of the 10 data blocks are equal or approximately equal (in actual applications, the data amounts of the 10 data blocks may have very small differences, but they are within the allowable error range. For example, when the total number of samples cannot be divided by the fold number, the data amounts of the 10 data blocks are approximately equal). For example, the data in the above-mentioned first-level training set are divided into A1, A2, A3, A4, A5, A6, A7, A8, A9, and A 10 data blocks, in the first cycle, the RFECV algorithm uses, for example, A2, A3, A4, A5, A6, A7, A8, A9, and A10 data blocks to train the gradient boosting classifier algorithm, and uses the A1 data block to calculate the F1 score of the first cycle and record the F1 score of the first cycle; in the second cycle, the RFECV algorithm uses, for example, A1, A3, A4, A5, A6, A7, A8, A9, and A10 data blocks to train the gradient boosting classifier algorithm, and uses the A2 data block to calculate the F1 score of the second cycle and record the F1 score of the second cycle; and so on, until the RFECV algorithm completes the 10th cycle, and uses the A10 data block to calculate the F1 score of the 10th cycle and record the F1 score of the 10th cycle; calculate the arithmetic mean of the F1 scores of the 10 cycles of 10-fold cross validation, and record the arithmetic mean as the F1 score of the first round of iteration;
[0057] S2.5: Repeat step S2.4 until the last iteration (e.g., the 28th iteration) is completed and only one feature remains in the current feature set (e.g., the 29th feature set).
[0058] S2.6: The RFECV algorithm sorts the F1 scores from the first iteration to the last iteration in descending order and obtains the maximum F1 score. The iteration that produces the maximum F1 score and the number of features in the corresponding feature set are then determined. The features contained in the feature set that produces the maximum F1 score are used as the key features most relevant to predicting high-frequency hearing loss.
[0059] In this embodiment, 18 key features most relevant to the prediction of high-frequency hearing loss are screened out from the initial feature set. The 18 key features are specifically: province (Province), race (Race), marital status (Marital), gender (Sex), age group (Age_rank), education level (Education), ear examination record (Earcheck), smoking status (Smoke), daily activities (Activities), dairy product intake (Dairy), bean intake (Bean), fruit intake (Fruit), platelet count (PLT), alkaline phosphatase (ALP), total cholesterol (CHO), triglyceride (TG), tinnitus (Tinnitus) and hearing level (Hearing).
[0060] S3: Constructing a high-frequency hearing loss prediction model, using the key features most relevant to the prediction of high-frequency hearing loss (i.e., the 18 key features above) selected in step S2 above to construct a high-frequency hearing loss prediction model; obtaining a secondary training set and a secondary test set;
[0061] In the above step S3, "obtaining a secondary training set and a secondary test set" specifically includes: modifying the above primary training set according to the key features most relevant to the prediction of high-frequency hearing loss obtained in the above step S2 (i.e., the above 18 key features), retaining the relevant data in the primary training set, and deleting the remaining irrelevant data, to obtain a secondary training set; modifying the above primary test set according to the key features most relevant to the prediction of high-frequency hearing loss obtained in the above step S2 (i.e., the above 18 key features), retaining the relevant data in the primary test set, and deleting the remaining irrelevant data, to obtain a secondary test set;
[0062] S4: Training a high-frequency hearing loss prediction model using the key features most relevant to the prediction of high-frequency hearing loss (i.e., the 18 key features above) selected in step S2 and the secondary training set, and training the high-frequency hearing loss prediction model based on the XGBoost model;
[0063] S5: Use the secondary test set to evaluate the high-frequency hearing loss prediction model trained in step S4 above.
[0064] The above step S4 further includes the following sub-steps:
[0065] S4.1: Select the XGBoost model as the basic machine learning model;
[0066] S4.2: Build the hyperparameter grid of the XGBoost model (param_grid);
[0067] S4.3: Initialize the Grid Search Cross Validator (Grid Search CV) and configure the key parameters of the Grid Search Cross Validator;
[0068] S4.4: Train the XGBoost model by traversing all hyperparameter combinations in the XGBoost model's hyperparameter grid (param_grid) using a Grid Search cross validator (Grid Search CV).
[0069] S4.5: Best Parameter Selection: The grid search cross-validator compares the F1 scores of all hyperparameter combinations in the hyperparameter grid of the XGBoost model and selects the hyperparameter combination with the highest F1 score as the best hyperparameter combination and stores it in the grid_search.best_params_ attribute.
[0070] S4.6: Best Model Training. The grid search cross-validator reconfigures a new XGBoost model using the best hyperparameter combination obtained in step S4.5 above, and then trains the new XGBoost model using the above secondary training set (grid_search.fit(X_train_selected,y_train)). The best XGBoost model is obtained and stored (stored in the grid_search.best_estimator_ attribute and assigned to the best_model variable, i.e., best_model = grid_search.best_estimator_).
[0071] The above step S4.2 further includes the following sub-steps:
[0072] S4.2.1: Define (user-defined) the hyperparameter grid (param_grid) of the XGBoost model, including the number of trees (n_estimators), the learning rate (learning_rate), the maximum tree depth (max_depth), and the minimum number of samples required to split an internal node (min_child_weight). Set two candidate values (user-defined) for each hyperparameter.
[0073] In this embodiment, for example, the candidate values of the number of trees (n_estimators) are set to 100 and 200, the candidate values of the learning rate (learning_rate) are set to 0.01 and 0.1, the candidate values of the maximum depth of the tree (max_depth) are set to 3 and 5, and the candidate values of the minimum number of samples required to split an internal node (min_child_weight) are set to 2 and 5;
[0074] S4.2.2: List all possible combinations of the four hyperparameters in the hyperparameter grid of the XGBoost model based on the candidate values.
[0075] Table 1 Combinations of four hyperparameters of the hyperparameter grid of the XGBoost model
[0076]
[0077] In this embodiment, for example, based on the candidate values, the combinations of four hyperparameters of the hyperparameter grid of the XGBoost model defined in step S4.2.1 above (the candidate values of each combination are arranged in the order of the number of trees, learning rate, maximum depth of the tree, and minimum number of samples required to split the internal nodes) include: G1: {100, 0.01, 3, 2}, G2: {100, 0.01, 3, 5}, G3: {100, 0.01, 5, 2}, G4: {100, 0.01, 5, 5}, G5: {100, 0.1, 3, 2}, G6: {100, 0.1, 3, 5}, G7: {100, 0.1, 5, 2}, G8: {100, 0.1, 5, 5}, G9: {100, 0.1, 5, 2} ,3,5}, G15: {200,0.1,5,2}, G16: {200,0.1,5,5}, a total of 16 combinations, of which G1, G2, G3, G4, G5, G6, G7, G8, G9, G10, G11, G12, G13, G14, G15, and G16 are the numbers of the combinations of the four hyperparameters of the hyperparameter grid of the XGBoost model, as listed in Table 1.
[0078] The above step S4.3 further includes the following sub-steps:
[0079] S4.3.1: Specify the estimator used by the high-frequency hearing loss prediction model as the XGBoost model (xgb) (estimator = xgb) in step S4.1 above;
[0080] S4.3.2: The Grid Search Cross Validator (Grid Search CV) uses and obtains the hyperparameter grid (param_grid) of the XGBoost model constructed in step S4.2 above (i.e., the hyperparameter grid of the XGBoost model constructed in step S4.2 above is passed into the Grid Search Cross Validator, param_grid = param_grid);
[0081] S4.3.3: Set the cross-validation fold to 10-fold cross-validation (cv=10);
[0082] S4.3.4: Specify the F1-score as the cross-validation evaluation metric (scoring = 'f1').
[0083] S4.3.5: Set parallel computing (n_jobs = -1, -1 means using all available CPU cores) to use all available CPU cores to speed up the grid search process.
[0084] The above step S4.4 further includes the following sub-steps:
[0085] S4.4.1: For any hyperparameter combination in the XGBoost model's hyperparameter grid (param_grid) (e.g., hyperparameter combination numbered G1), perform a 10-fold cross-validation on the grid search cross-validator and obtain the corresponding F1 score for the current hyperparameter combination (e.g., hyperparameter combination numbered G1) and record it.
[0086] S4.4.2: Repeat the above step S4.4.1 until the grid search cross validator obtains the F1 score corresponding to each hyperparameter combination in the hyperparameter grid of the XGBoost model and records it.
[0087] The above step S4.4.1 further includes the following content: for any hyperparameter combination in the hyperparameter grid (param_grid) of the XGBoost model (for example, the hyperparameter combination numbered G1), all the data in the above secondary training set are evenly divided into 10 mutually exclusive data blocks (here, "mutually exclusive data blocks" means that any data in the above secondary training set can and can only be divided into one data block, and cannot be divided into two or more data blocks), and the data amounts of the 10 data blocks are equal or approximately equal (in actual applications, the data amounts of the 10 data blocks may have very slight differences, but they are within the allowable error range. For example, when the total number of samples cannot be divided by the fold, the data amounts of the 10 data blocks are approximately equal). For example, the data in the above secondary training set are divided into B1, B2, B3, B4, B5, B6, B7, B8, B9, and B10 data blocks. In the first cycle, the grid search cross validator uses, for example, B2, B3, B4, B5, B6, The B7, B8, B9, and B10 data blocks train the XGBoost model configured with the current hyperparameter combination (e.g., the hyperparameter combination numbered G1) (internal training), and use the B1 data block to calculate the F1 score of the first cycle and record the F1 score of the first cycle (internal evaluation); in the second cycle, for example, the B1, B3, B4, B5, B6, B7, B8, B9, and B10 data blocks are used to train the XGBoost model configured with the current hyperparameter combination (e.g., the hyperparameter combination numbered G1) oost model (internal training), and use the B2 data block to calculate the F1 score of the second cycle and record the F1 score of the second cycle (internal evaluation); and so on, until the grid search cross validator completes the 10th cycle, and uses the B10 data block to calculate the F1 score of the 10th cycle and record the F1 score of the 10th cycle; calculate the arithmetic mean of the F1 scores of the 10-fold cross validation cycles, and use the arithmetic mean as the F1 score corresponding to the current hyperparameter combination and record it.
[0088] In step S5, the high-frequency hearing loss prediction model trained in step S4 is evaluated using the secondary test set, and the following evaluation results are obtained:
[0089] Accuracy: 80.17%. Accuracy indicates the proportion of correct predictions made by the model. An accuracy of 80.17% means that the model correctly predicts whether approximately 80 out of every 100 people have high-frequency hearing loss.
[0090] Precision: 81.59%. Precision indicates the probability that the model is correct when it predicts someone has high-frequency hearing loss. A precision of 81.59% means that when the model predicts someone has high-frequency hearing loss, the prediction is correct 81.59% of the time.
[0091] Recall: 90.25%. Recall indicates the percentage of people the model can identify who actually has high-frequency hearing loss. A recall of 90.25% means that the model can identify 90.25% of people who actually have high-frequency hearing loss.
[0092] F1 score: 85.70%. The F1 score is a metric that takes into account both precision and recall. A higher F1 score indicates better overall model performance.
[0093] Area Under the Circumference (AUC): 86.04%. The AUC value indicates the model's ability to distinguish between people with and without high-frequency hearing loss. This is a measure of the model's discriminatory power, indicating how well the model can distinguish between people with and without high-frequency hearing loss. The closer the AUC value is to 100%, the stronger the model's discriminatory power.
[0094] Application scenarios: The high-frequency hearing loss prediction model provided in this embodiment can be used for large-scale health data analysis in medical institutions, helping doctors quickly screen people at high risk of high-frequency hearing loss. It can also be applied to health management platforms to provide users with personalized health advice.
[0095] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications based on the above description are possible. It is not possible to enumerate all embodiments here. Any obvious variations or modifications arising from the technical solution of the present invention remain within the scope of protection of the present invention.
Claims
1. A method for predicting high-frequency hearing loss, characterized in that: The prediction method includes the following steps: S1: Data preparation: collect multi-dimensional data related to human hearing health as raw data, pre-process the raw data to obtain basic data, and divide the basic data into two parts: a first-level training set and a first-level test set; S2: Construct an initial feature set and use the RFECV algorithm for feature selection to screen out several key features that are most relevant to predicting high-frequency hearing loss from the initial feature set; S3: constructing a high-frequency hearing loss prediction model, using the key features most relevant to the prediction of high-frequency hearing loss screened in step S2 to construct the high-frequency hearing loss prediction model; obtaining a secondary training set and a secondary test set; S4: training a high-frequency hearing loss prediction model, using the key features and secondary training set most relevant to the prediction of high-frequency hearing loss screened out in step S2, and training the high-frequency hearing loss prediction model based on an XGBoost model; S5: Using the secondary test set, the high-frequency hearing loss prediction model trained in step S4 is evaluated.
2. The high-frequency hearing loss prediction method according to claim 1, wherein: In step S1, the preprocessing of the original data includes filling missing values, data transformation and / or standardization; the specific method of dividing the basic data into two parts, a first-level training set and a first-level test set, is as follows: taking each tested person as the basic unit of division, each tested person corresponds to a piece of data, and the piece of data corresponding to each tested person is determined to enter the first-level training set or the first-level test set by random selection, and the ratio of the data volume of the first-level training set to the first-level test set is 8:2; the piece of data corresponding to each tested person includes all indicator data associated with the tested person.
3. The high-frequency hearing loss prediction method according to claim 1, wherein: The step S2 further includes the following sub-steps: S2.1: An initial feature set is artificially constructed based on previous experience. The initial feature set includes several features related to the prediction of high-frequency hearing loss. S2.2: The RFECV algorithm uses the gradient boosting classifier algorithm as its evaluator; S2.3: The gradient boosting classifier algorithm is trained using all features of the initial feature set. The gradient boosting classifier algorithm selects the least important feature from the initial feature set and removes it, thereby completing the zeroth iteration and obtaining the first feature set. S2.4: The gradient boosting classifier is trained using all features of the first feature set. The RFECV algorithm performs 10-fold cross-validation and calculates the average F1 score of the 10-fold cross-validation. This is recorded as the F1 score for the first iteration. The gradient boosting classifier then selects the least important feature from the first feature set and removes it, completing the first iteration and obtaining the second feature set. S2.5: Repeat step S2.4 until the last iteration is completed and only one feature remains in the current feature set; S2.6: The RFECV algorithm sorts the F1 scores from the first iteration to the last iteration in descending order and obtains the maximum F1 score. It then obtains the iteration that produces the maximum F1 score and the number of features in the corresponding feature set. The features contained in the feature set that produces the maximum F1 score are used as the key features most relevant to predicting high-frequency hearing loss.
4. The high-frequency hearing loss prediction method according to claim 3, wherein: In step S2.3, "the gradient boosting classifier algorithm selects and removes the feature with the lowest importance from the initial feature set" specifically includes: the gradient boosting classifier algorithm scores the importance of each feature in the initial feature set according to its own feature importance scorer, and then selects and removes the feature with the lowest importance score from the initial feature set; In step S2.4, "the gradient boosting classifier algorithm selects and removes the feature with the lowest importance from the first feature set" specifically includes: the gradient boosting classifier algorithm scores the importance of each feature in the first feature set according to its own feature importance scorer, and then selects and removes the feature with the lowest importance score from the first feature set; In step S2.4, "the RFECV algorithm performs 10-fold cross validation, calculates the average of the F1 scores of the 10-fold cross validation and records it as the F1 score of the first round of iteration" is specifically as follows: all the data in the first-level training set are evenly divided into 10 mutually exclusive data blocks, and the data amounts of the 10 data blocks are equal or approximately equal. Specifically, the data in the above-mentioned first-level training set are evenly divided into A1, A2, A3, A4, A5, A6, A7, A8, A9, and A10 data blocks. In the first cycle, the RFECV algorithm uses the A2, A3, A4, A5, A6, A7, A8, A9, and A10 data blocks to train the gradient boosting classifier algorithm, and uses The A1 data block calculates the F1 score of the first cycle and records the F1 score of the first cycle; in the second cycle, the RFECV algorithm uses the A1, A3, A4, A5, A6, A7, A8, A9, and A10 data blocks to train the gradient boosting classifier algorithm, and uses the A2 data block to calculate the F1 score of the second cycle and record the F1 score of the second cycle; and so on, until the RFECV algorithm completes the 10th cycle, and uses the A10 data block to calculate the F1 score of the 10th cycle and record the F1 score of the 10th cycle; calculate the arithmetic mean of the F1 scores of the 10 cycles of 10-fold cross validation, and record the arithmetic mean as the F1 score of the first round of iteration.
5. The high-frequency hearing loss prediction method according to claim 3, wherein: In step S2.6, 18 key features most relevant to predicting high-frequency hearing loss are screened out from the initial feature set. The 18 key features are specifically: province, race, marital status, gender, age group, education level, ear examination record, smoking status, daily activities, dairy product intake, bean intake, fruit intake, platelet count, alkaline phosphatase, total cholesterol, triglycerides, tinnitus and hearing level.
6. The high-frequency hearing loss prediction method according to claim 1, wherein: In step S3, "obtaining a secondary training set and a secondary test set" specifically includes: modifying the primary training set based on the key features most relevant to the prediction of high-frequency hearing loss obtained in step S2, retaining relevant data in the primary training set, and deleting remaining irrelevant data, to obtain a secondary training set; and modifying the primary test set based on the key features most relevant to the prediction of high-frequency hearing loss obtained in step S2, retaining relevant data in the primary test set, and deleting remaining irrelevant data, to obtain a secondary test set.
7. The high-frequency hearing loss prediction method according to claim 1, wherein: The step S4 further includes the following sub-steps: S4.1: Select the XGBoost model as the basic machine learning model; S4.2: Build the hyperparameter grid for the XGBoost model; S4.3: Initialize the grid search cross validator and configure the key parameters of the grid search cross validator; S4.4: Train the XGBoost model by traversing all hyperparameter combinations in the XGBoost model's hyperparameter grid using a grid search cross-validator. S4.5: Select the best hyperparameters. The grid search cross-validator compares the F1 scores of all hyperparameter combinations in the hyperparameter grid of the XGBoost model, selects the hyperparameter combination with the highest F1 score as the best hyperparameter combination, and stores it. S4.6: Best model training: The grid search cross validator uses the best hyperparameter combination obtained in step S4.5 to reconfigure a new XGBoost model, and then uses the secondary training set to train the new XGBoost model to obtain the best XGBoost model and store it.
8. The high-frequency hearing loss prediction method according to claim 7, wherein: The step S4.2 further includes the following sub-steps: S4.2.1: Define the hyperparameter grid for the XGBoost model, including the number of trees, the learning rate, the maximum tree depth, and the minimum number of samples required to split an internal node. Set two candidate values for each hyperparameter. S4.2.2: List all possible combinations of the four hyperparameters in the hyperparameter grid of the XGBoost model based on the candidate values. In step S4.2.1, the candidate values of the number of trees are set to 100 and 200, the candidate values of the learning rate are set to 0.01 and 0.1, the candidate values of the maximum depth of the tree are set to 3 and 5, and the candidate values of the minimum number of samples required to split internal nodes are set to 2 and 5.
9. The high-frequency hearing loss prediction method according to claim 7, wherein: The step S4.3 further includes the following sub-steps: S4.3.1: Specify the estimator used by the high-frequency hearing loss prediction model as the XGBoost model in step S4.1; S4.3.2: A grid search cross validator uses and obtains the hyperparameter grid of the XGBoost model constructed in step S4.2; S4.3.3: Set the number of cross-validation folds to 10; S4.3.4: Specify the F1 score as the evaluation metric for cross-validation. S4.3.5: Set up parallel computing to use all available CPU cores to speed up the grid search process.
10. The high-frequency hearing loss prediction method according to claim 7, wherein: The step S4.4 further includes the following sub-steps: S4.4.1: For any hyperparameter combination in the hyperparameter grid of the XGBoost model, perform a 10-fold cross-validation with the grid search cross-validator and record the F1 score corresponding to the current hyperparameter combination. S4.4.2: Repeat step S4.4.1 until the grid search cross validator obtains the corresponding F1 score for each hyperparameter combination in the hyperparameter grid of the XGBoost model and records it; The step S4.4.1 further includes the following contents: for any hyperparameter combination in the hyperparameter grid of the XGBoost model, all the data in the secondary training set are evenly divided into 10 mutually exclusive data blocks, and the data amounts of the 10 data blocks are equal or approximately equal. Specifically, the data in the secondary training set are divided into B1, B2, B3, B4, B5, B6, B7, B8, B9, and B10 data blocks. In the first cycle, the grid search cross validator uses B2, B3, B4, B5, B6, B7, B8, B9, and B10 data blocks to train the XGBoost model configured with the current hyperparameter combination, and uses the B1 data block to calculate the first The F1 score of the first cycle is calculated and the F1 score of the first cycle is recorded; in the second cycle, the B1, B3, B4, B5, B6, B7, B8, B9, and B10 data blocks are used to train the XGBoost model configured with the current hyperparameter combination, and the B2 data block is used to calculate the F1 score of the second cycle and the F1 score of the second cycle is recorded; and so on, until the grid search cross validator completes the 10th cycle, and the B10 data block is used to calculate the F1 score of the 10th cycle and the F1 score of the 10th cycle is recorded; the arithmetic mean of the F1 scores of the 10-fold cross validation cycles is calculated, and the arithmetic mean is used as the F1 score corresponding to the current hyperparameter combination and is recorded.
Citation Information
Cited By
Hearing impairment recovery condition evaluation method based on artificial intelligence
CN121366732A