Image data classification method, apparatus, and storage medium
Patent Information
- Application Number
- CN202510344862.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2026-09-22
AI Technical Summary
[0005]本申请实施例提供了一种影像数据分类方法、装置和存储介质,该影像数据分类方法、装置和存储介质,能够自动选取最优的基学习器组合与最优超参数,解决了已有研究中最优的基学习器组合和超参数难以确定的问题,提高了影像数据分类的准确性和可靠性
[0018]本申请实施例通过对原始影像数据进行预处理,并提取影像组学特征,并对影像组学特征进行筛选后,根据分类器模型以及筛选出的影像组学特征对影像数据进行分类。由于所述分类器模型包括能够自动选取最优的基学习器组合与最优超参数的自适应集成学习模型,因此提高了影像数据分类的准确性和可靠性,降低了计算成本和时间成本。
Smart Images

Figure CN122799147A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of medical image processing and machine learning technology, and more specifically, to an image data classification method, apparatus, and storage medium. Background Technology
[0002] Tumor subtyping is a crucial step in clinical diagnosis and treatment planning. Traditional tumor subtyping relies primarily on histopathological examination; however, this method is invasive and complex. In recent years, with the rapid development of medical imaging technology, radiomics methods have gradually become an important tool for tumor subtyping.
[0003] Radiomics methods first extract quantitative image features from medical image data, then combine them with machine learning algorithms for disease classification and prediction, providing crucial support for clinical decision-making. However, in practical applications, the extraction and classification of image features are often influenced by various factors, such as image quality, disease type, and individual differences, leading to unstable model performance. Furthermore, single learners are prone to overfitting or underfitting when faced with complex and variable medical image data, limiting their clinical application value. Therefore, effectively integrating multiple learners has become a current research hotspot in the field of medical image analysis.
[0004] While ensemble learning methods can combine the advantages of multiple base learners to improve classification performance, their performance is often affected by the combination of base learners and the choice of hyperparameters. Traditional methods such as grid search and random search, although feasible, suffer from drawbacks such as high computational cost and low efficiency, and it is difficult to guarantee finding the global optimum. Summary of the Invention
[0005] This application provides an image data classification method, apparatus, and storage medium that can automatically select the optimal combination of base learners and the optimal hyperparameters, solving the problem that the optimal combination of base learners and hyperparameters is difficult to determine in existing studies, and improving the accuracy and reliability of image data classification.
[0006] This application provides an image data classification method, including: Acquire and preprocess raw image data; Extracting radiomics features from preprocessed image data and filtering the radiomics features; Image data is classified based on a classifier model and selected radiomics features; wherein, the classifier model is configured to classify based on radiomics features; the classifier model is a model established based on the optimal combination of base learners, the optimal hyperparameters, and a preset learning algorithm; the optimal combination of base learners and the optimal hyperparameters are obtained based on an adaptive ensemble learning model; the adaptive ensemble learning model is configured to automatically select the optimal combination of base learners and the optimal hyperparameters from the candidate base learner set and the hyperparameter space.
[0007] In one exemplary embodiment, the classifier model is constructed as follows: The classifier model is established as follows: A first initial classifier model is established based on the optimal combination of base learners, the optimal hyperparameters, and the preset learning algorithm. The first initial classifier model is trained and evaluated using the first training set and the first test set, respectively; The first initial classifier model that has completed training is used as the classifier model.
[0008] In one exemplary embodiment, the adaptive ensemble learning model is constructed in the following manner: Multiple initial candidate points are determined, each initial candidate point being a combination of base learners and hyperparameter settings randomly selected from the candidate base learner set and hyperparameter space; the objective function value of each initial candidate point is calculated based on the objective function and each initial candidate point, the objective function being used to evaluate the classification performance under different combinations of base learners and hyperparameter settings; A surrogate model is trained based on multiple initial candidate points and the objective function value of each initial candidate point. The trained surrogate model is then used to predict candidate points in the candidate space other than the initial candidate points, or a predetermined number of randomly sampled candidate points. Based on the predicted objective function values of the candidate points other than the initial candidate points or the predetermined number of randomly sampled candidate points, and a collection function, the next candidate point to be evaluated is selected and added to the existing candidate points for surrogate model training. The surrogate model is iteratively updated. When the iteration ends, the base learner combination and hyperparameter settings that result in the highest predicted objective function value are selected as the optimal base learner combination and optimal hyperparameters. The collection function is used to obtain the expected improvement amount of the next candidate point to be evaluated based on the predicted objective function value of the candidate point and the currently known optimal objective function value.
[0009] In one exemplary embodiment, the method further includes: The range of the hyperparameter space is adjusted based on the predicted objective function value and prediction variance of candidate points in the candidate space according to the trained surrogate model. The surrogate model is retrained based on the candidate points after adjusting the hyperparameter space.
[0010] In one exemplary embodiment, adjusting the range of the hyperparameter space based on the predicted objective function value and prediction variance of candidate points in the candidate space using the trained surrogate model includes: The hyperparameter space includes multiple regions; If there exists a region in the current hyperparameter space where the average value of the predicted objective function of the surrogate model is higher than a first preset threshold in other regions, the range of the hyperparameter space is narrowed to that region. If there is no region in the current hyperparameter space where the average value of the predicted objective function of the surrogate model is higher than a first preset threshold in other regions, or if the average value of the predicted variance of the current hyperparameter space calculated according to the surrogate model exceeds a second preset threshold, the range of the current hyperparameter space will be expanded.
[0011] In one exemplary embodiment, calculating the objective function value for each initial candidate point based on the objective function and each initial candidate point includes: A second initial classifier model is built based on each initial candidate point and a preset learning algorithm; The second initial classifier model is trained and evaluated using the first training set and the first test set, respectively; Use the trained second initial classifier model as the second classifier model; The objective function is used to evaluate the second classifier model to obtain the objective function value corresponding to each initial candidate point.
[0012] In one exemplary embodiment, selecting the next candidate point to be evaluated based on the predicted objective function value and acquisition function of candidate points in the candidate space other than the initial candidate point includes: The candidate point that maximizes the output value of the acquisition function will be selected as the next candidate point to be evaluated.
[0013] In one exemplary embodiment, the candidate base learner set includes one or more of the following algorithms: support vector machine, random forest, decision tree; The hyperparameter space includes one or more of the following: the hyperparameter space of support vector machines, the hyperparameter space of random forests, and the hyperparameter space of decision trees.
[0014] In one exemplary embodiment, the objective function is a classification performance evaluation index function; The proxy model is a Gaussian process model; The acquisition function is the function to be improved.
[0015] In one exemplary embodiment, the iteration termination condition includes one of the following: the number of iterations equals a preset number of iterations; the difference in prediction results of multiple consecutive iterations of the surrogate model is less than a third preset threshold.
[0016] This application also provides an image data classification device, including a memory and a processor. The memory is used to store the program for the image data classification method; The processor is configured to read the program that executes the image data classification method and execute the method described in any of the above embodiments.
[0017] This application also provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to cause the computer to perform the method described in any of the above embodiments.
[0018] This application embodiment preprocesses the raw image data, extracts radiomics features, and filters these features before classifying the image data based on a classifier model and the selected radiomics features. Since the classifier model includes an adaptive ensemble learning model that automatically selects the optimal combination of base learners and optimal hyperparameters, the accuracy and reliability of image data classification are improved, while reducing computational and time costs.
[0019] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the description and the accompanying drawings. Attached Figure Description
[0020] The accompanying drawings are used to provide a further understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.
[0021] Figure 1 This is a schematic diagram of an image data classification method according to an embodiment of this application; Figure 2 This is a schematic diagram of another image data classification method according to an embodiment of this application; Figure 3 This is a schematic diagram of an adaptive ensemble learning model according to an embodiment of this application; Figure 4 This is a schematic diagram of another adaptive ensemble learning model according to an embodiment of this application; Figure 5 This is a schematic diagram of an image data classification device according to an embodiment of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in detail below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be arbitrarily combined with each other.
[0023] Figure 1 This is a schematic diagram of an image data classification method according to an embodiment of this application. The image data classification method is as follows: Figure 1 As shown, it includes the following steps 11 to 13: Step 11: Acquire and preprocess the raw image data; Step 12: Extract radiomics features from the preprocessed image data and filter the radiomics features; Step 13: Classify the image data according to the classifier model and the selected radiomics features; The classifier model is configured to classify image data based on radiomics features; the classifier model is a model established based on the optimal combination of base learners, the optimal hyperparameters, and a preset learning algorithm; the optimal combination of base learners and the optimal hyperparameters are obtained based on an adaptive ensemble learning model; the adaptive ensemble learning model is configured to automatically select the optimal combination of base learners and the optimal hyperparameters from the candidate base learner set and the hyperparameter space.
[0024] This application embodiment preprocesses the raw image data, extracts radiomics features, and filters these features before classifying the data based on a classifier model and the selected features. The classifier model is built using the optimal combination of base learners, optimal hyperparameters, and a preset learning algorithm; while the adaptive ensemble learning model can automatically select the optimal combination of base learners and optimal hyperparameters from the candidate base learner set and hyperparameter space. Therefore, the accuracy and reliability of image data classification are improved, while computational and time costs are reduced.
[0025] In step 11, for example, preprocessing the raw image data may include operations such as format standardization, noise reduction, and normalization to ensure the consistency and comparability of the raw image data. It may also include a professional radiologist delineating the lesion area layer by layer; the lesion area is the region of interest (ROI). The classification of the raw image data is known.
[0026] In step 12, radiomics features can be extracted from the preprocessed raw image data using radiomics methods. Alternatively, wavelet transform can be applied to the preprocessed image data to extract radiomics features from the wavelet-transformed image data. Furthermore, different LoG (also known as Laplacian of Gaussian) filters can be applied to the preprocessed image data to extract radiomics features from the filtered image data. Radiomics features can include first-order features, morphological features, and texture features.
[0027] Screening radiomics features can include using statistical methods such as Pearson correlation coefficient, Least Absolute Convergence and Selection Operator (LASSO) to select the most representative subset of features.
[0028] In step 13, in one exemplary embodiment, the classifier model can be built as follows: A first initial classifier model is established based on the optimal combination of base learners, the optimal hyperparameters, and the preset learning algorithm. The first initial classifier model is trained and evaluated using the first training set and the first test set, respectively; The first initial classifier model that has completed training is used as the first classifier model.
[0029] For example, the labeled image data can be divided into two parts, one part as the first training set and the other part as the first test set.
[0030] For example, the preset learning algorithm can be a bagging ensemble learning algorithm. In some other embodiments, other learning methods may also be used as the preset learning algorithm.
[0031] In one exemplary embodiment, constructing the adaptive ensemble learning model may include: Multiple initial candidate points are determined, each initial candidate point being a combination of base learners and hyperparameter settings randomly selected from the candidate base learner set and hyperparameter space; the objective function value of each initial candidate point is calculated based on the objective function and each initial candidate point, the objective function being used to evaluate the classification performance under different combinations of base learners and hyperparameter settings; A surrogate model is trained based on multiple initial candidate points and the objective function value of each initial candidate point. The trained surrogate model is then used to predict candidate points in the candidate space other than the initial candidate points, or a predetermined number of randomly sampled candidate points. Based on the predicted objective function values of the candidate points other than the initial candidate points or the predetermined number of randomly sampled candidate points, and a collection function, the next candidate point to be evaluated is selected and added to the existing candidate points for surrogate model training. The surrogate model is iteratively updated. When the iteration ends, the base learner combination and hyperparameter settings that result in the highest predicted objective function value are selected as the optimal base learner combination and optimal hyperparameters. The collection function is used to obtain the expected improvement amount of the next candidate point to be evaluated based on the predicted objective function value of the candidate point and the currently known optimal objective function value.
[0032] The objective function can be a classification performance evaluation metric (AUC) function. The candidate base learner set can include one or more of the following: Support Vector Machine (SVM), Random Forest (RF), Decision Tree (DT), Logistic Regression (LR), etc.
[0033] The hyperparameter space may include one or more of the following: the hyperparameter space of support vector machines, the hyperparameter space of random forests, and the hyperparameter space of decision trees.
[0034] For example, the surrogate model is a Gaussian process model. In some embodiments, the surrogate model may also be a random forest model or a neural network model.
[0035] In one exemplary embodiment, calculating the objective function value for each initial candidate point based on the objective function and each initial candidate point includes: A second initial classifier model is built based on each initial candidate point and a preset learning algorithm; The second initial classifier model is trained and evaluated using the first training set and the first test set, respectively; Use the trained second initial classifier model as the second classifier model; The objective function is used to evaluate the second classifier model to obtain the objective function value corresponding to each initial candidate point.
[0036] In one exemplary embodiment, the preset conditions include: the number of iterations equals a preset number of iterations; or the difference in prediction results of multiple consecutive iterations of the surrogate model is less than a preset threshold.
[0037] In one exemplary embodiment, the acquisition function is a desired improvement function. In other embodiments, the acquisition function may also be an upper confidence bound (UCB), probability of improvement (PI), or entropy search (ES) function.
[0038] In one exemplary embodiment, the method further includes: The range of the hyperparameter space is adjusted based on the predicted objective function value and prediction variance of the candidate points in the candidate space using the trained surrogate model. The surrogate model is retrained based on the candidate points after adjusting the hyperparameter space.
[0039] By adjusting the range of the hyperparameter space, the efficiency and accuracy of iterative search can be improved.
[0040] In one exemplary embodiment, adjusting the range of the hyperparameter space based on the predicted objective function value and prediction variance of candidate points in the candidate space using the trained surrogate model includes: The hyperparameter space includes multiple regions; If there exists a region in the current hyperparameter space where the average value of the predicted objective function of the surrogate model is higher than a first preset threshold in other regions, the range of the hyperparameter space is narrowed to that region. If there is no region in the current hyperparameter space where the average value of the predicted objective function of the surrogate model is higher than a first preset threshold in other regions, or if the average value of the predicted variance of the current hyperparameter space calculated according to the surrogate model exceeds a second preset threshold, the range of the current hyperparameter space will be expanded.
[0041] That is, if the surrogate model predicts a high objective function value in a certain region of the current hyperparameter space, the search range, i.e., the hyperparameter space, is narrowed to this region; if the surrogate model fails to find a region in the current hyperparameter space where the prediction result is significantly improved, or if the prediction result fluctuates greatly, the search range is expanded.
[0042] The following section uses the classification of pituitary tumors as an example to provide a detailed explanation of this application.
[0043] Different subtypes of pituitary adenomas (such as PIT1, SF1, and TPIT) require different treatment approaches. Accurate classification of the pituitary adenoma subtype before treatment can provide patients with a more precise treatment plan. This application aims to determine the type of pituitary adenoma based on the patient's preoperative imaging examinations, thereby assisting physicians in developing an initial treatment plan.
[0044] A specific implementation plan for classifying pituitary tumor subtypes based on the classifier model of this application (e.g.) Figure 2 (As shown), including the following steps 21 to 25: Step 21, MRI image data acquisition and preprocessing; Step 22, Image data feature extraction and feature selection; Step 23, Construction of the Adaptive Ensemble Learning Model; Step 24, classifier model construction; Step 25, Application of the classifier model.
[0045] In step 21, preoperative image data of N pituitary tumor patients are acquired using magnetic resonance imaging (MRI), which may include sequences such as T1WI, CE-T1WI, and T2WI. These N pituitary tumor patients have been classified into two subtypes by immunohistochemical staining, with the number of patients in each subtype being approximately equal. Then, the raw image data undergoes preprocessing such as image format standardization, noise reduction, and normalization. A professional radiologist then uses image annotation tools to delineate the pituitary tumor lesion area layer by layer, extracting the region of interest, i.e., the tumor region.
[0046] In step 22, radiomics features can be extracted from the preprocessed image using PyRadiomics software. Alternatively, wavelet transform can be performed on the preprocessed image data to extract radiomics features from the wavelet-transformed image data. Different LoG (LoG, also known as Laplacian of Gaussian) filters can also be applied to the preprocessed image data to extract radiomics features from the filtered image data. The extracted features include first-order features, morphological features, and texture features. First-order features include, for example, histogram statistics. Morphological features include, for example, tumor area and surface area. Texture features include, for example, gray-level co-occurrence matrix, gray-level dependency matrix, and gray-level run matrix.
[0047] Feature selection can include calculating Pearson correlation coefficients to eliminate redundant features, using one-way ANOVA to select statistically significant features, and then further filtering using the LASSO algorithm to select the most representative features. This avoids information redundancy and overfitting caused by excessively high dimensionality of radiomics features. Redundant features can be excluded by setting a threshold, such as 0.85.
[0048] In step 23, as Figure 3 As shown, the construction of an adaptive ensemble learning model may include the following steps 231 to 234: Step 231: Define the candidate space and construct the proxy model.
[0049] Defining the candidate space includes defining the objective function, defining candidate basis learners, and defining the hyperparameter space.
[0050] This application employs a bagging ensemble learning method for sample classification. The classification performance evaluation metric AUC can be used as the objective function to evaluate the classification performance under different combinations of base learners and hyperparameter settings. The candidate base learners include various machine learning algorithms, such as Support Vector Machine (SVM), Random Forest (RF), and Gradient Boosting Decision Tree (GBDT). Simultaneously, a hyperparameter space is defined for each base learner, including the range of values and their types. For example, the hyperparameter space of SVM includes the penalty parameter C (0.1, 1, 10, 100), the kernel function parameter gamma (1e-3, 1e-2, 1e-1, 1, 10, 100), and the kernel function type (e.g., linear kernel, RBF kernel). The hyperparameter space of RF includes the number of trees n_estimators (10, 50, 100) and the maximum depth max_depth (None, 10, 20, 30), etc.
[0051] Gaussian processes are used as surrogate models.
[0052] The basic form of the Gaussian process model is as follows: f(x)~GP(μ(x), k(x,x′)); Here, x represents candidate points in the candidate space, i.e., different machine learning model configurations (including base learner combinations and hyperparameter settings). f(x) is the output of the Gaussian process, representing the predicted objective function value of the input x obtained through the Gaussian process. μ(x) is the mean function, usually assumed to be 0 or a constant; k(x,x′) is the kernel function, used to measure the correlation between the outputs corresponding to two input points x and x′. Commonly used kernel functions include the RBF kernel (Radial Basis Function kernel), the Matern kernel, etc.
[0053] Step 232: Train the surrogate model using the initial sample data.
[0054] The steps for training the surrogate model using initial sample data are as follows: 1) From the candidate space, randomly select 10 candidate points. Each candidate point includes: a specific combination of base learners and hyperparameter settings (used to determine the configuration of the classifier model). Then calculate the true objective function value (AUC) of the initial classifier model under each candidate point configuration. The true objective function value is obtained by training the initial classifier model on the training set (e.g., the first training set mentioned above) and evaluating it on the test set (e.g., the first test set mentioned above). Use these 10 candidate points and their corresponding true objective function values (AUC) as the initial sample data for the surrogate model. 2) Select the kernel function of the Gaussian process model, and use the maximum likelihood function and the selected initial sample data to estimate the hyperparameters of the Gaussian process to construct the Gaussian process model.
[0055] This step involves optimizing the likelihood function, with the goal of finding hyperparameter values that maximize the likelihood function.
[0056] 3) After determining the hyperparameters of the Gaussian process, the Gaussian process model can be used to predict new input points.
[0057] During prediction, the existing sample data and estimated hyperparameters are used to calculate the predicted mean and predicted variance of the new input point, and the predicted objective function value of the new input point is obtained.
[0058] For new input points The Gaussian process uses the following formulas to calculate its prediction mean and prediction variance: ; ; Where K is the kernel function mentioned above, X=[x1,x2,...,xn] is the existing training sample data, and y=[y1,y2,...yn] is the objective function value corresponding to X. (Corresponding to the aforementioned μ(x)) is the predicted mean of the new input point. It is the prediction variance of the new input point.
[0059] Step 233: Iteratively search the surrogate model trained on the initial sample data to obtain new candidate points, update the sample data, and retrain the surrogate model using the updated sample data.
[0060] In each iteration, firstly, the next candidate point to be evaluated is selected based on the prediction results of the surrogate model and the acquisition function. Specifically, the surrogate model can be used to predict the remaining candidate points in the candidate space to obtain the predicted objective function values corresponding to the remaining candidate points. Alternatively, a certain number (e.g., 1000) of candidate points can be randomly sampled from the candidate space, and the surrogate model can be used to predict the corresponding predicted objective function values for these candidate points. The candidate point with the largest acquisition function value is selected as the next candidate point to be evaluated. Then, the candidate point and its corresponding true objective function value are added to the sample data to update the sample data, and the surrogate model is retrained using the updated sample data.
[0061] Expected Improvement (EI) can be used as the acquisition function, and its basic form is: ; Here, x represents candidate points in the candidate space, i.e., different machine learning model configurations (including combinations of base learners and hyperparameter settings). f(x) is the predicted objective function value of candidate point x obtained through the surrogate model, and f_best is the currently known optimal objective function value, i.e., the optimal true objective function value corresponding to the sample data. Initially, f_best is the optimal true objective function value corresponding to the 10 initial candidate points. EI(x) measures the expected improvement by obtaining a better solution than the current best solution at candidate point x.
[0062] The true objective function value of the candidate point can be determined as follows: A classifier model is built based on candidate points and the bagging ensemble learning method. Then, the classifier model is trained and evaluated using real data (including image data and corresponding subtypes) to obtain the true objective function value corresponding to the candidate point.
[0063] Step 234 continues until a predetermined number of iterations is reached or the prediction results of the surrogate model no longer improve significantly. The base learner combination and hyperparameter settings that yield the highest prediction objective function value when the preset conditions are met are taken as the optimal base learner combination and optimal hyperparameters.
[0064] The predetermined number of iterations can be set to 100; the surrogate model's prediction no longer improves significantly, meaning that in multiple consecutive iterations, the improvement in the predicted objective function value is less than a specified threshold.
[0065] In step 23, as Figure 4 As shown, the construction of an adaptive ensemble learning model may include the following steps 271 to 274: Step 271: Define the candidate space and construct the proxy model.
[0066] Defining the candidate space includes defining the objective function, defining candidate basis learners, and defining the hyperparameter space.
[0067] This application employs a bagging ensemble learning method for sample classification. The classification performance evaluation metric AUC can be used as the objective function to evaluate the classification performance under different combinations of base learners and hyperparameter settings. The candidate base learners include various machine learning algorithms, such as Support Vector Machine (SVM), Random Forest (RF), and Gradient Boosting Decision Tree (GBDT). Simultaneously, a hyperparameter space is defined for each base learner, including the range of values and their types. For example, the hyperparameter space of SVM includes the penalty parameter C (0.1, 1, 10, 100), the kernel function parameter gamma (1e-3, 1e-2, 1e-1, 1, 10, 100), and the kernel function type (e.g., linear kernel, RBF kernel). The hyperparameter space of RF includes the number of trees n_estimators (10, 50, 100) and the maximum depth max_depth (None, 10, 20, 30), etc.
[0068] Gaussian processes are used as surrogate models.
[0069] The basic form of the Gaussian process model is as follows: f(x)~GP(μ(x), k(x,x′)); Here, x represents candidate points in the candidate space, i.e., different machine learning model configurations (including base learner combinations and hyperparameter settings). f(x) is the output of the Gaussian process, representing the predicted objective function value of the input x obtained through the Gaussian process. μ(x) is the mean function, usually assumed to be 0 or a constant; k(x,x′) is the kernel function, used to measure the correlation between the outputs corresponding to two input points x and x′. Commonly used kernel functions include the RBF kernel (Radial Basis Function kernel), the Matern kernel, etc.
[0070] Step 272: Train the surrogate model using the initial sample data.
[0071] The steps for training the surrogate model using initial sample data are as follows: 1) From the candidate space, randomly select 10 candidate points. Each candidate point includes: a specific combination of base learners and hyperparameter settings (used to determine the configuration of the classifier model). Then, calculate the true objective function value (AUC) of the initial classifier model under each candidate point configuration. The true objective function value is obtained by training the initial classifier model on the training set (e.g., the first training set mentioned above) and evaluating it on the test set (e.g., the first test set mentioned above). Use these 10 candidate points and their corresponding true objective function values (AUC) as the initial sample data for the surrogate model.
[0072] 2) Select the kernel function of the Gaussian process model, use the maximum likelihood function and the selected initial sample data to estimate the hyperparameters of the Gaussian process, and construct the Gaussian process model.
[0073] This step involves optimizing the likelihood function, with the goal of finding hyperparameter values that maximize the likelihood function.
[0074] 3) After determining the hyperparameters of the Gaussian process, the Gaussian process model can be used to predict new input points.
[0075] During prediction, the existing sample data and the estimated hyperparameters of the Gaussian process are used to calculate the prediction mean and prediction variance of the new input point, and the prediction objective function value of the new input point is obtained.
[0076] For new input points The Gaussian process uses the following formulas to calculate its prediction mean and prediction variance: ; ; Where K is the kernel function mentioned above, X=[x1,x2,...,xn] is the existing training sample data, and y=[y1,y2,...yn] is the objective function value corresponding to X. It is the predicted mean of the new input points. It is the prediction variance of the new input point.
[0077] Step 273: Iteratively search the surrogate model trained on the initial sample data to obtain new candidate points, update the sample data, retrain the surrogate model using the updated sample data, and dynamically adjust the range of the hyperparameter space based on the prediction results of the trained surrogate model for the remaining candidate points in the candidate space, so as to search for the optimal solution more efficiently in the next iteration.
[0078] In each iteration, firstly, the next candidate point to be evaluated is selected based on the prediction results of the surrogate model and the acquisition function. Specifically, the surrogate model can be used to predict the remaining candidate points in the candidate space to obtain the predicted objective function values corresponding to the remaining candidate points. Alternatively, a predetermined number of candidate points (e.g., 1000) can be randomly sampled from the candidate space, and the surrogate model can be used to predict the corresponding predicted objective function values for these candidate points. The candidate point with the largest acquisition function value is selected as the next candidate point to be evaluated. Then, the candidate point and its corresponding true objective function value are added to the sample data to update the sample data, and the surrogate model is retrained using the updated sample data.
[0079] Expected Improvement (EI) can be used as the acquisition function, and its basic form is: ; Here, x represents candidate points in the candidate space, i.e., different machine learning model configurations (including combinations of base learners and hyperparameter settings). f(x) is the predicted objective function value of candidate point x obtained through the surrogate model, and f_best is the currently known optimal objective function value, i.e., the optimal true objective function value corresponding to the sample data. Initially, f_best is the optimal true objective function value corresponding to the 10 initial candidate points. EI(x) measures the expected improvement by obtaining a better solution than the current best solution at candidate point x.
[0080] The true objective function value of the candidate point can be determined as follows: A classifier model is built based on candidate points and the bagging ensemble learning method. Then, the classifier model is trained and evaluated using real data (including image data and corresponding subtypes) to obtain the true objective function value corresponding to the candidate point.
[0081] In one exemplary embodiment, dynamically adjusting the range of the hyperparameter space based on the prediction results of the trained surrogate model for the remaining candidate points in the candidate space may include: Based on the prediction results of the surrogate model, the performance distribution and potential optimal solution regions in the current hyperparameter space are analyzed. Then, the range of the hyperparameter space is dynamically adjusted to search for the optimal solution more efficiently in the next iteration. Specifically, if the surrogate model predicts a high objective function value within a certain region of the current hyperparameter space, the search range can be narrowed to this region for a more refined search for the optimal solution. If the surrogate model fails to find a region with significantly improved performance within the current hyperparameter space, or if the prediction results fluctuate greatly, the search range can be expanded to explore more potential high-quality candidate points. The region with a high predicted objective function value in the current hyperparameter space can be determined as follows: By analyzing the prediction results f(x) of the surrogate model, regions with high objective function values can be identified.
[0082] In each iteration, the surrogate model predicts the value of each candidate point, f(x), for all other candidate points in the candidate space or a randomly sampled number of candidate points. This prediction can be visualized across the entire hyperparameter space. If the predicted values f(x) of multiple candidate points within a certain hyperparameter region are significantly higher than those in other regions, then the objective function value of that region is considered to be high. For low-dimensional problems, the distribution of predicted values can be visually observed through visualization (e.g., heatmaps). For high-dimensional problems, dimensionality reduction techniques (e.g., PCA or t-SNE) can be used to map the high-dimensional space to a low-dimensional space before visualization. For example, assuming the hyperparameter space is two-dimensional (e.g., C and γ), a heatmap of the predicted values can be used to visually observe regions with high predicted values. Alternatively, the hyperparameter space can be divided into several sub-regions, and the average predicted value within each sub-region can be calculated. If the average predicted value of a certain region is significantly higher than that of other regions, then the objective function value of that region is considered to be high. This narrows the search scope to the region with the highest predicted objective function value.
[0083] The following method is used to determine whether the prediction results in the current hyperparameter space fluctuate significantly: For example, the volatility of prediction results can be examined through visualization. For instance, by plotting a heatmap of the predicted values, if the colors change drastically, it may indicate that the prediction results for that hyperparameter space are highly volatile.
[0084] For example, a Gaussian process (GP) can be used as a surrogate model, which can not only predict the objective function value f(x) of candidate points, but also provide the uncertainty of the prediction (prediction variance). By analyzing the prediction variance of the surrogate model, the volatility of the prediction results can be determined. The larger the variance, the more uncertain the model's prediction for that region. Specifically: In each iteration, the surrogate model predicts for all remaining candidate points in the candidate space or a randomly sampled number of candidate points, obtaining the prediction variance for each candidate point. The average prediction variance of the current hyperparameter space can be calculated to determine if it is significantly larger than the variance of the hyperparameter space before narrowing. If the current average variance is significantly higher than historical values (statistical tests such as t-tests can be used to compare the variance distribution of the current search range with the historical variance distribution, and the significance of the test results can be used to determine if the fluctuation is large), or if the current average variance exceeds a set threshold, then the fluctuation can be considered large. Then, a decision is made on whether to expand the search range.
[0085] The following example of dynamically adjusting the hyperparameter space illustrates the method of adjusting the hyperparameter space.
[0086] Suppose there is a two-dimensional hyperparameter space with hyperparameters C and γ, and initial search ranges of [0.1, 100] and [1e-3, 1e3], respectively.
[0087] Initial search phase: Randomly select 10 candidate points within the initial search range to construct a Gaussian process model.
[0088] First iteration: A surrogate model is used to predict candidate points in the candidate space. It is found that the predicted values are higher and the variance is lower in the regions where C is in [1, 10] and γ is in [1e-1, 1e1]. The search range is then narrowed down to C in [1, 10] and γ in [1e-1, 1e1].
[0089] Second iteration: Prediction was performed within the narrowed search range. It was found that the predicted values of C in the region [5, 10] and γ in the region [1e0, 1e1] were higher and the variance was lower. The search range was further narrowed to C in the region [5, 10] and γ in the region [1e0, 1e1].
[0090] The third iteration: A search was conducted within the further narrowed search range, revealing significant fluctuations in the prediction results (i.e., high variance). The search range was then expanded to C = [1, 100] and γ = [1e-1, 1e2] to explore more potential high-quality candidate points.
[0091] Step 274 continues until a predetermined number of iterations is reached or the prediction results of the surrogate model no longer improve significantly. The base learner combination and hyperparameter settings that yield the highest prediction objective function value when the preset conditions are met are taken as the optimal base learner combination and optimal hyperparameters.
[0092] The predetermined number of iterations can be set to 100; the surrogate model's prediction no longer improves significantly, meaning that in multiple consecutive iterations, the improvement in the predicted objective function value is less than a specified threshold.
[0093] In step 24, the dataset (including image data and corresponding subtypes) can be randomly divided into a training set and a test set in an 8:2 ratio. A classifier model is built using the optimal base learner combination and hyperparameters selected by the adaptive ensemble learning model, along with the bagging ensemble learning algorithm. The classifier model is trained using the training set, and its performance is evaluated on the test set. This application uses AUC as the model performance evaluation metric to reflect the model's classification and generalization abilities.
[0094] In step 25, the finally trained classifier model is applied to the new pituitary tumor imaging data. First, the input patient images are preprocessed and features are extracted. Then, the trained classifier model is used to classify the extracted features and output the classification results. Finally, doctors can formulate personalized quality strategies and assess prognosis based on the classification results.
[0095] The image data classification method of this application can automatically select the optimal combination of base learners and hyperparameter settings, improving the accuracy and reliability of pituitary tumor classification and solving the problem of difficulty in determining the optimal combination of base learners and hyperparameters in existing studies. In the adaptive iterative search process, this application can dynamically adjust the range of the hyperparameter space based on the prediction results of the surrogate model, achieving continuous optimization of the search space and improving search efficiency and accuracy. Compared with traditional grid search, random search, and other methods, it can converge to the optimal solution faster, reducing computational and time costs.
[0096] This application discloses an image data classification device, such as... Figure 5 As shown, it includes memory and processor. The memory 100 is used to store programs for image data classification methods; The processor 200 is configured to read the program that executes the image data classification method and execute the method described in any of the above embodiments.
[0097] This application also provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to cause the computer to perform the methods described in any of the above embodiments.
[0098] This application describes several embodiments, but these descriptions are exemplary and not restrictive, and it will be apparent to those skilled in the art that many more embodiments and implementations are possible within the scope of the embodiments described herein. Although many possible combinations of features are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically limited, any feature or element of any embodiment may be used in combination with, or may replace, any feature or element of any other embodiment.
[0099] This application includes and contemplates combinations of features and elements known to those skilled in the art. The embodiments, features, and elements disclosed in this application can also be combined with any conventional features or elements to form unique inventive solutions. Any feature or element of any embodiment can also be combined with features or elements from other inventive solutions to form another unique inventive solution. Therefore, it should be understood that any feature shown and / or discussed in this application can be implemented individually or in any suitable combination. Therefore, the embodiments are not limited except by the limitations imposed by the appended claims and their equivalents. Furthermore, various modifications and changes can be made within the scope of the appended claims.
[0100] Furthermore, in describing representative embodiments, the specification may have presented methods and / or processes as a specific sequence of steps. However, the method or process should not be limited to the specific order of steps described herein, to the extent that it does not depend on such a specific order. As will be understood by those skilled in the art, other sequences of steps are also possible. Therefore, the specific order of steps set forth in the specification should not be construed as a limitation of the claims. Moreover, the claims concerning the method and / or process should not be limited to the steps performed in the written order, and those skilled in the art will readily understand that these orders can be varied and still remain within the spirit and scope of the embodiments of this application.
[0101] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term "computer storage medium" includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0102] Furthermore, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include at least one of those features.
[0103] In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise expressly and specifically limited.
[0104] In this application, unless otherwise expressly specified and limited, the terms "installation," "connection," "joining," "fixing," etc., should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral part; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be the internal communication of two components or the interaction between two components, unless otherwise expressly limited. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0105] In this application, unless otherwise expressly specified and limited, "above" or "below" the second feature can mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Furthermore, "above," "over," and "on top" of the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.
[0106] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0107] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. An image data classification method, comprising: Acquire and preprocess raw image data; Extracting radiomics features from preprocessed image data and filtering the radiomics features; Image data is classified based on a classifier model and selected radiomics features; wherein, the classifier model is configured to classify based on radiomics features; the classifier model is a model established based on the optimal combination of base learners, the optimal hyperparameters, and a preset learning algorithm; the optimal combination of base learners and the optimal hyperparameters are obtained based on an adaptive ensemble learning model; the adaptive ensemble learning model is configured to automatically select the optimal combination of base learners and the optimal hyperparameters from the candidate base learner set and the hyperparameter space.
2. The image data classification method as described in claim 1, characterized in that, The classifier model is established as follows: A first initial classifier model is established based on the optimal combination of base learners, the optimal hyperparameters, and the preset learning algorithm. The first initial classifier model is trained and evaluated using the first training set and the first test set, respectively; The first initial classifier model that has completed training is used as the classifier model.
3. The image data classification method as described in claim 1 or 2, characterized in that, The adaptive ensemble learning model is constructed in the following manner: Multiple initial candidate points are determined, each initial candidate point being a combination of base learners and hyperparameter settings randomly selected from the candidate base learner set and hyperparameter space; the objective function value of each initial candidate point is calculated based on the objective function and each initial candidate point, the objective function being used to evaluate the classification performance under different combinations of base learners and hyperparameter settings; A surrogate model is trained based on multiple initial candidate points and the objective function value of each initial candidate point. The trained surrogate model is then used to predict candidate points in the candidate space other than the initial candidate points, or a predetermined number of randomly sampled candidate points. Based on the predicted objective function values of the candidate points other than the initial candidate points or the predetermined number of randomly sampled candidate points, and a collection function, the next candidate point to be evaluated is selected and added to the existing candidate points for surrogate model training. The surrogate model is iteratively updated. When the iteration ends, the base learner combination and hyperparameter settings that result in the highest predicted objective function value are selected as the optimal base learner combination and optimal hyperparameters. The collection function is used to obtain the expected improvement amount of the next candidate point to be evaluated based on the predicted objective function value of the candidate point and the currently known optimal objective function value.
4. The image data classification method as described in claim 3, characterized in that, The method further includes: The range of the hyperparameter space is adjusted based on the predicted objective function value and prediction variance of candidate points in the candidate space according to the trained surrogate model. The surrogate model is retrained based on the candidate points after adjusting the hyperparameter space.
5. The image data classification method as described in claim 4, characterized in that, The step of adjusting the range of the hyperparameter space based on the predicted objective function value and prediction variance of candidate points in the candidate space according to the trained surrogate model includes: The hyperparameter space includes multiple regions; If there exists a region in the current hyperparameter space where the average value of the predicted objective function of the surrogate model is higher than a first preset threshold in other regions, the range of the hyperparameter space is narrowed to that region. If there is no region in the current hyperparameter space where the average value of the predicted objective function of the surrogate model is higher than a first preset threshold in other regions, or if the average value of the predicted variance of the current hyperparameter space calculated according to the surrogate model exceeds a second preset threshold, the range of the current hyperparameter space will be expanded.
6. The image data classification method as described in claim 3, characterized in that, The step of calculating the objective function value for each initial candidate point based on the objective function and each initial candidate point includes: A second initial classifier model is built based on each initial candidate point and a preset learning algorithm; The second initial classifier model is trained and evaluated using the first training set and the first test set, respectively; Use the trained second initial classifier model as the second classifier model; The objective function is used to evaluate the second classifier model to obtain the objective function value corresponding to each initial candidate point.
7. The image data classification method as described in claim 3, characterized in that, The step of selecting the next candidate point to be evaluated based on the predicted objective function value and acquisition function of candidate points other than the initial candidate point in the candidate space includes: The candidate point with the largest output value of the acquisition function is selected as the next candidate point to be evaluated.
8. The image data classification method as described in claim 3, characterized in that, The candidate base learner set includes one or more of the following algorithms: support vector machine, random forest, decision tree; The hyperparameter space includes one or more of the following: the hyperparameter space of support vector machines, the hyperparameter space of random forests, and the hyperparameter space of decision trees.
9. The image data classification method as described in claim 3, characterized in that, The objective function is a classification performance evaluation index function; The proxy model is a Gaussian process model; The acquisition function is the function to be improved.
10. The image data classification method as described in claim 3, characterized in that, The iteration termination condition includes one of the following: the number of iterations equals the preset number of iterations; the difference in the predicted objective function values of the surrogate model after multiple consecutive iterations is less than a third preset threshold.
11. An image data classification device, comprising a memory and a processor, characterized in that, The memory is used to store the program for the image data classification method; The processor is configured to read the program that executes the image data classification method and execute the method according to any one of claims 1 to 10.
12. A computer-readable storage medium storing computer-executable instructions, wherein, The computer-executable instructions are used to cause the computer to perform the method according to any one of claims 1 to 10.