Traditional Chinese Medicine Syndrome Classification Method and Device Based on Improved Harris Hawk Optimization Algorithm
Through the improved Harris Eagle optimization algorithm and genetic algorithm, a Chinese medicine proof classification model was established, which solved the problems of complex operations and insufficient generalization capabilities of existing algorithms in traditional Chinese medicine proof classification, and improved the classification accuracy and generalization capabilities of the model.
Patent Information
- Application Number
- CN202311870483.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-09
- Filing Date
- 2023-12-29
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-12-29
AI Technical Summary
Existing machine learning and deep learning algorithms have problems with complex computing and insufficient generalization capabilities in the traditional Chinese medicine proof classification, which limits the better development of traditional Chinese medicine.
The improved Harris Eagle optimization algorithm combined with genetic algorithm is used to perform feature selection and parameter optimization to establish a traditional Chinese medicine proof classification model. The specific steps include obtaining the traditional Chinese medicine proof data set, standardizing processing, using genetic algorithms for feature selection, improving the Harris Eagle optimization algorithm for parameter optimization, and finally establishing and training the traditional Chinese medicine proof classification model.
It improves the accuracy of traditional Chinese medicine proof type classification, enhances the generalization ability of the model, reduces the complexity of the operation, and achieves faster and more accurate traditional Chinese medicine proof type recognition.
Smart Images

Figure CN118116574B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly to a traditional Chinese medicine syndrome type classification method and device based on an improved Harris hawk optimization algorithm. Background Art
[0002] Traditional Chinese medicine has gone through thousands of years of development and precipitation, and plays an irreplaceable role in preventing and treating diseases. It has unique advantages especially in the prevention and treatment of endocrine diseases, cardiovascular diseases, tumors, plagues and other diseases. Especially under the guidance of the theory of adapting treatment to individual conditions, traditional Chinese medicine pays attention to the personalized management of chronic diseases. It can not only achieve good treatment effects, but also has the advantages of few side effects and reducing the medical economic burden. It is a huge treasure for the treatment of chronic diseases such as diabetes and its complications. In 2019, the World Health Organization officially released the 11th edition of the International Classification of Diseases (ICD-11), and for the first time included traditional Chinese medicine in the global disease diagnosis standard, marking that traditional Chinese medicine has achieved parallel development with modern medicine. Especially after the epidemic, traditional Chinese medicine has received more and more attention worldwide.
[0003] Traditional Chinese medicine is a sharp sword. However, how to use it correctly is a huge challenge. Syndrome differentiation and treatment is the basic principle for traditional Chinese medicine to understand and deal with diseases. Among them, syndrome differentiation is the premise and basis of treatment, and also the soul of traditional Chinese medicine. In fact, most doctors in clinical practice conduct syndrome differentiation and treatment based on their own clinical experience. Inevitably, it takes a long process for doctors to learn and understand traditional Chinese medicine, and it is easy to have subjective biases in clinical practice. It is also these reasons that limit the better development of traditional Chinese medicine. How to quickly utilize the experience of experts to identify traditional Chinese medicine syndrome types more accurately, objectively and quickly is a very worthy exploration issue.
[0004] With the development of artificial intelligence technology, many machine learning and deep learning algorithms have been applied to syndrome type classification, combining traditional Chinese medicine with computer-aided diagnosis. However, the existing machine learning and deep learning algorithms face problems such as complex operations and insufficient generalization ability to varying degrees. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a traditional Chinese medicine syndrome type classification method and device based on an improved Harris hawk optimization algorithm.
[0006] A traditional Chinese medicine syndrome type classification method based on an improved Harris hawk optimization algorithm, the method includes:
[0007] Obtain a traditional Chinese medicine syndrome type data set of a target disease, standardize the traditional Chinese medicine syndrome type data set, and construct a training data set.
[0008] Feature selection is performed using a genetic algorithm based on the traditional Chinese medicine syndrome type dataset to obtain the optimal feature subset.
[0009] A traditional Chinese medicine syndrome type classification model is established based on the optimal feature subset.
[0010] An improved Harris hawk optimization algorithm is used to optimize the parameters of the traditional Chinese medicine syndrome type classification model.
[0011] The trained traditional Chinese medicine syndrome type classification model is retrained using the training dataset to obtain the trained traditional Chinese medicine syndrome type classification model.
[0012] The features to be recognized are input into the trained traditional Chinese medicine syndrome type classification model to obtain the traditional Chinese medicine syndrome type recognition result.
[0013] In one embodiment, the selection operator of the genetic algorithm uses the roulette wheel selection method, the crossover operator uses the single-point crossover method, and the mutation operator uses binary mutation; each chromosome in the genetic algorithm represents a feature subset; the encoding method of the chromosome gene positions is binary encoding, and the value on the chromosome gene positions is 1 or 0, representing the presence or absence of the feature column.
[0014] Feature selection is performed using a genetic algorithm based on the traditional Chinese medicine syndrome type dataset to obtain the optimal feature subset, including:
[0015] The traditional Chinese medicine syndrome type dataset is encoded in binary.
[0016] A number of individuals are randomly generated according to the obtained encoding results as the initial population.
[0017] A fitness function is constructed.
[0018] The individual fitness function values of each individual in the initial population are calculated.
[0019] Individual selection is performed using the roulette wheel selection method according to the fitness function values of each individual, the selected individuals are crossed using the single-point crossover method, and the crossover results are mutated using the binary mutation operator to obtain a new population. Continue to solve the individual fitness of each individual in the new population, and perform individual selection, crossover, and mutation according to the fitness values until the preset termination condition is met to obtain the optimal feature subset.
[0020] Among them, the steps of the binary mutation operator include:
[0021] A random number is generated for each gene position of the selected parental chromosome and compared with the mutation probability.
[0022] Judge whether the gene position needs to mutate, and perform mutation operations on the gene positions that need to mutate.
[0023] In one embodiment, the fitness function is constructed as follows:
[0024]
[0025] where f is the fitness value, acc(Classifier) is the accuracy of the traditional Chinese medicine syndrome classification model, α and β respectively represent the weights of the accuracy and the length of the feature subset selected by the algorithm; n represents the length of the selected feature subset, and N is the total number of feature attributes in the traditional Chinese medicine syndrome dataset.
[0026] In one embodiment, the traditional Chinese medicine syndrome classification model established according to the optimal feature subset is any classifier model based on machine learning; the classifier model includes but is not limited to: random forest model, XGBoost model, support vector machine model, and K-nearest neighbor model.
[0027] In one embodiment, the improved Harris hawk optimization algorithm is used to optimize the parameters of the traditional Chinese medicine syndrome classification model, including:
[0028] Set the population size, the maximum number of iterations, and the problem space dimension; each individual in the population represents a parameter combination of the traditional Chinese medicine syndrome classification model.
[0029] Set the current iteration number to 1.
[0030] Set the value range of the parameters of the traditional Chinese medicine syndrome classification model, and initialize the Harris hawk population using the Bernoulli chaotic map and the reverse learning strategy.
[0031] Calculate the fitness function value of the individual according to the training dataset.
[0032] Calculate the escape energy using the non-linear decay escape energy update strategy.
[0033] When the absolute value of the escape energy is less than 1, the individual position update strategy after adding Gaussian mutation in the exploitation stage is used for individual position update; when the absolute value of the escape energy is greater than or equal to 1, the individual position update strategy in the exploration stage is used for individual position update.
[0034] Judge whether the number of iterations reaches the maximum number of iterations. If it reaches the maximum number of iterations, output the optimal parameter combination and the fitness function value; if it does not reach the maximum number of iterations, calculate the individual fitness function value and the escape energy, increment the current iteration number by 1, and continue the next round of individual position update.
[0035] In one embodiment, the individual position update strategy after adding Gaussian mutation in the exploitation stage is as follows:
[0036]
[0037] Among them, is the optimal solution of the Harris hawk population at the k-th iteration after adding Gaussian mutation, is the optimal solution of the Harris hawk population obtained by measuring the individual position update in the exploitation phase of the classical Harris hawk algorithm at the k-th iteration, τ is a random number between [0, 1], and Gauss(0, 1) represents a Gaussian distribution function with a mean of 0 and a variance of 1.
[0038] In one embodiment, the non-linear decay escape energy update strategy is:
[0039] E = E1 × E0
[0040]
[0041] Among them, E is the escape energy of the prey, E0 is the initial energy of the prey, is a random number between [-1, 1], T represents the maximum number of iterations, and t represents the current number of iterations.
[0042] In one embodiment, setting the value range of the parameters of the traditional Chinese medicine syndrome type classification model and initializing the Harris hawk population using the Bernoulli chaotic map and the opposition-based learning strategy includes:
[0043] Setting the value range of the parameters of the traditional Chinese medicine syndrome type classification model to obtain the upper and lower bounds of the solution space.
[0044] Obtaining a chaotic sequence using the Bernoulli chaotic map.
[0045] Mapping the generated chaotic sequence into the solution space according to the upper and lower bounds of the solution space to obtain the Harris hawk population. The mapping formula for mapping the chaotic sequence into the solution space is:
[0046] X(t) = X lb +(X ub -X lb )×X'(t)
[0047] Among them, X lb is the lower bound of the search space, X ub is the upper bound of the search space, X(t) is the t-th individual in the Harris hawk population, t = 1, 2, 3,..., S, S is the number of individuals in the Harris hawk population, and X'(t) is the t-th particle in the chaotic sequence.
[0048] Optimizing the Harris hawk population using the opposition-based learning strategy.
[0049] In one embodiment, obtaining the traditional Chinese medicine syndrome type dataset of the target disease, standardizing the traditional Chinese medicine syndrome type dataset, and constructing a training dataset includes:
[0050] Obtain the traditional Chinese medicine syndrome type dataset of the target disease; the dataset includes input features and syndrome type classification labels.
[0051] Use the Z-Score method to standardize the traditional Chinese medicine syndrome type dataset to obtain the training dataset.
[0052] A traditional Chinese medicine syndrome type classification device based on an improved Harris hawk optimization algorithm, the method includes:
[0053] A training dataset determination module, configured to obtain the traditional Chinese medicine syndrome type dataset of the target disease, standardize the traditional Chinese medicine syndrome type dataset, and construct a training dataset.
[0054] A feature selection module, configured to perform feature selection using a genetic algorithm according to the traditional Chinese medicine syndrome type dataset to obtain an optimal feature subset.
[0055] A traditional Chinese medicine syndrome type classification model establishment module, configured to establish a traditional Chinese medicine syndrome type classification model according to the optimal feature subset.
[0056] A model parameter optimization module, configured to optimize the parameters of the traditional Chinese medicine syndrome type classification model using an improved Harris hawk optimization algorithm.
[0057] A traditional Chinese medicine syndrome type classification model retraining module, configured to retrain the traditional Chinese medicine syndrome type classification model with optimized parameters using the training dataset to obtain a trained traditional Chinese medicine syndrome type classification model.
[0058] A traditional Chinese medicine syndrome type recognition module, configured to input the features to be recognized into the trained traditional Chinese medicine syndrome type classification model to obtain the traditional Chinese medicine syndrome type recognition result.
[0059] The above-mentioned traditional Chinese medicine syndrome type classification method and device based on an improved Harris hawk optimization algorithm, the method includes: obtaining the traditional Chinese medicine syndrome type dataset of the target disease, standardizing the traditional Chinese medicine syndrome type dataset, and constructing a training dataset; performing feature selection using a genetic algorithm according to the traditional Chinese medicine syndrome type dataset to obtain an optimal feature subset; establishing a traditional Chinese medicine syndrome type classification model according to the optimal feature subset; optimizing the parameters of the traditional Chinese medicine syndrome type classification model using an improved Harris hawk optimization algorithm; retraining the traditional Chinese medicine syndrome type classification model with optimized parameters using the training dataset to obtain a trained traditional Chinese medicine syndrome type classification model; inputting the features to be recognized into the trained traditional Chinese medicine syndrome type classification model to obtain the traditional Chinese medicine syndrome type recognition result. This method improves the accuracy of traditional Chinese medicine syndrome type classification. Description of the Drawings
[0060] Figure 1 It is a flow schematic diagram of the traditional Chinese medicine syndrome type classification method based on an improved Harris hawk optimization algorithm in an embodiment;
[0061] Figure 2Schematic diagram of the improved Harris hawk optimization algorithm in another embodiment;
[0062] Figure 3 Frequency histogram generated by observing the Bernoulli chaotic map 1000 times in another embodiment;
[0063] Figure 4 Block diagram of a traditional Chinese medicine syndrome type classification device based on the improved Harris hawk optimization algorithm in one embodiment. Detailed implementation manners
[0064] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0065] In one embodiment, as Figure 1 shown, a traditional Chinese medicine syndrome type classification method based on the improved Harris hawk optimization algorithm is provided, and the method includes the following steps:
[0066] Step 100: Obtain a traditional Chinese medicine syndrome type data set of a target disease, standardize the traditional Chinese medicine syndrome type data set, and construct a training data set.
[0067] Step 102: Perform feature selection on the traditional Chinese medicine syndrome type data set by using a genetic algorithm to obtain an optimal feature subset.
[0068] Specifically, the genetic algorithm (GA) is a heuristic search method inspired by the genetic mechanism in nature and the theory of biological evolution, and can search for the optimal solution of a problem within a given space range.
[0069] The appearance of redundant and irrelevant features not only increases the dimension of the feature vector, but also reduces the performance of machine learning. Therefore, by feature selection to eliminate redundant features in the disease diagnosis data set, the model can have lower complexity and better classification performance.
[0070] Each chromosome in the genetic algorithm represents a feature subset, and binary coding is used. The value on the chromosome gene locus is 1 or 0, representing the presence or absence of the feature column.
[0071] Step 104: Establish a traditional Chinese medicine syndrome type classification model according to the optimal feature subset.
[0072] Specifically, the traditional Chinese medicine syndrome type classification model can be a support vector machine (SVM) model, or a classifier model such as a random forest or XGBoost model.
[0073] Step 106: Optimize the parameters of the traditional Chinese medicine syndrome type classification model by using an improved Harris hawk optimization algorithm.
[0074] Specifically, the improved Harris hawk optimization algorithm means that in the entire search space, based on the initialized Harris hawk population generated by the Bernoulli chaotic map, reverse learning is used to expand the search space and improve the quality of the population in the initial stage, so as to enhance the global search ability. In addition, Gaussian mutation is introduced in the exploitation stage, and a random perturbation conforming to the normal distribution is added to the optimal position of the Harris hawk, so that it can get rid of the bondage of local extrema, achieve global convergence, and use a non-linear decay escape energy update strategy to better simulate the physical energy change when the prey escapes.
[0075] Step 108: Retrain the traditional Chinese medicine syndrome type classification model with optimized parameters using the training data set to obtain a trained traditional Chinese medicine syndrome type classification model.
[0076] Step 110: Input the features to be recognized into the trained traditional Chinese medicine syndrome type classification model to obtain the traditional Chinese medicine syndrome type recognition result.
[0077] In the above traditional Chinese medicine syndrome type classification method based on the improved Harris hawk optimization algorithm, the method includes: obtaining the traditional Chinese medicine syndrome type data set of the target disease, standardizing the traditional Chinese medicine syndrome type data set, and constructing a training data set; performing feature selection by using a genetic algorithm according to the traditional Chinese medicine syndrome type data set to obtain an optimal feature subset; establishing a traditional Chinese medicine syndrome type classification model according to the optimal feature subset; optimizing the parameters of the traditional Chinese medicine syndrome type classification model by using an improved Harris hawk optimization algorithm; retraining the traditional Chinese medicine syndrome type classification model with optimized parameters using the training data set to obtain a trained traditional Chinese medicine syndrome type classification model; inputting the features to be recognized into the trained traditional Chinese medicine syndrome type classification model to obtain the traditional Chinese medicine syndrome type recognition result. This method improves the accuracy of traditional Chinese medicine syndrome type classification.
[0078] In one embodiment, the selection operator of the genetic algorithm adopts the roulette wheel selection method, the crossover operator adopts the single-point crossover method, and the mutation operator adopts binary mutation; each chromosome in the genetic algorithm represents a feature subset; the encoding method of the chromosome gene positions is binary encoding, and the values on the chromosome gene positions are 1 or 0, representing the presence or absence of feature columns; step 102 includes: performing binary encoding on the traditional Chinese medicine syndrome type dataset; randomly generating a number of individuals according to the obtained encoding results as the initial population; constructing a fitness function; calculating the individual fitness function values of each individual in the initial population; performing individual selection on the selected individuals by using the roulette wheel selection method, performing crossover on the selected individuals by using the single-point crossover method, and performing mutation on the crossover results by using the binary mutation operator to obtain a new population, continuing to solve the individual fitness of each individual in the new population, and performing individual selection, crossover, and mutation according to the fitness values until the preset termination condition is met to obtain the optimal feature subset; wherein, the steps of the binary mutation operator include: generating a random number for each gene position of the selected parent chromosome and comparing it with the mutation probability; determining whether the gene position needs to mutate, and performing a mutation operation on the gene position that needs to mutate.
[0079] Specifically, the selection operator in the genetic algorithm design is implemented by the roulette wheel selection method, and the crossover operator is implemented by the single-point crossover method, which are also common implementation methods in the genetic algorithm. In the design stage of the mutation operator, since feature selection is performed, this method adopts binary mutation, generates a random number for each gene position of the selected parent chromosome and compares it with the mutation probability, determines whether the gene position needs to mutate, and performs a mutation operation on the gene position that needs to mutate.
[0080] In one embodiment, to balance the number of selected optimal features and the classification accuracy, the constructed fitness function is:
[0081]
[0082] where f is the fitness value, acc(Classifier) is the accuracy of the traditional Chinese medicine syndrome type classification model, α and β respectively represent the weight of the accuracy and the weight of the length of the feature subset selected by the algorithm. Preferably, α is set to 0.9 and β is set to 0.1; n represents the length of the selected feature subset, and N is the total number of feature attributes in the traditional Chinese medicine syndrome type dataset.
[0083] In one embodiment, a traditional Chinese medicine syndrome type classification model is established according to the optimal feature subset as any machine learning classifier model; the classifier model includes but is not limited to: random forest model, XGBoost model, support vector machine model, and K-nearest neighbor model.
[0084] Specifically, the support vector machine is based on the VC dimension and the principle of structural risk minimization in statistical theory, aiming to find the optimal dividing hyperplane in the sample space to achieve the robustness and generalization ability of the optimal classification effect. Specifically, the SVM searches for support vectors in the sample space to construct the dividing hyperplane, thereby realizing the classification of data, and has been widely adopted in practical applications. The support vector machine algorithm can usually obtain better results than other classifiers during the experiment. When dealing with linearly inseparable problems, a support vector machine with a radial basis kernel function can be used, which can expand the feature space and thus solve the non-linearly separable problem.
[0085] The random forest model is an ensemble algorithm classifier, and all its base classifiers are decision trees, which are then integrated through the bagging method. From an intuitive perspective, each decision tree is a classifier (assuming it is a classification problem now). Then, for an input sample, N trees will have N classification results. The random forest model integrates all the classification voting results and designates the category with the most votes as the final output, which is the simplest bagging idea.
[0086] The XGBoost model (Adaptive Boosting) is an effective and practical Boosting algorithm that trains weak learners sequentially in a highly adaptive manner. The core idea of the XGBoost model is to adjust the weights of misclassified samples and then iterate and upgrade.
[0087] The K-nearest neighbor model classifies by measuring the distances between different feature values.
[0088] It should be noted that the syndrome classification model in this application can also be a classifier model of other types of machine learning in addition to the above classifier models.
[0089] In one embodiment, step 106 includes: setting the population size, the maximum number of iterations, and the problem space dimension; where each individual in the population represents a parameter combination of the traditional Chinese medicine syndrome classification model; setting the current iteration number to 1; setting the value range of the parameters of the traditional Chinese medicine syndrome classification model, initializing the Harris hawk population using the Bernoulli chaotic map and the opposition-based learning strategy; calculating the fitness function value of each individual according to the training data set; calculating the escape energy using the non-linear decay escape energy update strategy; when the absolute value of the escape energy is less than 1, then use the individual position update strategy after adding Gaussian mutation in the exploitation phase to update the individual position; when the absolute value of the escape energy is greater than or equal to 1, use the individual position update strategy in the exploration phase to update the individual position; determine whether the number of iterations reaches the maximum number of iterations, if it reaches the maximum number of iterations, then output the optimal parameter combination and the fitness function value; if it does not reach the maximum number of iterations, then calculate the individual fitness function value and the escape energy, increment the current iteration number by 1, and continue with the next round of individual position updates.
[0090] In one embodiment, the individual position update strategy in the exploitation phase after adding Gaussian mutation is:
[0091]
[0092] Where, is the optimal solution of the Harris hawk population at the k-th iteration after adding Gaussian mutation, is the optimal solution of the Harris hawk population obtained by measuring the individual position update in the exploitation phase of the classical Harris hawk algorithm at the k-th iteration, τ is a random number between [0, 1], and Gauss(0, 1) represents the Gaussian distribution function with a mean of 0 and a variance of 1.
[0093] Specifically, Gaussian mutation is an optimization strategy, which uses a random vector subject to a normal distribution to act on the original individual to generate a new position. The Gaussian probability density formula is as follows: Where μ is the expectation of the distribution and σ is the standard deviation.
[0094] To address the problem that the optimal position of the Harris hawk falls into a local optimal solution, Gaussian mutation is introduced in the exploitation phase of the Harris hawk optimization algorithm (HHO), and a random perturbation conforming to the normal distribution is added to the optimal position of the Harris hawk, enabling it to break free from the bondage of local extrema and achieve global convergence. The optimal solution position update strategy after adding Gaussian mutation is shown in formula (2).
[0095] In one embodiment, the non-linear decay escape energy update strategy is:
[0096] E = E1 × E0
[0097]
[0098] Among them, E is the escape energy of the prey, E0 is the initial energy of the prey, which is a random number between [-1, 1] and is automatically updated each iteration. T represents the maximum number of iterations, and t represents the current iteration number.
[0099] Specifically, in the classical Harris hawk optimization algorithm, the exploration or exploitation phase of the algorithm is determined according to the escape energy of the prey. The escape energy E linearly decreases from 2 to 0. In the later stage of iteration, the value of E is always less than 1, and only local exploitation is performed, without the ability of global exploration. The random exponential decay function is more suitable for simulating the physical energy change of the prey when escaping.
[0100] In one embodiment, the value range of the parameters of the traditional Chinese medicine syndrome classification model is set, and the Harris hawk population is initialized by using the Bernoulli chaotic mapping and the opposition-based learning strategy, including: setting the value range of the parameters of the traditional Chinese medicine syndrome classification model to obtain the upper and lower bounds of the solution space; obtaining the chaotic sequence by using the Bernoulli chaotic mapping; according to the upper and lower bounds of the solution space, mapping the generated chaotic sequence into the solution space to obtain the Harris hawk population. The mapping formula for mapping the chaotic sequence into the solution space is:
[0101] X(t) = X lb +(X ub -X lb )×X'(t) (4)
[0102] Where X lb is the lower bound of the search space, X ub is the upper bound of the search space, X(t) is the t-th individual in the Harris hawk population, t = 1, 2, 3,..., S, S is the number of individuals in the Harris hawk population, and X'(t) is the t-th particle in the chaotic sequence.
[0103] The Harris hawk population is optimized by using the opposition-based learning strategy.
[0104] In a specific embodiment, the flow of the improved Harris hawk optimization algorithm is as Figure 2 shown.
[0105] Aiming at the problem that the entire search space cannot be covered during the random initialization of the population, resulting in falling into the local optimum, population initialization based on Bernoulli chaotic mapping is proposed to increase the difference between population individuals, so as to improve the global search ability of the algorithm. The Bernoulli chaotic mapping distribution is very uniform, and the frequency histogram generated by mapping 1000 times is as Figure 3 shown.
[0106] The Bernoulli mapping expression is as follows:
[0107]
[0108] Among them, X(t) is the t-th particle in the population. Preferably, λ = 0.4.
[0109] According to the upper and lower bounds of the solution space, the generated chaotic sequence is mapped into the solution space using formula (4).
[0110] In the entire search space, based on the initialized Harris hawk population generated by the Bernoulli chaotic map, the opposition-based learning strategy is used to expand the search space and improve the quality of the population in the initial stage, so as to enhance the global search ability. The opposition-based learning strategy (OBL) refers to within the search space, based on the original solution, finding its opposite solution, and from the set of the original solution and the opposite solution, determining a better candidate solution by calculating the fitness value for the next iteration.
[0111] Currently, opposition-based learning has been used in the improvement of various optimization algorithms and achieved good results. The mathematical model of opposition-based learning is as follows:
[0112] x = [x1, x2, x3, …, x D (6)
[0113]
[0114]
[0115] Among them, D represents the dimension of the data, is the opposite solution of x, x j ∈[lb j , ub j , lb j is the lower bound of the solution space, and ub j is the upper bound of the solution space.
[0116] The classical Harris hawk optimization algorithm (HHO) consists of three stages, and obtains the optimal solution through multiple rounds of search and exploitation. The first stage is the search stage, where the Harris hawk is in the state of looking for prey; the second stage is the search and exploitation transition stage, where the Harris hawk is in the state of finding prey and can be in the state of transitioning from the exploration stage to the exploitation stage; the third stage is the exploitation stage, in which the Harris hawk attacks the prey and captures it through four attack strategies: soft siege, hard siege, soft siege with progressive rapid dive, and hard siege with progressive rapid dive.
[0117] The process of the improved Harris hawk optimization algorithm includes:
[0118] In the search phase, Harris hawks can track and detect prey through their powerful eyes. In HHO, Harris hawks are candidate solutions, and the best candidate solution in each step is regarded as the target prey. In this phase, Harris hawks randomly perch somewhere and find the prey through two strategies:
[0119]
[0120] where t is the number of iterations, X(t) is the position vector of an individual Harris hawk at the t-th iteration of the algorithm, X(t + 1) is the position vector of an individual Harris hawk at the (t + 1)-th iteration, X rand (t) is the position vector of a randomly selected individual Harris hawk, X rabbit (t) is the position vector of the prey, that is, the position vector of the Harris hawk individual with the optimal fitness. r1, r2, r3, r4, and q are all random numbers between [0, 1]. q is used to randomly select the strategy to be adopted. lb and ub represent the upper and lower bounds of the algorithm search space respectively, and X m (t) is the average position vector of the current Harris hawk population, N is the size of the hawk group, and X i (t) represents the position vector of the i-th hawk at present, and the expression of X m (t) is:
[0121]
[0122] In the search and development transition phase, the HHO algorithm can switch from the exploration phase to the development phase and then switch between different development behaviors according to the escape energy of the prey. The escape energy of the prey continuously decreases during the escape process. To simulate this fact, the definition of the escape energy E of the prey is shown in formula (3). The value of |E| decreases from 2 to 0. When |E| ≥ 1, the escape energy of the prey is large, and Harris hawks enter the search phase to search for the prey position. When |E| < 1, the escape energy of the prey is small, and Harris hawks enter the development phase to capture the prey through four strategies (soft siege, hard siege, soft siege with progressive rapid dive, hard siege with progressive rapid dive).
[0123] In the development phase, Harris hawks capture the target prey detected in the exploration phase. However, the prey often tries to escape from the dangerous environment. Therefore, different chasing styles occur in real life. According to the escape behavior of the prey and the chasing strategy of Harris hawks, four possible strategies are proposed in HHO to simulate the development phase. The four strategies are determined by two control variables, E and r. E represents the escape energy of the prey, and r represents the probability of successful escape. The following is the mathematical model representation of the four development strategies:
[0124] (1) Soft siege strategy
[0125] Define \(r\) as a random number between \([0, 1]\) for selecting different development strategies. When \(0.5\leq|E|\lt1\) and \(r\geq0.5\), the prey still has enough energy and tries to escape by some random misleading jumps, but finally cannot escape. The Harris hawk makes the prey exhausted through a soft siege strategy and then launches a surprise attack. The position vector update rule of the soft siege strategy is as follows:
[0126] \(X(t + 1)=\Delta X(t)-E|J X rabbit (t)-X(t)| (11)
[0127] \(\Delta X(t)=X rabbit (t)-X(t) (12)
[0128] where \(\Delta X(t)\) represents the difference between the prey's position and the current position of the Harris hawk individual, and \(J\) is a random number between \([0, 2]\), representing the random jump intensity of the prey during the entire escape process.
[0129] (2) Hard siege strategy
[0130] When \(|E|\lt0.5\) and \(r\geq0.5\), the prey is exhausted and has very low escape energy. The position vector update rule of the hard siege strategy is as follows:
[0131] \(X(t + 1)=X rabbit (t)-E|\Delta X(t)| (13)
[0132] (3) Soft encirclement strategy of asymptotic rapid dive
[0133] When \(0.5\leq|E|\lt1\) and \(r\lt0.5\), the prey has enough energy to escape successfully and will still form a soft encirclement before the surprise attack. To mathematically model the escape pattern and jumping movement of the prey, Lévy flight is used in the HHO algorithm to simulate the prey's escape and the irregular and rapid dive movement of the Harris hawk around the escaping prey. The position update rule of the soft encirclement strategy of asymptotic rapid dive is as follows:
[0134] \(Y = X rabbit (t)-E|J X rabbit (t)-X(t)| (14)
[0135] \(X = Y+S\times LF(D)\) (15)
[0136] where \(D\) is the problem dimension, \(S\) is a \(1\times D\) random vector, and \(LF()\) is the mathematical expression of Lévy flight, which is calculated using the following formula:
[0137]
[0138] Among them, μ and v are random numbers between [0, 1], and β is a default constant set to 1.5.
[0139] The final result of position update using the soft encirclement strategy of asymptotic rapid dive is as follows:
[0140]
[0141] Among them, Y and Z are obtained from equations (14) and (15), f() is the fitness function, and the results of the two positions of Y and Z are compared with the solution generated in the previous time to determine the position of the Harris hawk group after this iteration.
[0142] (4) Hard encirclement strategy of asymptotic rapid dive
[0143] When |E| < 0.5 and r < 0.5, the prey does not have enough escape energy. In this strategy, Harris hawks try to reduce their distance from the average position of the escaping prey. The position update rule using the hard encirclement strategy of asymptotic rapid dive is as follows:
[0144]
[0145] Among them, the expressions of Y and Z are as follows:
[0146] Y = X rabbit (t) - E|JX rabbit (t) - X m (t)| (19)
[0147] Z = Y + S × LF(D) (20)
[0148] Aiming at the problem that the optimal position of Harris hawks falls into a local optimal solution, Gaussian mutation is introduced during the development stage of HHO. A random perturbation conforming to the normal distribution is added to the optimal position of Harris hawks, enabling them to break free from the bondage of local extrema and achieve global convergence. The updated formula for the optimal solution position after adding Gaussian mutation is as shown in (2).
[0149] The optimization performance of the improved Harris hawk algorithm (BGOHHO) is tested using 4 typical unconstrained minimization problems shown in Table 1.
[0150] Table 1 4 typical unconstrained minimization problems
[0151]
[0152]
[0153] Well-known meta-heuristic algorithms, namely Particle Swarm Optimization (PSO), Grey Wolf Optimization (GWO), Whale Optimization Algorithm (WOA), and the basic Harris Hawks Optimization (HHO) are selected as comparative algorithms. In this experiment, the population size is uniformly set to 30, and the maximum number of iterations is 500. In the PSO algorithm, the learning factors are set as c1 = c2 = 2, and the inertia weight w varies between 0.2 and 0.9. To avoid the randomness of the experiment, for each test function, all algorithms are independently run 30 times, and the optimal value (Best), average value (Mean), and standard deviation (Std) obtained from the experiment are recorded as the measurement criteria for the algorithm performance. The experimental simulation results are shown in Table 2 below.
[0154] Table 2 Test Results of Objective Functions
[0155]
[0156] As can be seen from Table 2, compared with HHO, PSO, WOA, and GWO, the proposed BGOHHO in this application has obtained the optimal convergence performance. Although none of the functions have obtained the theoretical optimal value, the optimal value obtained by the convergence of BGOHHO is far better than that of other algorithms. It can also be seen from the average value and standard deviation that the optimized BGOHHO has better solving ability and stronger robustness.
[0157] In one embodiment, step 100 includes: obtaining a traditional Chinese medicine syndrome type dataset of the target disease; the dataset includes input features and syndrome type classification labels; the traditional Chinese medicine syndrome type dataset is standardized by the Z-Score method to obtain a training dataset.
[0158] In a specific embodiment, the improved Harris Hawks Optimization algorithm is used to optimize the parameters of the traditional Chinese medicine syndrome type classification model based on the support vector machine to obtain the optimal penalty factor C and kernel parameter G, so as to improve the classification performance of the model. The classification accuracy of the model is used as the fitness evaluation function. The specific steps are as follows:
[0159] Step 1: Initialize the relevant parameters of BGOHHO: the dimension of the problem space, the number of the population, the maximum number of iterations. Each individual in the population represents a parameter combination (C, G);
[0160] Step 2: Set the value range of the penalty factor C and the kernel parameter G, initialize the population using the Bernoulli chaotic mapping and the reverse learning mechanism, and import the training dataset to calculate the fitness function value of the individual;
[0161] Step 3: Calculate the escape energy E. If |E| < 1, update the individual position using the formula in the exploitation stage, otherwise update the individual position using the formula for updating the position in the exploration stage;
[0162] Step 4: Determine whether the algorithm has reached the maximum number of iterations. If so, output the optimal parameter combination (C, G). Otherwise, calculate the individual fitness value and go back to Step 3 to continue the iterative operation;
[0163] Step 5: Obtain the optimal parameter combination and retrain the traditional Chinese medicine syndrome type classification model based on the support vector machine using the training dataset;
[0164] Step 6: Use the trained traditional Chinese medicine syndrome type classification model to conduct tests on the test set and evaluate the performance of the model according to the set evaluation metrics.
[0165] For the traditional Chinese medicine syndrome type classification model based on the random forest model: Use the optimized improved Harris hawk optimization algorithm to optimize the parameters of the random forest syndrome type classification model, and obtain the optimal maximum depth of the tree (max_depth), the number of decision trees n_estimators, the maximum number of separating features (i.e., the number of feature variables to consider when finding the best node split) max_features, and the minimum number of samples required for each leaf node (min_samples_leaf) to improve the classification performance of the model.
[0166] For the traditional Chinese medicine syndrome type classification model based on the XGBoost model, use the optimized improved Harris hawk optimization algorithm to optimize the parameters of the XGBoost model, and obtain the optimal number of weak classifiers (n_estimators), the maximum depth of the decision tree (max_depth), and the learning rate (learning_rate) to improve the classification performance of the model.
[0167] For the traditional Chinese medicine syndrome type classification model based on the k-nearest neighbor model, use the optimized improved Harris hawk optimization algorithm to optimize the parameters of the k-nearest neighbor model, and obtain the optimal k value to improve the classification performance of the model.
[0168] It should be understood that although Figure 1 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in
[0169] In a verification embodiment, first, the genetic algorithm is used to perform feature selection on the dataset to reduce the redundant features in the dataset, obtain the key features required by the model, and improve the efficiency of model establishment. After obtaining the optimal feature subset, the improved Harris hawk optimization algorithm is used to optimize the penalty factor C and kernel parameter G of the traditional Chinese medicine syndrome type classification model based on the support vector machine, obtain the optimal parameter combination, and further improve the classification performance of the model.
[0170] The dataset of this embodiment selects the traditional Chinese medicine syndrome type data of type 2 diabetes from the information center platform of the Hunan Provincial Health Commission. There are 744 samples in total, among which the number of samples with qi and yin deficiency is 264, and the number of samples with liver and kidney yin deficiency is 480. The dataset includes 16 input features and 1 syndrome type classification label. To avoid the influence of dimensional differences in the dataset on the experimental results, the dataset is standardized using the Z-Score method before the experiment to eliminate dimensional differences and make the results obtained by the model more stable and reliable.
[0171] The system platform used in this embodiment is the Windows10 (64-bit) operating system, the programming language is Python, and the programming software is PyCharm. The number of populations is set to 10, and the number of iterations is set to 50. By randomly generating 70% of the training set and 30% of the test set, it runs independently 10 times, and the average result of the 10 operations on the test set is taken as the final experimental result. Finally, the classification accuracy, precision, recall rate, and F1 value are used to evaluate the classification performance of the model. The results of running 10 times on the traditional Chinese medicine syndrome type dataset of type 2 diabetes are shown in Table 3 below. It can be seen that among the 10 running results, the optimal classification accuracy of the model reaches 86.6071%, and at the same time, the corresponding precision of the model is 89.1892%, the recall rate is 90.4110%, and the F1 value is 89.7959%.
[0172] Table 3 Results of running the traditional Chinese medicine syndrome type classification model 10 times
[0173] Classification accuracy (%) Precision rate (%) Recall rate (%) F1 value (%) First run 84.8214 85.8974 91.7808 88.7417 Second run 85.2679 90.0709 86.9863 88.5017 Third run 85.7143 86.5384 92.4658 89.404 Fourth run 85.7143 85.625 93.8356 89.5425 Fifth run 86.6071 89.1892 90.4110 89.7959 Sixth run 85.2679 87.4172 90.4110 88.8889 Seventh run 85.7143 85.625 93.8356 89.5425 Eighth run 84.8214 84.5679 93.8356 88.961 Ninth run 84.8214 83.7349 95.2055 89.1026 Tenth run 86.1607 88.0795 91.0959 89.5623 Average value 85.4911 86.6745 91.9863 89.2043
[0174] In this embodiment, a control group experiment is also set up. The fusion model that uses both feature selection and parameter optimization is compared and analyzed with the GA_FS that only performs feature selection and the BGOHHO_SVM that only performs parameter optimization in terms of experimental results. The experimental results are shown in Table 4 below. The fusion model obtains the best experimental effect. In terms of classification accuracy, the classification effect of the fusion model is the best, reaching 85.4464%. The classification effect of GA_FS ranks second, with an accuracy of 84.0625%, which is 1.3839% lower than that of the fusion model. The effect of BGOHHO_SVM is the second best, 1.7410% lower than that of the fusion model. The classification accuracy of the original SVM model is the lowest, at 82.5893%, 2.8571% lower than that of the fusion model. In terms of classification precision, the fusion model also achieves the best result, with the classification precision reaching 86.5754%. Compared with it, the classification precisions of SVM, GA_FS, and BGOHHO_SVM are 1.6081%, 0.2558%, and 2.4776% respectively. In terms of classification recall rate, the evaluation effect of BGOHHO_SVM ranks first, with a recall rate of 92.5324%. The fusion model ranks second, with a difference of not much, at 92.0548%. GA_FS and the original SVM model rank behind. Just looking at precision and recall rate cannot well evaluate the performance of a model. As the harmonic mean of precision and recall rate, the F1 value can take both into account and better evaluate the model. In terms of the F1 value, the effect of the fusion model is the best, with the F1 value reaching 89.1826%. The effect of BGOHHO_SVM ranks second, at 88.0929%, 1.0897% lower than that of the fusion model. The evaluation effect of GA_FS follows closely, with an F1 value of 88.0096%, 1.1730% lower than that of the fusion model. The F1 value of the original SVM model is the lowest, at 86.9565%, 2.2261% lower than that of the fusion model. It can be seen that by using the genetic algorithm for feature selection to reduce the redundant features in the dataset, establishing a model with the optimal feature subset, and then optimizing the parameters of the model based on the improved Harris optimization algorithm BGOHHO, the classification performance of the final model is well improved.
[0175] Table 4 Comparison Table of Evaluation Indicators of Different Models
[0176]
[0177] In addition, the BGOHHO_GA_SVM model was experimentally compared with the fusion model optimized by the PSO algorithm (PSO_GA_SVM), the fusion model optimized by the WOA (WOA_GA_SVM), the fusion model optimized by the GWO (GWO_GA_SVM), and the unimproved fusion model optimized by the HHO (HHO_GA_SVM). The BGOHHO_GA_SVM model (the model of this method) achieved the best experimental results.
[0178] In summary: Using the genetic algorithm for feature selection on the traditional Chinese medicine syndrome type dataset of diabetes effectively removed redundant features, reduced the complexity of model establishment, and improved the classification performance of the model. To enable the Harris hawk optimization algorithm to be better applied to the parameter optimization of the model and further improve the classification performance of the model, a variety of strategies were proposed to improve it. Experiments were conducted on standard test functions, and compared with classical intelligent optimization algorithms, the best optimization performance was obtained, indicating that the improvement strategies proposed in this application can effectively help the Harris hawk optimization algorithm jump out of the local optimum. The final fusion model achieved the optimal classification accuracy, fully demonstrating that the feature selection and parameter optimization experiments designed based on intelligent optimization algorithms can effectively improve the classification performance of the model.
[0179] In one embodiment, as Figure 4 shown, a traditional Chinese medicine syndrome type classification device based on an improved Harris hawk optimization algorithm is provided, including: a training dataset determination module, a feature selection module, a traditional Chinese medicine syndrome type classification model establishment module, a model parameter optimization module, a traditional Chinese medicine syndrome type classification model retraining module, and a traditional Chinese medicine syndrome type recognition module, where:
[0180] The training dataset determination module is used to obtain the traditional Chinese medicine syndrome type dataset of the target disease, standardize the traditional Chinese medicine syndrome type dataset, and construct a training dataset;
[0181] The feature selection module is used to perform feature selection on the traditional Chinese medicine syndrome type dataset using the genetic algorithm to obtain an optimal feature subset;
[0182] The traditional Chinese medicine syndrome type classification model establishment module is used to establish a traditional Chinese medicine syndrome type classification model based on the optimal feature subset;
[0183] The model parameter optimization module is used to optimize the parameters of the traditional Chinese medicine syndrome type classification model using the improved Harris hawk optimization algorithm;
[0184] The traditional Chinese medicine syndrome type classification model retraining module is used to retrain the traditional Chinese medicine syndrome type classification model with optimized parameters using the training dataset to obtain a trained traditional Chinese medicine syndrome type classification model;
[0185] Traditional Chinese Medicine (TCM) syndrome type recognition module, which is used to input the features to be recognized into the trained TCM syndrome type classification model to obtain the TCM syndrome type recognition result.
[0186] In one embodiment, the selection operator of the genetic algorithm adopts the roulette wheel selection method, the crossover operator adopts the single-point crossover method, and the mutation operator adopts binary mutation; each chromosome in the genetic algorithm represents a feature subset; the encoding method of the chromosome gene positions is binary encoding, and the values on the chromosome gene positions are 1 or 0, representing the presence or absence of feature columns; the feature selection module is also used to perform binary encoding on the TCM syndrome type data set; randomly generate a number of individuals as the initial population according to the obtained encoding result; construct a fitness function; calculate the individual fitness function values of each individual in the initial population; perform individual selection on the selected individuals using the roulette wheel selection method, perform crossover on the selected individuals in a single-point crossover manner, and perform mutation on the crossover result using the binary mutation operator to obtain a new population, continue to solve the individual fitness of each individual in the new population, and perform individual selection, crossover and mutation according to the fitness values until the preset termination condition is met to obtain the optimal feature subset; among them, the steps of the binary mutation operator include: generating a random number for each gene position of the selected parent chromosome and comparing it with the mutation probability; determining whether the gene position needs to mutate, and performing mutation operations on the gene positions that need to mutate.
[0187] In one embodiment, the feature selection module is also used to construct a fitness function as shown in formula (1).
[0188] In one embodiment, the TCM syndrome type classification model establishment module, the TCM syndrome type classification model is any machine learning classifier model; the classifier model includes a random forest model, an XGBoost model, a support vector machine model, and a K-nearest neighbor model.
[0189] In one embodiment, the model parameter optimization module is further configured to set the population size, the maximum number of iterations, and the problem space dimension; where each individual in the population represents a parameter combination of the traditional Chinese medicine syndrome type classification model; set the current iteration number to 1; set the value range of the parameters of the traditional Chinese medicine syndrome type classification model, and initialize the Harris hawk population by using the Bernoulli chaotic map and the opposition-based learning strategy; calculate the fitness function value of the individual according to the training data set; calculate the escape energy by using the non-linear decay escape energy update strategy; when the absolute value of the escape energy is less than 1, update the individual position by using the individual position update strategy in the exploitation stage after adding Gaussian mutation; when the absolute value of the escape energy is greater than or equal to 1, update the individual position by using the individual position update strategy in the exploration stage; determine whether the number of iterations reaches the maximum number of iterations, if it reaches the maximum number of iterations, output the optimal parameter combination and the fitness function value; if it does not reach the maximum number of iterations, calculate the individual fitness function value and the escape energy, increment the current iteration number by 1, and continue the next round of individual position update.
[0190] In one embodiment, the individual position update strategy in the exploitation stage in the model parameter optimization module is shown in Equation (2).
[0191] In one embodiment, the non-linear decay escape energy update strategy in the model parameter optimization module is shown in Equation (3).
[0192] In one embodiment, the model parameter optimization module is further configured to set the value range of the parameters of the traditional Chinese medicine syndrome type classification model to obtain the upper and lower bounds of the solution space; obtain a chaotic sequence by using the Bernoulli chaotic map; according to the upper and lower bounds of the solution space, map the generated chaotic sequence into the solution space by using Formula (4) to obtain the Harris hawk population, and then optimize the Harris hawk population by using the opposition-based learning strategy.
[0193] In one embodiment, the training data set determination module is further configured to obtain the traditional Chinese medicine syndrome type data set of the target disease; the data set includes input features and syndrome type classification labels; perform standardization processing on the traditional Chinese medicine syndrome type data set by using the Z-Score method to obtain the training data set.
[0194] For the specific limitations of the traditional Chinese medicine syndrome type classification device based on the improved Harris hawk optimization algorithm, reference can be made to the limitations of the traditional Chinese medicine syndrome type classification method based on the improved Harris hawk optimization algorithm in the above text, which will not be elaborated here. Each module in the above-mentioned traditional Chinese medicine syndrome type classification device based on the improved Harris hawk optimization algorithm can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0195] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0196] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A traditional Chinese medicine syndrome classification method based on an improved Harris hawk optimization algorithm, characterized in that, The method includes: Obtain the traditional Chinese medicine syndrome type dataset of the target disease, standardize the traditional Chinese medicine syndrome type dataset, and construct a training dataset; Perform feature selection using a genetic algorithm according to the traditional Chinese medicine syndrome type dataset to obtain an optimal feature subset; Establish a traditional Chinese medicine syndrome type classification model according to the optimal feature subset; Use an improved Harris hawk optimization algorithm to optimize the parameters of the traditional Chinese medicine syndrome type classification model; Retrain the traditional Chinese medicine syndrome type classification model with optimized parameters using the training dataset to obtain a trained traditional Chinese medicine syndrome type classification model; Input the features to be recognized into the trained traditional Chinese medicine syndrome type classification model to obtain the traditional Chinese medicine syndrome type recognition result; Among them, the selection operator of the genetic algorithm adopts the roulette wheel selection method, the crossover operator adopts the single-point crossover method, and the mutation operator adopts binary mutation; each chromosome in the genetic algorithm represents a feature subset; the coding method of the chromosome gene position is binary coding, and the value on the chromosome gene position is 1 or 0, representing the presence or absence of a feature column; the fitness function is: Among them, f is the fitness value, acc(Classifier) is the accuracy of the traditional Chinese medicine syndrome type classification model, α and β respectively represent the weight of the accuracy and the weight of the length of the feature subset selected by the algorithm; n represents the length of the selected feature subset, and N is the total number of feature attributes in the traditional Chinese medicine syndrome type dataset; Among them, using an improved Harris hawk optimization algorithm to optimize the parameters of the traditional Chinese medicine syndrome type classification model includes: Set the population size, the maximum number of iterations, and the problem space dimension; each individual in the population represents a parameter combination of the traditional Chinese medicine syndrome type classification model; Set the current iteration number to 1; Set the value range of the parameters of the traditional Chinese medicine syndrome type classification model, and initialize the Harris hawk population using the Bernoulli chaotic map and the reverse learning strategy; Calculate the fitness function value of the individual according to the training dataset; Calculate the escape energy using the non-linear decay escape energy update strategy; When the absolute value of the escape energy is less than 1, update the individual position using the individual position update strategy in the exploitation stage after adding Gaussian mutation; when the absolute value of the escape energy is greater than or equal to 1, update the individual position using the individual position update strategy in the exploration stage; Judge whether the number of iterations reaches the maximum number of iterations. If it reaches the maximum number of iterations, output the optimal parameter combination and the fitness function value; if it does not reach the maximum number of iterations, calculate the individual fitness function value and the escape energy, increment the current iteration number by 1, and continue the next round of individual position update.
2. The method according to claim 1, characterized in that, The selection operator of the genetic algorithm adopts the roulette wheel selection method, the crossover operator adopts the single-point crossover method, and the mutation operator adopts binary mutation; each chromosome in the genetic algorithm represents a feature subset; the coding method of the chromosome gene position is binary coding, and the value on the chromosome gene position is 1 or 0, representing the presence or absence of a feature column; Performing feature selection using a genetic algorithm according to the traditional Chinese medicine syndrome type dataset to obtain an optimal feature subset includes: Perform binary coding on the traditional Chinese medicine syndrome type dataset; Randomly generate several individuals as the initial population according to the obtained coding results; Calculate the individual fitness function values of each individual in the initial population according to the fitness function; Perform individual selection on the selected individuals using the roulette wheel selection method according to the fitness function values of each individual, perform crossover on the selected individuals in a single-point crossover manner, and perform mutation on the crossover results using a binary mutation operator to obtain a new population. Continue to solve the individual fitness of each individual in the new population, and perform individual selection, crossover, and mutation according to the fitness values until the preset termination condition is met to obtain the optimal feature subset; Among them, the steps of the binary mutation operator include: Generate a random number for each gene position of the selected parental chromosome and compare it with the mutation probability; Judge whether the gene position needs to mutate, and perform mutation operations on the gene positions that need to mutate.
3. The method according to claim 1, characterized in that, Establish a traditional Chinese medicine syndrome type classification model according to the optimal feature subset as any classifier model based on machine learning; the classifier model includes but is not limited to: random forest model, XGBoost model, support vector machine model, and K-nearest neighbor model.
4. The method according to claim 1, characterized in that, The individual position update strategy in the exploration stage after adding Gaussian mutation is: Among them, is the optimal solution of the Harris hawk population at the k-th iteration after adding Gaussian mutation, is the optimal solution of the Harris hawk population obtained by measuring the individual position update in the exploitation stage of the classical Harris hawk algorithm at the k-th iteration. τ is a random number in the interval [0, 1], and Gauss(0, 1) represents the Gaussian distribution function with a mean of 0 and a variance of 1.
5. The method according to claim 1, characterized in that, The non-linear decay escape energy update strategy is: E = E1×E0 Among them, E is the escape energy of the prey, E0 is the initial energy of the prey, which is a random number between [-1, 1], T represents the maximum number of iterations, and t represents the current number of iterations.
6. The method according to claim 5, wherein, Set the value range of the parameters of the traditional Chinese medicine syndrome type classification model, and initialize the Harris hawk population using the Bernoulli chaotic map and reverse learning strategy, including: Set the value range of the parameters of the traditional Chinese medicine syndrome type classification model to obtain the upper and lower bounds of the solution space; Obtain a chaotic sequence using the Bernoulli chaotic map; According to the upper and lower bounds of the solution space, map the generated chaotic sequence into the solution space to obtain the Harris hawk population. The mapping formula for mapping the chaotic sequence into the solution space is: X(t) = X lb +(X ub -X lb )×X'(t) Among them, X lb is the lower bound of the search space, X ub is the upper bound of the search space, X(t) is the t-th individual in the Harris hawk population, t = 1, 2, 3, …, S, where S is the number of individuals in the Harris hawk population, and X'(t) is the t-th particle in the chaotic sequence; Optimize the Harris hawk population using the reverse learning strategy.
7. The method according to claim 1, wherein, Obtain the traditional Chinese medicine syndrome type dataset of the target disease, standardize the traditional Chinese medicine syndrome type dataset, and construct a training dataset, including: Obtain the traditional Chinese medicine syndrome type dataset of the target disease; the dataset includes input features and syndrome type classification labels; Perform standardization processing on the traditional Chinese medicine syndrome type dataset using the Z-Score method to obtain a training dataset.
8. A traditional Chinese medicine syndrome classification device based on an improved Harris hawk optimization algorithm, wherein, The device includes: A training dataset determination module, configured to obtain the traditional Chinese medicine syndrome type dataset of the target disease, standardize the traditional Chinese medicine syndrome type dataset, and construct a training dataset; A feature selection module, configured to perform feature selection using a genetic algorithm according to the traditional Chinese medicine syndrome type dataset to obtain an optimal feature subset; A traditional Chinese medicine syndrome type classification model establishment module, configured to establish a traditional Chinese medicine syndrome type classification model according to the optimal feature subset; A model parameter optimization module, configured to optimize the parameters of the traditional Chinese medicine syndrome type classification model using an improved Harris hawk optimization algorithm; The traditional Chinese medicine syndrome type classification model retraining module is used to retrain the optimized traditional Chinese medicine syndrome type classification model with the training data set to obtain a trained traditional Chinese medicine syndrome type classification model; The traditional Chinese medicine syndrome type recognition module is used to input the features to be recognized into the trained traditional Chinese medicine syndrome type classification model to obtain the traditional Chinese medicine syndrome type recognition result; Among them, the selection operator of the genetic algorithm in the feature selection module adopts the roulette wheel selection method, the crossover operator adopts the single-point crossover method, and the mutation operator adopts binary mutation; each chromosome in the genetic algorithm represents a feature subset; the encoding method of the chromosome gene position is binary encoding, and the value on the chromosome gene position is 1 or 0, representing the presence or absence of the feature column; the fitness function is: Among them, f is the fitness value, acc(Classifier) is the accuracy of the traditional Chinese medicine syndrome type classification model, α and β respectively represent the weights of the accuracy and the length of the feature subset selected by the algorithm; n represents the length of the selected feature subset, and N is the total number of feature attributes in the traditional Chinese medicine syndrome data set; The model parameter optimization module is also used to set the population size, the maximum number of iterations, and the problem space dimension; each individual in the population represents a parameter combination of the traditional Chinese medicine syndrome type classification model; set the current number of iterations to 1; set the value range of the parameters of the traditional Chinese medicine syndrome type classification model, and initialize the Harris hawk population using the Bernoulli chaotic map and the reverse learning strategy; calculate the fitness function value of the individual according to the training data set; calculate the escape energy using the non-linear decay escape energy update strategy; when the absolute value of the escape energy is less than 1, update the individual position using the individual position update strategy in the exploitation stage after adding Gaussian mutation; when the absolute value of the escape energy is greater than or equal to 1, update the individual position using the individual position update strategy in the exploration stage; determine whether the number of iterations reaches the maximum number of iterations. If the maximum number of iterations is reached, output the optimal parameter combination and the fitness function value; if the maximum number of iterations is not reached, calculate the individual fitness function value and the escape energy, add 1 to the current number of iterations, and continue the next round of individual position update.