Tobacco leaf grade recognition method and system based on MHHO algorithm and SVM model
By introducing the MHHO algorithm to optimize the SVM model, the problems of low manual assessment efficiency and local optimal parameter selection in tobacco leaf grade recognition are solved, and more efficient and accurate tobacco leaf grade recognition are achieved.
Patent Information
- Application Number
- CN202111490607.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-12-08
AI Technical Summary
In the prior art, the tobacco leaf grade recognition method relies on manual assessment, and there are problems such as low efficiency and difficult to guarantee accuracy. In addition, there are local optimal problems in parameter selection of existing SVM models, which affects the recognition accuracy.
The SVM model is optimized based on the MHHO algorithm, and the Harris Hawk algorithm is updated through chaotic perturbation convergence and nonlinear time-varying strategies, the parameter selection of support vector machines is optimized, and the search ability and convergence performance of the algorithm are enhanced.
It improves the efficiency and accuracy of tobacco leaf grade recognition, reduces data errors, and achieves higher classification accuracy and stability.
Smart Images

Figure CN114266931B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tobacco recognition, and particularly to a method and system for identifying tobacco leaf grades based on the MHHO algorithm and the SVM model. Background Art
[0002] In the tobacco industry, the grade of tobacco leaves directly affects the quality and taste of cigarettes. Therefore, the classification of tobacco leaf grades is of great significance. The traditional classification of tobacco leaf grades mainly relies on professionals to identify the quality of tobacco leaves through visual, tactile, olfactory and other senses, and then comprehensively evaluate the tobacco leaf grades. This method involves more subjective factors and has a great correlation with the experience of professionals, resulting in low efficiency and difficult to guarantee the accuracy.
[0003] Currently, features such as color, size, shape and surface texture of tobacco leaf images are extracted and analyzed, and these feature data are input into a neural network integrated with a fuzzy set, and then the tobacco leaf grades are estimated and predicted. In the identification of tobacco grades by image, a large number of tobacco leaf images reflecting the characteristics of various types of tobacco leaves need to be collected. The quality of tobacco leaf images is interfered by factors such as color, brightness and clarity, resulting in difficulty in obtaining an ideal tobacco leaf classification effect.
[0004] In the prior art, it is found that 22 chemical components are screened out from the chemical components of tobacco leaves by means of ReliefF and particle swarm optimization algorithms, and then input into a support vector machine to classify the quality grades of tobacco leaves by analyzing the chemical components of tobacco leaves; since the support vector machine can improve the generalization ability of the classifier, a classifier with high accuracy can also be trained using a small sample data set, and the non-linear characteristics can be used to fit the deviation of the existing error data. The support vector machine has been proved to be an economical and efficient tobacco leaf chemical component data classification technology.
[0005] The parameters of the support vector machine affect its classification accuracy, and there are certain difficulties in the selection of parameters. Therefore, the Harris hawk optimization algorithm is currently used to select the optimal parameters of the support vector machine. At different stages, Harris hawk individuals adopt different strategies to hunt. This multi-strategy search method makes the Harris hawk optimization algorithm have good optimization accuracy and convergence performance. However, when the algorithm iterates to the later stage, the population position update becomes convergent and the population diversity is single, making it difficult for the algorithm to jump out of the local optimum, thus affecting the accuracy and speed of algorithm convergence.
[0006] In view of this, there is an urgent need to provide a method that can enhance the search ability of the algorithm, enable the algorithm to jump out of the local optimum, improve the convergence performance of the algorithm, and thus effectively improve the efficiency of tobacco leaf grade recognition. Summary of the Invention
[0007] To solve the above technical problems, the technical solution adopted by the present invention is to provide a tobacco leaf grade recognition method based on the MHHO algorithm and the SVM model, including the following steps:
[0008] Obtain the chemical composition dataset of tobacco leaf grades, and train the initial SVM model in combination with the MHHO algorithm to obtain an optimized SVM model;
[0009] Obtain the chemical composition dataset of the tobacco leaf grades to be classified;
[0010] Input the dataset into the optimized SVM model;
[0011] Determine and classify the tobacco leaf grades according to the output results of the SVM model;
[0012] Among them, the optimization of the SVM model is specifically to optimize the parameter selection of the SVM model through the MHHO algorithm to achieve the optimization of the SVM model; the MHHO algorithm is specifically:
[0013] Use the nonlinear time-varying strategy of chaotic perturbation convergence to update the calculation strategy of the escape energy in the exploration-to-exploitation phase conversion of the Harris hawk algorithm;
[0014] Use the nonlinear time-varying strategy of chaotic perturbation convergence to update the position update strategy in the exploitation phase of the Harris hawk algorithm.
[0015] The present invention also provides a tobacco leaf grade recognition system based on the MHHO algorithm and the SVM model, including
[0016] Data input unit: used to input the chemical composition dataset of the tobacco leaf grades to be classified;
[0017] Tobacco leaf grade classification unit: According to the input chemical composition dataset of the tobacco leaf grades to be classified, use the optimized SVM model to classify the chemical composition dataset of the tobacco leaf grades;
[0018] Output unit: Determine the tobacco leaf grades according to the data classification results of the tobacco leaf grade classification unit;
[0019] The tobacco leaf grade classification unit includes an initial SVM model training module, which is used to train the initial SVM model in combination with the MHHO algorithm according to the obtained chemical composition dataset of the tobacco leaf grades to obtain an optimized SVM model; among them, the MHHO algorithm is specifically:
[0020] Use the nonlinear time-varying strategy of chaotic perturbation convergence to update the calculation strategy of the escape energy in the exploration-to-exploitation phase conversion of the Harris hawk algorithm;
[0021] Use the nonlinear time-varying strategy of chaotic perturbation convergence to update the position update strategy in the exploitation phase of the Harris hawk algorithm.
[0022] The present invention improves the Harris hawk optimization algorithm by introducing chaotic perturbation convergence and non-linear time-varying update strategies, and uses it to optimize the support vector machine model and select the optimal parameters of the support vector machine. The improved Harris hawk algorithm maintains the randomness and diversity of the population during the iteration process, enhances the search ability of the algorithm, enables the algorithm to jump out of the local optimum, improves the convergence performance of the algorithm, and at the same time enables the support vector machine to obtain better classification accuracy and stability, thereby reducing the error of the data, effectively improving the efficiency and accuracy of tobacco leaf grade recognition, and solving the problem that the selection of SVM model parameters in the prior art affects the accuracy of tobacco leaf grade recognition.
[0023] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other objects, features and advantages of the present invention more obvious and understandable, the following specifically describes the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0025] Figure 1 It is a schematic flowchart of the method provided by the present invention;
[0026] Figure 2 It is a flowchart of the MHHO algorithm for optimizing the SVM model provided by the present invention;
[0027] Figure 3 It is a convergence curve graph of the optimization of SVM parameters by each algorithm in six UCI data sets in the case of the present invention; specifically
[0028] Figure (a) is the convergence curve graph on Diagnostic; (b) is the convergence curve graph on heart; (c) is the convergence curve graph on Ionosphere; (d) is the convergence curve graph on iris; (e) is the convergence curve graph on seed; (f) is the convergence curve graph on sonar;
[0029] Figure 4 It is a system block diagram provided by the present invention;
[0030] Figure 5 It is a computer device structure block diagram provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0033] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined. In addition, the terms "mounted", "connected" and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0034] The basic idea of the present invention is that the dataset of tobacco leaf grade chemical components has the characteristics of a small sample and there are certain errors in the collection process. Support vector machines are suitable for small sample datasets and can obtain high-precision prediction values. In addition, the non-linear characteristics of support vector machines can fit the errors of the dataset. Therefore, the present invention identifies tobacco leaf grades based on support vector machines in the tobacco leaf chemical component dataset. Since when using the support vector machine prediction SVR for modeling, the penalty parameter C and the kernel parameter g have an important impact on the prediction performance of the model, the penalty parameter C is used to weigh the weight of the loss, and the kernel parameter g affects the radial range of the kernel function and determines the range and distribution characteristics of the training sample data. The fitness value is selected as the mean square error (MSE) for evaluation. Therefore, the present invention improves it to the MHHO algorithm by introducing the Harris hawk optimization algorithm, that is, adding a chaotic perturbation convergence and a non-linear time-varying update strategy to optimize the existing Harris hawk algorithm, searching for the optimal parameters C and g of the SVM algorithm, effectively increasing the population diversity, improving the local optimization ability and the global optimization ability of the algorithm, and thus improving the classification accuracy of the SVM model.
[0035] The following makes a detailed description of the present invention in combination with the specific embodiments and the drawings of the specification.
[0036] System embodiment
[0037] According to an embodiment of the present invention, a method for identifying tobacco leaf grades based on the MHHO algorithm and the SVM model is provided. As Figure 1 shown, it is a flowchart of the method for identifying tobacco leaf grades based on the MHHO algorithm and the SVM model provided by the present invention. The method includes:
[0038] Step S1: Obtain the dataset of tobacco leaf grade chemical components, and train the initial SVM model in combination with the MHHO algorithm to obtain an optimized SVM model;
[0039] Step S2: Obtain the dataset of tobacco leaf grade chemical components to be classified;
[0040] Step S2: Input the dataset into the optimized SVM model;
[0041] Step S3: Determine the tobacco leaf grades according to the output results of the SVM model and classify them.
[0042] In this embodiment, in step S1:
[0043] The parameters of the SVM model are optimized by the improved Harris hawk algorithm, hereinafter referred to as the MHHO algorithm, to implement the optimization of the SVM model and improve the classification accuracy of tobacco leaf grades. The MHHO algorithm is specifically as follows:
[0044] The nonlinear time-varying strategy of chaotic perturbation convergence is used to update the Harris Eagle algorithm to explore the calculation strategy of escape energy in the transition to the development stage, ensuring that the algorithm has the ability to mainly develop local capabilities and supplement global development capabilities in the later stage of iteration, ensuring the accuracy and speed of algorithm convergence; the updated escape energy calculation is as follows:
[0045]
[0046] yt=1-2*yt 2 (2)
[0047] In the formula, yt is the chaotic disturbance factor, E1 initial and E1 final These are the parameters of the chaotic perturbation convergence nonlinear time-varying strategy, which are the initial parameter value and the final value respectively.
[0048] This embodiment also uses the nonlinear time-varying strategy of chaotic perturbation convergence to update the position change strategy in the development phase of the Harris Hawk algorithm, that is, to update the positions of the soft encirclement, hard encirclement, soft encirclement of the rapid dive mode, and hard encirclement of the rapid dive mode. In this way, the nonlinear time-varying strategy of chaotic perturbation convergence is used to fully search the area around the individual position of the Harris Hawk to increase the local search capability of the algorithm. The updated position is calculated as follows:
[0049] omega=omega_initial*abs(yt)-(omega_initial-omega_final)*((Tt) / T) 2 (3)
[0050] Where omega is the chaotic dynamic weight factor, omega_initial and omega_final are the initial parameter and final value respectively, T is the total number of iterations, and t is the current number of iterations.
[0051] Updates have been made to the various optimization methods of the Harris Hawk, specifically:
[0052] Update the position of the soft bracket to:
[0053] ΔX(t)=omega*X rabbit (t)-X(t) (4)
[0054] The position of the hard surround is updated as:
[0055] X(t+1)=omega*X rabbit (t)-E|ΔX(t)| (5)
[0056] The position of the soft encirclement progressive fast dive is updated to:
[0057] Y = omega * X rabbit (t) - E|JX rabbit (t) - X(t)| (6)
[0058] The position update of the hard surrounded progressive rapid dive is:
[0059] Y = omega * X rabbit (t) - E|JX rabbit (t) - X m (t)| (7)
[0060] Preferably in this embodiment, the method of this embodiment uses the above MHHO algorithm to optimize the SVM to select parameters, realizes the optimized training of the SVM model, and obtains the optimized SVM model. The specific process of the optimized training of the initial SVM model is as follows:
[0061] In order to avoid the phenomenon of overfitting in the training process of the SVM model and more objectively evaluate the generalization ability of the data set in this embodiment, the k-fold cross-validation method is introduced to train the initial SVM model. That is, the obtained tobacco leaf grade chemical composition data set is divided into k data sets, where 1 data set is the training set for training the SVM model, and the other k - 1 data sets are the test sets, which are respectively used to test the SVM model obtained by training. Then, the average value of the SVM classification accuracies obtained from k - 1 tests is taken as the classification accuracy of the SVM. As Figure 2 shown, the flow chart of optimizing the SVM model by the MHHO algorithm in this embodiment includes the following steps:
[0062] Step S21, set the population position X of the MHHO algorithm = {X1,..., X i ,..., X N}, the individual position X i = [X i1 , X i2 , where X i1 , X i2 are respectively the penalty parameter C and the kernel parameter g of the support vector machine;
[0063] Step S22, randomly initialize the population position within the value ranges of the two parameters C and g, and set the maximum number of iterations Maxit of the algorithm as the end condition;
[0064] Step S23, use the K-fold cross-validation method to divide the obtained tobacco leaf grade chemical composition data set into 1 training set and K - 1 test sets;
[0065] Step S24: Initialize one individual of the population. Substitute the parameters C and g values corresponding to the position of this individual into the initial SVM model, train the initial SVM model with the training set, then use the test set to test the trained SVM model respectively, and obtain K - 1 classification accuracy values. Calculate the average of the K - 1 classification accuracy values to obtain the classification accuracy value of the corresponding individual, which is the fitness value of this individual; and calculate the fitness value of each individual in the population one by one;
[0066] Step S25: Use the MHHO algorithm to update the positions of the population individuals. Call Step S24 to obtain the fitness value of the corresponding individual at the new position. If the fitness value of the individual at the new position is better than the fitness value of the original individual position, then replace the original individual position with the new individual position, and complete the positions of all updated population individuals in turn;
[0067] Step S26: Find the optimal individual fitness in the population, and record the individual position corresponding to this fitness as the current population optimal individual position;
[0068] Step S27: Determine whether the current iteration number meets the maximum iteration number Maxit. If it meets, determine the values of parameters C and g according to the current population optimal individual position, and go to Step S28; otherwise, go to Step S25 to continue the iteration;
[0069] Step S28: Input the optimal parameters C and g into the trained SVM model for training to obtain an optimized SVM model.
[0070] In Step S27, the optimal individual position of the current population is the best values of parameters C and g. Therefore, inputting the two best parameters into the trained SVM model can obtain an optimized SVM model.
[0071] In this embodiment, K is taken as 5, which ensures that while training the SVM model, the computer operation amount is not too large.
[0072] For the method of this embodiment, in order to ensure that each chemical component in the obtained tobacco leaf grade chemical composition dataset has an equal status in the training process, the following formula is used to normalize the tobacco leaf grade chemical composition data. In the formula, x is one of the values of a certain chemical component, x max 、x min are the maximum and minimum values of the chemical component data. The normalization processing of the dataset can effectively improve the performance of the classifier. The specific calculation formula is
[0073]
[0074] In addition, since the tobacco leaf grade data is non-numerical and cannot be input into the SVM model for effective classification, in this embodiment, it is numerically encoded, where 1 represents the B2F tobacco leaf grade, 2 represents the C2F tobacco leaf grade, 3 represents the C3F tobacco leaf grade, and 4 represents the X2F tobacco leaf grade.
[0075] In the invention of this embodiment, the Harris hawk optimization algorithm is improved by introducing chaotic perturbation convergence and non-linear time-varying update strategies, and is used to optimize the support vector machine model and select the optimal parameters of the support vector machine. The improved Harris hawk algorithm maintains the randomness and diversity of the population during the iteration process, enhances the search ability of the algorithm, enables the algorithm to jump out of the local optimum, improves the convergence performance of the algorithm, and at the same time enables the support vector machine to obtain better classification accuracy and stability, thereby reducing data errors, effectively improving the tobacco leaf grade recognition efficiency and accuracy, and solving the problem that the selection of SVM model parameters in the prior art affects the accuracy of tobacco leaf grade recognition.
[0076] The effectiveness of this method is analyzed through specific cases below.
[0077] 1) The operating environment of the simulation experiment in this embodiment is as follows: In the Windows10 operating system, an Inter(R) Core(TM) i7-7500 CPU @ 2.70GHz 2.90GHz processor is used, and the simulation is carried out with matlab R2019b.
[0078] The following uses the Harris hawk optimization algorithm HHO, particle swarm optimization algorithm PSO, and satin bowerbird optimization algorithm SBO as comparison algorithms to compare with the MHHO algorithm proposed in this embodiment method in optimizing SVM parameters. The parameter settings of the above algorithms are described in Table 1 below.
[0079] Table 1 Parameter settings of each algorithm
[0080]
[0081] Six UCI datasets are selected as the experimental datasets, and the characteristics of the experimental datasets are described in Table 2. The 5-fold cross-validation method is used to split the dataset to train and test the initial SVM model.
[0082] Table 2 Description of the experimental datasets
[0083]
[0084] 2) Analysis of experimental results
[0085] As Figure 3As shown, the convergence curves of the MHHO algorithm and other comparison algorithms for optimizing the SVM parameters on six UCI datasets after 30 iterations are presented in sequence. The abscissa of the figure is the number of iterations, and the ordinate is the classification accuracy. The experimental results on these six datasets show that compared with the other three algorithms, the MHHO algorithm has a faster convergence speed of the classification accuracy of the data and higher precision. When each algorithm optimizes the SVM parameters, it runs independently 10 times on the dataset, with 30 iterations each time, and four statistical indicators of the maximum value, average value, minimum value, and variance of the SVM classification accuracy are obtained. The statistical results are shown in Table 3 below.
[0086] Table 3 Experimental statistical results of the dataset
[0087]
[0088] As can be seen from Table 3, compared with the HHO, SBO, and PSO algorithms, the parameters selected by the MHHO algorithm provided in this embodiment result in higher SVM classification accuracy and smaller variance, indicating that the given MHHO algorithm has high convergence accuracy and stability in optimizing the SVM parameters. The reason is that after introducing chaotic perturbation convergence and non-linear time-varying update strategies into the HHO algorithm, the MHHO algorithm effectively increases the population diversity, improves the local optimization ability and global optimization ability of the algorithm, can find better SVM parameters, and improves the SVM classification accuracy.
[0089] 3) Application case analysis of the method in this embodiment in tobacco leaf grade recognition
[0090] Data collection and processing: The chemical composition of tobacco leaves is an important factor affecting the taste and quality of cigarettes. Tobacco leaves mainly contain seven chemical components, namely total sugar, reducing sugar, total alkaloid, potassium, chlorine, total nitrogen, and starch, and the units of these index data are all (%). The main grades of tobacco leaves are B2F, C2F, C3F, and X2F, etc. The main factors determining the grade of tobacco leaves are the chemical composition of tobacco leaves. Table 4 shows the data sample information of the chemical composition datasets of some tobacco leaf grades. The data comes from flue-cured tobacco bases in Guangxi, Yunnan, Chongqing, Hunan, etc.
[0091] Table 4 Chemical components contained in different grades of tobacco leaves
[0092]
[0093] Table 4 shows that there are obvious differences in the data values among the chemical components of tobacco leaves. Usually, during the process of training a classifier, larger data can have more influence on the classifier than smaller data. To ensure that each chemical component has an equal status during training, the seven chemical component data of tobacco leaves are normalized using Equation (8) above.
[0094] The analysis of the tobacco leaf grade recognition results is as follows:
[0095] In the chemical composition dataset of tobacco leaf grades, 70% is set as the training set for training the MHHO-optimized SVM model, and the remaining 30% of the data is used as the test set.
[0096] Finally, a tobacco leaf grade recognition rate of 93.06% is obtained. At this time, the corresponding SVM parameter values are C = 42.33 and g = 0.94. In this 30% test set, the actual tobacco leaf grades are compared with the predicted grades. The prediction results are shown in Table 5 below. Among them, the recognition rate of tobacco leaf grade B2F reaches 92.86%, the recognition rate of tobacco leaf grade C2F reaches 92.31%, the recognition rate of tobacco leaf grade C3F reaches 93.94%, and the recognition rate of tobacco leaf grade X2F reaches 93.10%.
[0097] Table 5 MHHO-SVM Tobacco Leaf Grade Prediction Results
[0098]
[0099] After testing with three comparison algorithms on 6 UCI datasets, it is proved that the SVM parameters selected by the improved Harris hawk optimization algorithm proposed in this method enable the SVM to obtain better classification accuracy and stability. Applying the improved Harris hawk optimization algorithm-optimized support vector machine model proposed in this method to the tobacco leaf chemical composition dataset can obtain a tobacco leaf grade recognition rate of 93.06%. In the future, during the data collection process, more accurate instruments will be used to analyze the chemical composition of tobacco leaves to reduce data errors, thereby further improving the tobacco leaf grade recognition rate.
[0100] System Embodiment
[0101] According to an embodiment of the present invention, a tobacco leaf grade recognition system based on the MHHO algorithm and the SVM model is provided. As Figure 4 shown, it is the block diagram of the tobacco leaf grade recognition system based on the MHHO algorithm and the SVM model provided by the present invention. The system includes:
[0102] Data input unit: used to input the chemical composition dataset of the tobacco leaf grades to be classified;
[0103] Tobacco leaf grade classification unit: classifies the chemical composition dataset of the tobacco leaf grades using the optimized SVM model according to the input chemical composition dataset of the tobacco leaf grades to be classified.
[0104] Grade confirmation and output unit: determines the tobacco leaf grade according to the data classification result of the tobacco leaf grade classification unit.
[0105] In this embodiment, the tobacco grade classification unit preferably includes an initial SVM model training module, which is used to train the initial SVM model according to the acquired tobacco grade chemical composition data set in combination with the MHHO algorithm to obtain an optimized SVM model; wherein the MHHO algorithm is specifically:
[0106] The nonlinear time-varying strategy of chaotic perturbation convergence is used to update the Harris Eagle algorithm to explore the calculation strategy of escape energy in the transition to the development stage, ensuring that the algorithm has the ability to mainly develop local capabilities and supplement global development capabilities in the later stage of iteration, ensuring the accuracy and speed of algorithm convergence; the updated escape energy calculation is as follows:
[0107]
[0108] yt=1-2*yt 2 (10)
[0109] In the formula, yt is the chaotic disturbance factor, E1 initial and E1 final These are the parameters of the chaotic perturbation convergence nonlinear time-varying strategy, representing the initial parameter value and the final value respectively.
[0110] This embodiment also uses the nonlinear time-varying strategy of chaotic perturbation convergence to update the position change strategy in the development phase of the Harris Hawk algorithm, that is, to update the positions of the soft encirclement, hard encirclement, soft encirclement of the rapid dive mode, and hard encirclement of the rapid dive mode. In this way, the nonlinear time-varying strategy of chaotic perturbation convergence is used to fully search the area around the individual position of the Harris Hawk to increase the local search capability of the algorithm. The updated position is calculated as follows:
[0111] omega=omega_initial*abs(yt)-(omega_initial-omega_finak)*((Tt) / T) 2 (11)
[0112] Where omega_initial and omega_final are the initial parameters and final values respectively, T is the total number of iterations, and t is the current number of iterations.
[0113] Updates have been made to the various optimization methods of the Harris Hawk, specifically:
[0114] Update the position of the soft bracket to:
[0115] ΔX(t)=omega*X rabbit (t)-X(t) (12)
[0116] The position of the hard surround is updated as:
[0117] X(t + 1) = omega * X rabbit (t) - E|ΔX(t)| (13)
[0118] The position update of the soft - enclosed progressive rapid dive is as follows:
[0119] Y = omega * X rabbit (t) - E|JX rabbit (t) - X(t)| (14)
[0120] The position update of the hard - enclosed progressive rapid dive is as follows:
[0121] Y = omega * X rabbit (t) - E|JX rabbit (t) - X m (t)| (15)
[0122] This embodiment is a system embodiment corresponding to the above - mentioned method embodiment. The specific process of optimizing the SVM model according to the MHHO algorithm can be understood with reference to the description of the method embodiment, and will not be elaborated here.
[0123] As Figure 5 shown, the present invention also provides a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the tobacco leaf grade recognition method based on the MHHO algorithm and the SVM model in the above - mentioned embodiment, or when the computer program is executed by a processor, it implements the tobacco leaf grade recognition method based on the MHHO algorithm and the SVM model in the above - mentioned embodiment.
[0124] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0125] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0126] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0127] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. A tobacco leaf grade recognition method based on the MHHO algorithm and the SVM model, characterized in that, It includes the following steps: Obtain the dataset of tobacco leaf grade chemical components, and train the initial SVM model in combination with the MHHO algorithm to obtain an optimized SVM model; Obtain the dataset of tobacco leaf grade chemical components to be classified; Input the dataset into the optimized SVM model; Determine the tobacco leaf grade according to the output result of the SVM model and classify it; Among them, the optimization of the SVM model is specifically to optimize the parameter selection of the SVM model through the MHHO algorithm to achieve the optimization of the SVM model; the specific MHHO algorithm is as follows: Use the non-linear time-varying strategy of chaotic perturbation convergence to update the calculation strategy of the escape energy in the exploration-to-exploitation stage conversion of the Harris hawk algorithm; Use the non-linear time-varying strategy of chaotic perturbation convergence to update the position update strategy in the exploitation stage of the Harris hawk algorithm.
2. The tobacco leaf grade recognition method based on the MHHO algorithm and the SVM model according to claim 1, wherein The calculation strategy of the escape energy updated by using the non-linear time-varying strategy of chaotic perturbation convergence to update the exploration-to-exploitation stage conversion of the Harris hawk algorithm is as follows: yt = 1 - 2 * yt 2 where yt is the chaotic perturbation factor, E1 initial and E1 final are the initial maximum and minimum values respectively; Use the non-linear time-varying strategy of chaotic perturbation convergence to update the position update strategy in the exploitation stage of the Harris hawk algorithm, and the updated position calculation is as follows: omega = omega_initial * abs(yt) - (omega_initial - omega_final) * ((T - t) / T) 2 In the formula, omega_initial and omega_final are the initial parameter and the termination value respectively, T is the total number of iterations, and t is the current number of iterations; Update each optimization method of the Harris hawk, specifically as follows: Update the position of the soft encirclement to: ΔX(t) = omega * X rabbit (t) - X(t) Update the position of the hard encirclement to: X(t + 1) = omega * X rabbit (t) - E|ΔX(t)| Update the position of the soft encirclement progressive rapid dive to: Y = omega * X rabbit (X) - t|JX rabbit (X) - X(t)| Update the position of the hard encirclement progressive rapid dive to: Y = omega * X rabbit (t) - E|JX rabbit (t) - X m (t)|.
3. The tobacco leaf grade recognition method based on the MHHO algorithm and the SVM model according to claim 1, wherein, According to the obtained dataset of tobacco leaf grade chemical components, train the initial SVM model in combination with the k-fold cross-validation method.
4. The tobacco leaf grade recognition method based on the MHHO algorithm and the SVM model according to claim 1, characterized in that, The specific implementation steps of obtaining the dataset of tobacco leaf grade chemical components and training the initial SVM model in combination with the MHHO algorithm to obtain an optimized SVM model are as follows: Step S21. Set the population position X of the MHHO algorithm as X = {X1, …, X i , …, X N}, and the individual position X i = [X i1 , X i2 , where X i1 and X i2 are the penalty parameter C and the kernel parameter g of the support vector machine, respectively; Step S22: Randomly initialize the population position within the value ranges of two parameters C and g, and set the maximum number of algorithm iterations Maxit as the end condition; Step S23: Use the K-fold cross-validation method to divide the obtained dataset of tobacco leaf grade chemical components into 1 training set and K - 1 test sets; Step S24: Select one individual from the initialized population according to the selection. Substitute the values of parameters C and g corresponding to the position of this individual into the initial SVM model, train the initial SVM model with the training set, then use the test sets to test the trained SVM model respectively, and obtain K - 1 classification accuracy values. Calculate the average of the K - 1 classification accuracy values to obtain the fitness value of the corresponding individual; and calculate the fitness value of each individual in the population one by one; Step S25: Use the MHHO algorithm to update the positions of the population individuals, call Step S24 to obtain the fitness value of the corresponding individual at the new position, judge that if the fitness value of the individual at the new position is better than the fitness value of the original individual position, then replace the original individual position with the new individual position, and complete the positions of all updated population individuals in turn; Step S26: Search for the optimal individual fitness in the population, and record the individual position corresponding to this fitness as the current population optimal individual position; Step S27: Determine whether the current iteration count meets the maximum iteration count Maxit. If it does, determine the values of parameters C and g based on the position of the optimal individual in the current population, and go to step S28; otherwise, go back to step S25 to continue the iteration. Step S28: Input the optimal parameters C and g into the trained SVM model for training to obtain an optimized SVM model.
5. The tobacco leaf grade recognition method based on the MHHO algorithm and the SVM model according to claim 1, characterized in that Normalize the obtained tobacco leaf grade chemical composition dataset and the tobacco leaf grade chemical composition dataset to be classified.
6. A tobacco leaf grade recognition system based on the MHHO algorithm and the SVM model, characterized in that, Include Data input unit: used to input the tobacco leaf grade chemical composition dataset to be classified. Tobacco leaf grade classification unit: classify the tobacco leaf grade chemical composition dataset using the optimized SVM model according to the input tobacco leaf grade chemical composition dataset to be classified. Grade confirmation output unit: determine the tobacco leaf grade according to the data classification result of the tobacco leaf grade classification unit. The tobacco leaf grade classification unit includes an initial SVM model training module, which is used to train the initial SVM model using the obtained tobacco leaf grade chemical composition dataset in combination with the MHHO algorithm to obtain an optimized SVM model; where the MHHO algorithm is specifically: Update the calculation strategy of the escape energy in the exploration-to-exploitation phase of the Harris hawks optimization algorithm using a non-linear time-varying strategy with chaotic perturbation convergence. Update the position update strategy in the exploitation phase of the Harris hawks optimization algorithm using a non-linear time-varying strategy with chaotic perturbation convergence.
7. The tobacco leaf grade recognition system based on the MHHO algorithm and the SVM model according to claim 6, characterized in that, The calculation of the updated escape energy for updating the calculation strategy of the escape energy in the exploration-to-exploitation phase of the Harris hawks optimization algorithm using a non-linear time-varying strategy with chaotic perturbation convergence is as follows: yt = 1 - 2 * yt 2 where yt is the chaotic perturbation factor, E1 initial and E1 final are the initial maximum and minimum values respectively; The calculation of the updated position for updating the position update strategy in the exploitation phase of the Harris hawks optimization algorithm using a non-linear time-varying strategy with chaotic perturbation convergence is as follows: omega = omega_initial * abs(yt) - (omega_initial - omega_final) * ((T - t) / T) 2 In the formula, omega_initial and omega_final are the initial parameter and the termination value respectively, T is the total number of iterations, and t is the current iteration count. Update each optimization method of the Harris hawks, specifically: Update the position of the soft besiege to: ΔX(t) = omega * X rabbit (t) - X(t) Update the position of the hard besiege to: X(t + 1) = omega * X rabbit (t) - E|ΔX(t)| Update the position of the soft besiege progressive rapid dive to: Y = omega * X rabbit (t) - E|JX rabbit (t) - X(t)| Update the position of the hard besiege progressive rapid dive to: Y = omega * X rabbit (t) - E|JX rabbit (t) - X m (t)|.
8. The tobacco leaf grade recognition system based on the MHHO algorithm and the SVM model according to claim 6, wherein The specific training process of the SVM model training module is as follows: Step S21. Let the population position X of the MHHO algorithm be X = {X1, …, X i , …, X N}, and the individual position X i = [X i1 , X i2 , where X i1 and X i2 are the penalty parameter C and the kernel parameter g of the support vector machine, respectively; Step S22: Randomly initialize the population position within the value ranges of the two parameters C and g, and set the maximum iteration count Maxit of the algorithm as the termination condition. Step S23: Use the K-fold cross-validation method to divide the obtained tobacco leaf grade chemical composition dataset into 1 training set and K - 1 test sets. Step S24: Select one individual from the initialized population, substitute the values of parameters C and g corresponding to the position of this individual into the initial SVM model, train the initial SVM model using the training set, then use the test sets to test the trained SVM model respectively, and obtain K - 1 classification accuracy values. Calculate the average of the K - 1 classification accuracy values to obtain the classification accuracy value of the corresponding individual, which is the fitness value of this individual; and calculate the fitness value of each individual in the population one by one. Step S25: Update the positions of the population individuals using the MHHO algorithm, call Step S24 to obtain the fitness value of the corresponding new position of the individual, and judge that if the fitness value of the new position of the individual is better than the fitness value of the original individual position, then replace the original individual position with the new individual position, and complete the positions of all the updated population individuals in sequence; Step S26: Find the optimal individual fitness in the population, and record the individual position corresponding to this fitness as the current population optimal individual position; Step S27: Judge whether the current iteration number meets the maximum iteration number Maxit. If it meets, determine the values of parameters C and g according to the current population optimal individual position, and go to Step S28. Otherwise, go to Step S25 to continue the iteration; Step S28: Input the optimal parameters C and g into the trained SVM model for training to obtain an optimized SVM model.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the tobacco leaf grade recognition method based on the MHHO algorithm and the SVM model according to any one of claims 1 to 5.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the tobacco leaf grade recognition method based on the MHHO algorithm and the SVM model according to any one of claims 1 to 5.
Citation Information
Patent Citations
RBF (Radial Basis Function) neural network optimization method based on improved Harlisia eagle algorithm
CN113240069A
Bearing defect identification method based on SDAE and improved GWO-svm
WO2021128510A1