Tobacco ecological region classification method and device based on metabonomics and medium

The metabolites data of tobacco leaf samples were obtained through metabolomics, pre-processed and feature selection were performed, and the training and fusion of multiple classification algorithm models was combined to solve the subjectivity and generalization of traditional tobacco leaf ecological region classification, and achieve higher precision tobacco leaf ecological region classification.

CN120381145APending Publication Date: 2025-07-29CHINA TOBACCO ZHEJIANG IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510660188.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The traditional tobacco ecological zone classification method relies on artificial experience, has strong subjectivity, insufficient information utilization, and poor generalization of the model, resulting in poor consistency and accuracy of classification results.

Method used

Using a metabolomics-based method, by obtaining the metabolites data set of tobacco leaf samples, and after data preprocessing, multiple feature selection methods are used to screen the optimal features, and combined with multiple classification algorithm models for training and fusion to generate a fusion classification model for tobacco leaf ecological area.

Benefits of technology

It significantly improves the objectivity and accuracy of classification of tobacco leaf ecological zone, solves the problems of strong subjectivity, low information utilization rate and insufficient model generalization ability in traditional methods, and improves the accuracy and efficiency of classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120381145A_ABST
    Figure CN120381145A_ABST
Patent Text Reader

Abstract

The invention relates to a metabonomics-based tobacco ecological region classification method and device and a medium. The method comprises the following steps: acquiring a metabolite data set of a tobacco leaf sample set; the tobacco leaf sample set comprises a plurality of tobacco leaf samples of each ecological region and corresponding ecological region identifiers; performing data preprocessing on the metabolite data set to obtain a sample feature set; selecting a plurality of optimal features from the sample feature set by using a plurality of feature selection methods to form an optimal feature data set; using the optimal feature data set to train and fuse a plurality of classification algorithm models to obtain a trained tobacco ecological region fusion classification model; and obtaining the plurality of optimal features of to-be-classified tobacco leaves, and inputting the optimal features to the trained tobacco leaf ecological region fusion classification model to obtain an ecological region prediction result. According to the method, accurate and automatic classification of the tobacco ecological region can be realized, and the technical problems that a traditional method depends on artificial experience, information utilization is insufficient and model generalization is poor are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tobacco leaf ecological area classification, and particularly to a method, device and medium for classifying tobacco leaf ecological areas based on metabolomics. Background Art

[0002] In the tobacco industry, the classification of tobacco leaf ecological areas is of great significance for tobacco quality control, product standardization and market value assessment. Environmental factors such as climate conditions, soil characteristics and altitude in different ecological areas will directly affect the physiological metabolism process of tobacco leaves, resulting in significant differences in their chemical compositions and physical properties. These differences are ultimately reflected in the burning characteristics, aroma styles and sensory qualities of tobacco leaves, thus determining their industrial usability and commercial value. Therefore, establishing a scientific and accurate tobacco leaf ecological area classification system is an important prerequisite for improving the quality stability of tobacco products, ensuring brand characteristics and realizing high quality and high price.

[0003] However, traditional tobacco leaf ecological area classification methods mainly rely on manual experience judgment or simple physical and chemical index detection. They not only rely on the subjective judgment of tobacco tasters, are easily interfered by individual differences and external factors, but also only focus on a few known components such as nicotine and total sugar, resulting in relatively poor consistency and accuracy of classification results. Summary of the Invention

[0004] Based on this, it is necessary to provide a systematic and high-precision method, device and computer-readable storage medium for classifying tobacco leaf ecological areas in view of the above technical problems.

[0005] In a first aspect, the present application provides a method for classifying tobacco leaf ecological areas based on metabolomics, including:

[0006] Obtaining a metabolite data set of a tobacco leaf sample set; the tobacco leaf sample set includes multiple tobacco leaf samples from each ecological area and corresponding ecological area identifiers;

[0007] Performing data preprocessing on the metabolite data set to obtain a sample feature set; using multiple feature selection methods to respectively select several optimal features from the sample feature set to form an optimal feature data set;

[0008] Training and fusing multiple classification algorithm models by using the optimal feature data set to obtain a trained tobacco leaf ecological area fusion classification model;

[0009] Obtaining the several optimal features for the tobacco leaf to be classified, and inputting them into the trained tobacco leaf ecological area fusion classification model to obtain an ecological area prediction result.

[0010] In some of these embodiments, the training and fusion of the multiple classification algorithm models using the optimal feature dataset to obtain the trained tobacco leaf ecological area fusion classification model includes:

[0011] Select multiple classification algorithm models and set the hyperparameters of each of the classification algorithm models;

[0012] Use the optimal feature dataset to train the multiple classification algorithm models respectively;

[0013] Use the model fusion method to fuse the trained multiple classification algorithm models to obtain the fused tobacco leaf ecological area fusion classification model.

[0014] In some of these embodiments, the use of the model fusion method to fuse the trained multiple classification algorithm models to obtain the fused tobacco leaf ecological area fusion classification model includes:

[0015] Obtain the class prediction results of the multiple classification algorithm models for the optimal feature dataset;

[0016] Count the number of votes for each of the class prediction results to be predicted as the final class;

[0017] Determine the class prediction result with the highest number of votes as the final prediction class of the optimal feature dataset.

[0018] In some of these embodiments, the multiple classification algorithm models include at least two of a support vector machine (SVM) model, a random forest model, and a multi-layer perceptron (MLP) model.

[0019] In some of these embodiments, the use of multiple feature selection methods to respectively select several optimal features from the sample feature set to form the optimal feature dataset includes:

[0020] Calculate the variance of the sample feature set and select several first optimal features with the largest variance in the calculation results;

[0021] Calculate the linear correlation degree between each sample feature in the sample feature set through the Pearson correlation coefficient and select several second optimal features;

[0022] Calculate the feature importance of the sample feature set through the LightGBM algorithm and select several third optimal features with the highest importance;

[0023] Take the union of the first optimal features, the second optimal features, and the third optimal features to form the optimal feature dataset.

[0024] In some of these embodiments, the evaluation of the tobacco leaf ecological area classification fusion model includes:

[0025] Predict the tobacco leaf ecological area classification fusion model based on the test data set to generate an ecological area prediction result;

[0026] Calculate the evaluation index between the ecological area prediction result and the true ecological area identifier, and the evaluation index includes at least one of accuracy, precision, recall, or F1 score;

[0027] When the evaluation index does not reach the preset value, adjust the hyperparameters of the multiple classification algorithm models and retrain the multiple classification algorithm models;

[0028] Repeat the above steps until the evaluation index reaches the preset value or the maximum number of iterations.

[0029] In some embodiments, the data preprocessing of the metabolite data set to obtain the sample feature set includes:

[0030] Clean the metabolite data set;

[0031] Perform data normalization on the data after data cleaning to obtain the sample feature set.

[0032] In some embodiments, the metabolite data set of the tobacco leaf sample set is obtained by analyzing the tobacco leaf sample set through GC-MS and LC-MS analysis platforms.

[0033] In a second aspect, the present application also provides a tobacco leaf ecological area classification device, and the device includes:

[0034] An acquisition module for acquiring a metabolite data set of tobacco leaf samples;

[0035] A feature screening module for selecting a number of optimal features to form an optimal feature data set;

[0036] A training module for respectively training and fusing multiple classification algorithm models using the optimal feature data set to obtain a trained tobacco leaf ecological area fusion classification model;

[0037] An identification module for obtaining the number of optimal features for the tobacco leaf to be classified and inputting them into the trained tobacco leaf ecological area fusion classification model to obtain an ecological area prediction result.

[0038] In a third aspect, the present application also provides a computer-readable storage medium, and a computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the following steps are implemented:

[0039] Obtain a metabolite data set of a tobacco leaf sample set; the tobacco leaf sample set includes multiple tobacco leaf samples of each ecological area and corresponding ecological area identifiers;

[0040] performing data preprocessing on the metabolite dataset to obtain a sample feature set; using multiple feature selection methods to select a number of optimal features from the sample feature set to form an optimal feature dataset;

[0041] Using the optimal feature data set to train and fuse multiple classification algorithm models respectively, to obtain a trained tobacco ecological zone fusion classification model;

[0042] The optimal features are obtained for the tobacco leaves to be classified, and are input into the trained tobacco leaf ecological zone fusion classification model to obtain ecological zone prediction results.

[0043] Compared to related technologies, the metabolomics-based tobacco ecological zone classification method, device, and medium described above utilize multiple feature selection methods to select the optimal features from the sample feature set, forming an optimal feature dataset. This allows for the selection of useful key features from different perspectives, while also achieving feature dimensionality reduction. Multiple classification algorithm models are trained and fused using this optimal feature dataset to produce a trained tobacco ecological zone fusion classification model. This model, based on different principles, allows for targeted data modeling, thereby improving classification capabilities from various perspectives. By integrating metabolomics analysis, multidimensional feature screening, and multi-model fusion techniques, this method significantly improves the objectivity and accuracy of tobacco ecological zone classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0045] Figure 1 This is a computer hardware block diagram of the tobacco ecoregion classification method according to some embodiments of the present application;

[0046] Figure 2 A flowchart of a tobacco leaf ecological zone classification method according to some embodiments of the present application;

[0047] Figure 3 The data distribution of sample points after PCA dimensionality reduction in some embodiments of the present application;

[0048] Figure 4 The data distribution of sample points after t-SNE dimensionality reduction in some embodiments of the present application;

[0049] Figure 5Flowchart of the method for screening optimal features in some embodiments of the present application;

[0050] Figure 6 Flowchart of the model training and fusion method in some embodiments of the present application;

[0051] Figure 7 Flowchart of the model fusion method in some embodiments of the present application;

[0052] Figure 8 Flowchart of the evaluation method for the fusion model in some embodiments of the present application;

[0053] Figure 9 Block diagram of the structure of the tobacco leaf ecological area classification device in some embodiments of the present application. Detailed implementation manners

[0054] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0055] A tobacco leaf ecological area classification method provided by an embodiment of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through a network. The terminal 102 can obtain a tobacco leaf metabolite dataset and send the obtained tobacco leaf metabolite dataset to the server 104. The server 104 receives the tobacco leaf metabolite dataset sent by the terminal 102, preprocesses the data to obtain a sample feature set, obtains an optimal feature dataset after feature selection, trains and fuses multiple classification algorithm models using this dataset to obtain a tobacco leaf ecological area fusion classification model, and this model can be used to distinguish the ecological area of the tobacco leaf to be classified. The data storage system can store the data that the server 104 needs to process, as well as the trained tobacco leaf ecological area fusion classification model. The data storage system can be integrated on the server 104, or can be placed in the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and Internet of Things devices. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0056] In an exemplary embodiment, as Figure 2 shown, a method for classifying the tobacco leaf ecological area is provided. Taking the method applied to a Figure 1 computer device as an example, it includes the following steps S202 to step S206. Among them:

[0057] Step S201 : obtaining a metabolite dataset of a tobacco leaf sample set, wherein the tobacco leaf sample set includes a plurality of tobacco leaf samples from various ecological zones and corresponding ecological zone identifiers.

[0058] Tobacco production has a significant impact on tobacco quality. Different ecological regions, influenced by growing environmental factors such as temperature, humidity, and altitude, can lead to variations in the content and composition of metabolites in tobacco leaves. Therefore, when selecting ecological regions for tobacco leaf samples, consider representative regions that exhibit significant differences in growing conditions, in addition to geographic location.

[0059] Specifically, the tobacco leaf sample set collection method can be to select several typical central production areas and one upper comprehensive tobacco leaf module, selecting tobacco leaves that are representative of the region and quality and meet the expected sampling requirements. Each sampling is determined using the five-point method, that is, one point is selected at the center and four corners of the target area to ensure representativeness and spatial coverage. After sampling, the samples are stored in a tobacco box. Each time, block tobacco leaves are selected from the tobacco box, the surface leaves are removed, and after mixing, the samples are sampled as the tobacco leaf sample set.

[0060] Each collected tobacco leaf sample is labeled with its ecological region. The ecological region identifier can be the number of the different ecological regions and will be used as a label for model training. The tobacco leaf samples labeled with ecological region identifiers are aggregated to form a tobacco leaf sample set. Metabolite detection is performed on the tobacco leaf sample set to obtain a metabolite dataset for the tobacco leaf sample set.

[0061] Understandably, due to the complex synergy of thousands of metabolites in tobacco leaves, metabolomics offers unique advantages in qualitative and quantitative analysis of tobacco leaf components. Metabolite data from tobacco leaf samples can be obtained using a variety of methods, and the appropriate technique can be selected based on specific resource conditions and target scope.

[0062] The metabolite dataset for the tobacco leaf sample set includes a feature set and a label set. The feature set is a two-dimensional array of structured metabolite data, with each row representing a sample and each column representing a feature. The label set is a one-dimensional array, with each value representing the ecoregion identifier for each sample. The length of the label set is the same as the number of rows in the feature set matrix.

[0063] Step S202 : performing data preprocessing on the metabolite dataset to obtain a sample feature set.

[0064] Among them, data preprocessing is an important step before model training in machine learning. It usually includes removing low-quality data in the data, normalizing the data representation, encoding labels, etc., which can improve the analysis and classification effects of the model. The methods of data preprocessing include but are not limited to steps such as data cleaning and data normalization. The preprocessed data is used as a sample feature set for subsequent model training.

[0065] Step S203: Use multiple feature selection methods to respectively select several optimal features from the sample feature set to form an optimal feature data set.

[0066] In the obtained sample feature set after preprocessing, the high-dimensional data matrix composed of all detected metabolite features is called the feature space. Usually, when the number of feature dimensions is too high, there will be a problem of dimensionality disaster in model training, which will not only lead to unreliable prediction results of the model, but also cause the consumption of computing resources to increase exponentially. By screening features with high discriminative power, the feature dimension can be reduced, thereby improving the accuracy of the model, reducing the risk of overfitting, accelerating the model training speed, and reducing the consumption of computing resources.

[0067] In this embodiment, a multi-dimensional feature selection strategy is adopted to screen the most discriminative feature subset from the original feature space, and the selection results of different feature selection methods are integrated through feature set fusion to form an optimal feature data set containing key features for subsequent model training. This process can significantly reduce the data dimension while retaining the differential metabolic characteristics of tobacco leaves in different ecological regions.

[0068] Step S204: Use the optimal feature data set to train and fuse multiple classification algorithm models respectively to obtain a trained tobacco leaf ecological region fusion classification model.

[0069] Select multiple classification algorithm models with complementary characteristics to form a basic model set. These models have their own advantages in aspects such as feature space analysis and capturing non-linear relationships, and can comprehensively mine the discriminant information in the data.

[0070] Then, use the optimal feature data set to independently train each basic classification model, adapt to the characteristics of different models through hyperparameter tuning strategies to ensure that each model reaches the optimal performance state under the same data conditions. Use the method of model fusion to fuse the above classification algorithm models, and use the model evaluation method to evaluate the ability of the fusion model, and optimize the model accordingly to improve the classification accuracy.

[0071] Step S205: Obtain the several optimal features for the tobacco leaves to be classified and input them into the trained tobacco leaf ecological region fusion classification model to obtain the ecological region prediction result.

[0072] After the fusion classification model is trained, the following standardized processing flow is performed on the tobacco leaf samples to be classified: First, the same feature preprocessing process as in the training phase is used to extract the key indicator set determined by feature selection, i.e., the optimal features, to ensure that the input data maintains the same dimensionality as the feature space used during model training. The standardized feature vectors are then fed into the trained tobacco leaf ecoregion fusion classification model to generate predictions containing ecoregion labels.

[0073] In this example, metabolomics data collection improves the utilization of effective information from tobacco leaf components. Multidimensional feature selection enhances the model's ability to process high-dimensional features. Fusion of multiple models addresses issues of poor model stability and generalization. This approach addresses the strong subjectivity, low information utilization, and insufficient model generalization capabilities of traditional methods, significantly improving the accuracy and efficiency of tobacco leaf production region resolution.

[0074] In an exemplary embodiment, the metabolite dataset of the tobacco leaf sample set is obtained by analyzing the tobacco leaf sample set using GC-MS and LC-MS analysis platforms.

[0075] It's understandable that gas chromatography-mass spectrometry (GC-MS) is an analytical platform that combines the separation capabilities of gas chromatography with the identification capabilities of mass spectrometry, making it particularly suitable for the detection of volatile and semivolatile small molecule metabolites. Liquid chromatography-mass spectrometry (LC-MS) combines the separation capabilities of liquid chromatography with the identification capabilities of mass spectrometry, making it suitable for the analysis of non-volatile, thermally labile macromolecules. Analytical methods combining GC-MS and LC-MS platforms can cover a wider range of compounds and are widely used.

[0076] Exemplarily, the extraction and analysis of tobacco leaf metabolites using GC-MS and LC-MS platforms specifically comprises the following steps:

[0077] First, pre-treat the tobacco leaf sample. Quickly freeze 100 mg of tobacco leaf tissue in liquid nitrogen and grind. Add 500 μL of 80% methanol in water containing 0.1% formic acid, vortex, and incubate on ice for 5 minutes. Centrifuge at 15,000 rpm and 4°C for 10 minutes. Collect a certain amount of the supernatant. Dilute the supernatant with mass spectrometry-grade water to a methanol content of 60%, place it in a centrifuge tube with a 0.22 μm filter, and centrifuge at 15,000 g and 4°C for 10 minutes. Collect the filtrate.

[0078] The collected filtrates were then analyzed by LC-MS and GC-MS. For quality control (QC), equal volumes of each tobacco leaf sample were mixed and used as QC samples. Blank samples, consisting of 60% methanol-water solution containing 0.1% formic acid, replaced the experimental samples and underwent the same pretreatment procedures as the experimental samples.

[0079] After the test is completed, the raw data is exported to obtain tobacco leaf metabolite information. In this example, 100 tobacco leaf samples from five ecological zones were used, and a total of 7024 metabolite features were detected, of which 3200 were identified by GC-MS and 3824 by LC-MS.

[0080] By analyzing and detecting tobacco leaf metabolites through the above-mentioned GC-MS and LC-MS analysis platforms, the metabolite components of tobacco leaves can be obtained, and the potential value of metabolomics in tobacco leaf ecological zone classification can be fully utilized.

[0081] In some embodiments, performing data preprocessing on a metabolite dataset to obtain a sample feature set includes the following steps:

[0082] Step S301: performing data cleaning on the metabolite dataset.

[0083] As you can understand, the Z-score normalization method is a commonly used method in statistics to normalize data. It can measure the degree of deviation of a data point from the mean of the data set, and is therefore often used for data outlier detection. In this method, the Z-Score calculation formula is as follows:

[0084]

[0085] in, For a feature dataset The value of the sample, is the mean of the data set, is the standard deviation of the data set.

[0086] Specifically, each data in the metabolite dataset is taken as a sample data, and for each feature in the metabolite dataset, its mean in all sample data is calculated. and standard deviation Calculate the Z-score for each sample's feature value. Setting a threshold of 3 means that under a normal distribution, 99.7% of data points will fall within ±3 standard deviations of the mean. Therefore, the probability of a data point exceeding 3 standard deviations being considered an outlier is very small. After calculating the Z-score values for all sample features, remove any samples with a value greater than 3 to obtain the cleaned data.

[0087] Step S302: Perform data normalization on the data after data cleaning to obtain the sample feature set.

[0088] It can be understood that since the data has different dimensions and distribution value ranges on different feature indicators, if the data is not normalized, gradient disappearance or strong dependence on certain features will occur during the model training process. Therefore, it is necessary to scale all features into a unified interval.

[0089] In a specific embodiment, the MinMax normalization method is used, and its formula is as follows:

[0090]

[0091] Where, is the value of the th sample in a certain feature dataset, is the minimum value of this feature in the dataset, is the maximum value of this feature in the dataset, is the value after normalization, and this value is usually between [0, 1].

[0092] Specifically, the MinMax normalization process is as follows: First, traverse each feature to find its minimum value and maximum value in the dataset; for each feature of each sample value , use the formula to calculate the standardized value ; replace the original data with the standardized value .

[0093] Through the normalization operation, the data values can be scaled to the same range, which is convenient for feature comparison and model training.

[0094] It can be understood that through exploratory data analysis, the main features of the dataset can be understood and summarized through visualization and statistical methods, and potential patterns, trends, and outliers can be discovered. Moreover, by removing irrelevant features or performing dimensionality reduction on the data, the dimension of the feature space can be reduced, the amount of calculation can be reduced, and at the same time, the model can be helped to focus on the features highly correlated with the target variable. In an exemplary tobacco leaf metabolite dataset, the number of features of the metabolite components exceeds 7000, while the number of samples is only 96. Therefore, the principal component analysis PCA method can be used without specifying the number of principal components, and all principal components are calculated using PCA to perform dimensionality reduction on the metabolite features.

[0095] Set the number of principal components equal to the number of features of the input data, and thus the number of principal components with a cumulative variance explanation rate reaching 95% can be obtained. On this tobacco leaf metabolite sample dataset, the number of principal components is 45. Take the two principal components with the highest variance explanation rate to draw a scatter plot. After PCA dimensionality reduction, the data distribution of the sample points is as Figure 3 shown. Among them, the abscissa is the score of the first principal component, that is, the projection value of the data point in the direction of the largest variance (the first eigenvector). Similarly, the ordinate is the score of the second principal component. The coordinate value represents the position of the sample point in the direction of this principal component, and the numerical size reflects the contribution of the original feature to the principal component.

[0096] It can be seen that by retaining the two principal components with the highest variance explanation rate, effective classification of different ecological regions can be carried out, and the spatial differences between them are very obvious. In order to further verify the effectiveness of dimensionality reduction, t-SNE is used for verification again.

[0097] t-SNE (t-distributed Stochastic Neighbor Embedding) is a non-linear method for dimensionality reduction and visualization of high-dimensional data. It is particularly suitable for mapping high-dimensional data into two-dimensional or three-dimensional space for visualization. The main purpose of t-SNE is to try to retain the local structure and similarity of the data while reducing the dimension. After t-SNE dimensionality reduction, the data distribution of the sample points is as Figure 4 shown. Among them, the abscissa is the embedding position of the sample point in the first dimension of t-SNE, representing the embedding position of the sample point in the second dimension of t-SNE.

[0098] It can be seen that the sample points after dimensionality reduction by different methods have obvious distinguishability on the two-dimensional coordinates. Therefore, the original data samples can be classified by a classification model. However, the problem with methods such as PCA is that the features after dimensionality reduction cannot correspond one by one to the original features, and it is also impossible to intuitively judge the index interval of metabolites. Therefore, it is necessary to screen the features in the original data through feature selection.

[0099] In some of these embodiments, Figure 5 is a flowchart for screening the optimal features in some embodiments of the present application, as Figure 5 shown, and this process includes the following steps:

[0100] Step S401, calculate the variance of the sample feature set, and select a number of first optimal features with the largest variance in the calculation results.

[0101] As you can understand, variance-based feature selection is a fast and effective feature selection method used to remove features from a dataset that have little impact on the target variable. The core idea is that features with low variance contain little information and may contribute little to the model's predictions, so they can be removed. This is often because low-variance features have little variation between samples, resulting in high levels of information redundancy, making it difficult to effectively distinguish data categories or regression targets.

[0102] Specifically, the sample variance calculation formula is:

[0103]

[0104] in, is a dataset of a certain feature, This is the feature dataset The value of the sample, Characterized by The mean of .

[0105] Calculate the variance of all samples and select the features with the largest variance as the first optimal features. In this embodiment, 221 features are selected as examples. Some of the features and their variances are shown in Table 1:

[0106] Table 1

[0107] Feature Name variance Lucuminic acid 0.144413 p-Aminophenazone 0.141298 N-(3,5-Dichlorophenyl)-2-hydroxysuccinamic acid 0.134267 7-O-Acetylhorminone 0.133427 4-Hydroxybenzoic acid glucoside 0.131983 Methyl prednisolonate 0.131924 13,14-dihydro-15-keto-PGA2 0.129236 8-C-Methylquercetin 3-methylether 0.127715 Val-Gly-Val-Ala-Pro-Gly 0.127643 Niazimicin 0.127626

[0108] Step S402: Calculate the linear correlation between each sample feature in the sample feature set using the Pearson correlation coefficient, and select several second optimal features.

[0109] It can be understood that correlation-based feature selection identifies and selects the most important features by analyzing the correlation between features and the target variable. The basic idea is to select features that are closely related to the target variable while eliminating features with low correlation. The Pearson Correlation Coefficient is a statistic used to measure the degree of linear correlation between two variables. The range is [−1, 1], where 1 indicates a perfect positive correlation, meaning the two variables are in direct proportion; 0 indicates no correlation, meaning there is no linear relationship; and -1 indicates a perfect negative correlation, meaning the two variables are in inverse proportion.

[0110] Specifically, the Pearson correlation coefficient calculation formula between features is:

[0111]

[0112] in are two features in the sample feature set, for The covariance of , which indicates the extent to which they vary together, and They are respectively The variance of .

[0113] The Pearson correlation coefficient is calculated for the sample feature set, and several features with the smallest correlation among the sample features are selected as the second optimal features. In this embodiment, 2640 features are exemplarily selected as the second optimal features.

[0114] In step S403, the feature importance of the sample feature set is calculated using the LightGBM algorithm, and several third optimal features with the highest importance are selected.

[0115] As you can understand, LightGBM is an efficient gradient boosting framework based on decision trees. It can calculate feature importance and identify key discriminant features during training. By using the feature importance index of the LightGBM model, we can filter out features that contribute significantly to the model's prediction performance, simplify the model, reduce computing resources, and improve the model's generalization ability.

[0116] Specifically, a LightGBM model is trained using a sample feature set. LightGBM provides three feature importance metrics: Split (default): This measures the number of times a feature is used for splits across all trees; Gain: This measures the average information gain of a feature across all trees; and Cover: This measures the number of samples covered by a feature across all trees. You can choose a different metric based on your needs. For example, Gain is selected in this example.

[0117] After the LightGBM model training is completed, the features with importance scores greater than 0 are retained. These features contribute to the model prediction. Several key features with the highest importance scores are selected as the third best features. In this embodiment, 197 key features are selected as the third best features. Some of the most important features are shown in Table 2:

[0118] Table 2

[0119] Feature Importance Irene 6 Metyrapone 6 Mycomycin 5 Tyrosyl_Lysine 5 Sophoraflavone_B 5 Stigmasta_4_22_dien_3_one 4 9_OxoODE 4 4_Hydroxybenzoic_acid_glucoside 4 6_Hydroxykaempferol_3_rutinoside 4 Cer_t18_0_PGF2alpha_ 4 DGMG_16_0_0_0_ 4 Alanine_lactate_pyruvate 4

[0120] Step S404: The first optimal feature, the second optimal feature, and the third optimal feature are combined to form the optimal feature data set.

[0121] After the above feature selection steps, the filtered features are subjected to a union operation to obtain a final feature set for use, which is then used to train the classification model. In this embodiment, a total of 2667 features are selected from 7024 features.

[0122] In some of these embodiments, Figure 6 is a flow chart of the model training and fusion method of some embodiments of the present application, such as Figure 6 As shown, the process includes the following steps:

[0123] Step S501: Select multiple classification algorithm models and set hyperparameters of each classification algorithm model.

[0124] Among them, the classification algorithm model is a machine learning model used to divide data into different categories. Hyperparameters are parameters that need to be manually set before training the machine learning model. They control the structure and learning process of the model and affect the final training effect of the machine learning model.

[0125] It's understandable that different models have different hyperparameters, so the selection of each hyperparameter can be based on different priorities. When initializing the model, you can set the initial hyperparameters based on experience. During the subsequent classification algorithm model training process, you can optimize the model by adjusting the hyperparameters.

[0126] Exemplarily, the classification algorithm model may select at least two of a support vector machine (SVM) model, a random forest model, and a multi-layer perceptron (MLP) model.

[0127] Step S502: Use the optimal feature data set to train the multiple classification algorithm models respectively.

[0128] Before training the classification algorithm, the optimal feature dataset was divided into training and test sets according to a specific ratio. Stratified sampling was used to ensure that the sample ratio for each ecological zone was consistent across the training and test sets. For example, the training set to test set ratio was 2:8.

[0129] Specifically, when training a classification algorithm model, a grid search method is used to optimize hyperparameters. This method includes: defining a hyperparameter grid for each classification algorithm model; traversing all possible hyperparameter combinations on the hyperparameter grid; and using cross-validation to verify the accuracy and stability of sample classification. The best performing hyperparameter combination is selected as the optimal hyperparameter for the prediction model.

[0130] Step S503: Use a model fusion method to fuse the multiple classification algorithm models that have been trained to obtain a fused tobacco ecological zone fusion classification model.

[0131] Among them, model fusion is an ensemble learning method in machine learning. Since different models may be good at capturing different features of data, by combining the prediction results of multiple independently trained base classifiers, even if the classification results of a certain type of model are poor, other models can still provide reliable predictions. Therefore, it can improve the performance and robustness of the overall model. By integrating the advantages of different models and making up for the limitations of a single model, more stable and accurate prediction results can be obtained.

[0132] In an implementable manner, the classification algorithm model uses a support vector machine (SVM) model, a random forest (RF) model, and a multi-layer perceptron (MLP) model. Before training the model, an optimization table for the hyperparameters of each model is set respectively. Specifically, for the SVM model, three hyperparameters, namely the regularization parameter (C), the kernel function, and Gamma, are selected; for the RF model, the number of decision trees (n_estimators), the maximum depth (max_depth), and the minimum number of samples per leaf node (min_samples_leaf) are selected; for the MLP model, the number and layers of neurons in the hidden layer (hidden_layer_sizes), the activation function type (activation) that controls the activation function of each layer of neurons, the learning rate (learning_rate), and the initial value of the learning rate (learning_rate_init) are selected. The 10-fold cross-validation and grid search method is adopted to train the three classification algorithm models respectively using the training set, select the model with the optimal output result, and obtain the optimal hyperparameters. Exemplarily, the hyperparameter selection method can be as follows:

[0133] (1) SVM hyperparameter selection:

[0134] Regularization parameter C: C controls the tolerance of the model to misclassification, that is, it weighs maximizing the margin and the penalty for misclassification. A smaller C value encourages the model to choose a larger margin, that is, allows more misclassifications; a larger C value will try to minimize misclassifications as much as possible. In this embodiment, [0.001, 0.01, 0.1, 1, 10, 100] is traversed for optimal selection, and finally C = 1 is obtained.

[0135] Kernel function: The kernel function determines how to map low-dimensional input features to a high-dimensional feature space. Different kernel functions are suitable for different data distributions. In this embodiment, [linear, poly, rbf, sigmoid] is traversed for optimal selection, and finally the linear and rbf kernel functions are determined to be used.

[0136] Gamma: Gamma determines the influence of a single sample. Larger gamma values focus the model on a localized set of training samples, resulting in a more complex model. Smaller gamma values focus the model on the entire dataset, leading to smoother decision boundaries. In this example, we iterate through [0.001, 0.01, 0.1, 1, 10, 100] to find the optimal value, ultimately using gamma = 0.01.

[0137] (2) Random forest hyperparameter selection:

[0138] n_estimators: This determines the number of trees in the forest. More trees generally improve performance but increase training time. In this example, we selected [10, 20, 50, 100, 200] to find the optimal parameter, ultimately using n_estimators = 10.

[0139] max_depth: This parameter controls the maximum depth of each tree to avoid overfitting. A larger depth may lead to overfitting, while a smaller depth may lead to underfitting. In this example, we traverse [None, 10, 20, 30, 50] to find the optimal value, where None indicates no depth limit. Finally, we use max_depth=None.

[0140] min_samples_leaf: This value controls the minimum number of samples per leaf node. Larger values reduce model complexity. In this example, we traverse [1, 2, 5, 10] to find the optimal value. Finally, we use min_sample_leaf=5.

[0141] (3) MLP hyperparameter selection:

[0142] hidden_layer_sizes: The number of hidden layer neurons and the number of layers. For two and three hidden layers, the optimal number of neurons is selected from [32, 64, 128, 256]. Ultimately, [128, 64, 32] is used, resulting in three layers.

[0143] Activation: The activation function controls the type of activation function for each layer of neurons. It iterates through [identity, logistic, tanh, relu] to select the optimal parameters, and finally uses activation=relu.

[0144] learning_rate: The learning rate control parameter. In this embodiment, [constant, invscaling, adaptive] is traversed to select the optimal parameter, and finally learning_rate = adaptive is used.

[0145] learning_rate_init: The initial learning rate controls the initial value of the learning rate. A larger learning rate may lead to unstable optimization, while a smaller learning rate may lead to slow convergence. Traverse [0.0001, 0.001, 0.01, 0.1] to select the optimal parameter, and finally learning_rate_init = 0.001 is used.

[0146] After the classification models are trained respectively, the model fusion method is used to fuse the above multiple classification models. The finally obtained fusion model and the optimal hyperparameters are saved as a model file at the same time. The model file is stored in the storage medium.

[0147] It can be understood that in the model fusion step, the fusion methods include but are not limited to Voting, Soft Voting, or Stacking. In one of the embodiments, Voting is used for model fusion. Figure 7 is a flowchart of the model fusion method of an embodiment of the present application, as Figure 7 shown, this process includes the following steps:

[0148] Step S601, obtain the class prediction results of the multiple classification algorithm models for the optimal feature dataset.

[0149] Specifically, for the tobacco leaf samples to be classified, the optimal feature dataset is input, and the format is a two-dimensional array. The trained models are called separately for independent prediction. Among them, the predicted class is the prediction result of each model, and the data format is a one-dimensional array, and it is kept consistent with the format of the true ecological area class label of the tobacco leaf sample.

[0150] In this embodiment, the SVM model, RF model, and MLP model are used as the basic classification algorithm models. The SVM model outputs the predicted class as y_svm, the RF model outputs the predicted class as y_rf, and the MLP model outputs the predicted class as y_mlp. {y_svm, y_rf, y_mlp} generated in this step will be used as the input of step S602, and the final classification result is generated through the voting mechanism. At the same time, the original prediction records are retained for quality traceability.

[0151] Step S602, count the number of votes for each type of class prediction result to be predicted as the final class.

[0152] First, a voting counter is established and the number of votes in each ecological zone category is initialized to 0. In this embodiment, tobacco leaves are divided into five ecological zones as an example. The initial states are: Southwest production area (XN): 0, Huanghuai production area (HH): 0, middle and lower reaches of the Yangtze River (CJ): 0, Southeast production area (DN): 0, and imported tobacco leaves (JK): 0.

[0153] Next, the predictions of each classification model are tallied and output as ecoregion labels. Each classification model's predictions are processed sequentially. For example, if the prediction output y_svm from the SVM model is "HH," the value corresponding to "HH" in the vote counter is incremented by 1. If the prediction output y_rf from the random forest model is "XN," the number of votes for "XN" is incremented by 1. If the prediction output y_mlp from the MLP model is "HH," the number of votes for "HH" is incremented by 1 again. At this point, the vote counter status is updated to: HH: 2, XN: 1, and other categories: 0.

[0154] In the voting counter, the category with the highest vote count is compared across all categories to determine the category with the highest vote count. In the previous example, the highest value is 2 (corresponding to category "HH"). A check is performed to see if there are multiple categories with this highest value. If there is only one, that category is determined as the final result. If multiple categories have the same highest vote count (e.g., HH: 2, XN: 2), the process proceeds to tie-vote processing.

[0155] The ticket processing process can choose any of the following methods according to the preset strategy:

[0156] (1) Model priority method: According to the pre-defined model decision weights (such as RF>SVM>MLP), the prediction results of the model with higher weight are adopted first.

[0157] (2) Confidence comparison method: Extract the maximum predicted probability of each tie category in all models and select the category with the higher probability value. For example, compare max(P(HH)) and max(P(XN)).

[0158] Step S603: Determine the category prediction result that obtains the highest number of votes as the final prediction category of the optimal feature dataset.

[0159] By traversing each sample in the optimal feature dataset, reading the voting counter status generated in step S602, and taking the category with the most votes as the final prediction result of each sample, the final prediction category of the optimal feature dataset can be obtained.

[0160] In this embodiment, a voting fusion mechanism is established through a multi-model collaborative decision-making system, which overcomes the problem of limited ability to analyze complex metabolomics features when relying on a single classification algorithm and improves the accuracy of model prediction.

[0161] In some of these embodiments, Figure 8 is a flowchart for evaluating the fusion model in some embodiments of the present application, as Figure 8 shown. The process includes the following steps:

[0162] Step S701: Predict the tobacco leaf ecological area classification fusion model based on the test data set to generate an ecological area prediction result.

[0163] Among them, the test data set refers to an independent data subset reserved from the original tobacco leaf samples, usually accounting for 20%-30% of the total sample volume, and contains metabolomic feature data with known ecological area labels, which is used to objectively evaluate the model performance. Its data format is the same as that of the training set. The ecological area prediction result is a set of classification labels output by the fusion model for the test set samples, and the format is {"sample ID": "predicted category"}, for example, {"T001": "HH", "T002": "XN"}.

[0164] Specifically, load the preprocessed test data set and ensure that the feature dimensions are exactly the same as those of the training set. Input the test set into the trained fusion model to obtain the ecological area prediction result. In addition, the confidence score of each sample prediction result can also be obtained.

[0165] Step S702: Calculate the evaluation metrics between the ecological area prediction result and the true ecological area identifier, and the evaluation metrics include at least one of accuracy, precision, recall, or F1 score.

[0166] After obtaining the prediction result of the fusion model for the test data set, the core metric combination can be selected according to the business requirements of tobacco leaf classification.

[0167] Specifically, select accuracy as the global metric, and precision, recall, and F1 score as the category-level metrics. And visualize the details of the classification result by constructing a multi-dimensional confusion matrix. In one embodiment, the classification results of different models on the ecological area after selecting the optimal hyperparameters are shown in Table 3.

[0168] Table 3

[0169] Model Name Accuracy Precision Recall F1-score SVM_linear 1 1 1 1 SVM_rbf 1 1 1 1 RF 1 1 1 1 MLP 1 1 1 1

[0170] Step S703: When the evaluation metric does not reach the preset value, adjust the hyperparameters of the multiple classification algorithm models and retrain the multiple classification algorithm models.

[0171] If the evaluation metric in step S702 does not reach the preset threshold, the model can be adjusted accordingly using different evaluation metric values in the confusion matrix. For example, the model can be adjusted to identify frequently misclassified category pairs, such as when the Huanghuai region "HH" is often misclassified as the Southwest region "XN".

[0172] Additionally, you can optimize the model specifically based on the problem type. For example, if overfitting occurs, you can enhance the regularization of the base classification model, such as by increasing the C value of the SVM model or adding a dropout layer to the MLP. If underfitting occurs, you can improve the expressiveness of features, such as by increasing the number of neurons in the hidden layer of the MLP.

[0173] Step S704: Repeat the above steps until the evaluation index reaches a preset value or reaches a maximum number of iterations.

[0174] Training termination conditions typically include: all key metrics reaching preset values, reaching the maximum number of iterations, or performance degradation (abnormal termination) after multiple consecutive iterations. The maximum number of iterations is a safety upper limit set to prevent infinite loops in model training.

[0175] Based on the same inventive concept, the present application also provides an apparatus for implementing the aforementioned metabolomics-based tobacco ecological zone classification method. The solution provided by this apparatus is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more tobacco ecological zone classification apparatus embodiments provided below can be found in the aforementioned definitions of the metabolomics-based tobacco ecological zone classification method and will not be further elaborated here.

[0176] In an exemplary embodiment, Figure 9 As shown, a tobacco ecological zone classification device is provided, comprising: an acquisition module, a feature screening module, a training module and a recognition module, wherein:

[0177] An acquisition module 1001 is used to acquire a metabolite dataset of a tobacco leaf sample;

[0178] A feature screening module 1002 is used to select a number of optimal features to form an optimal feature data set;

[0179] A training module 1003 is configured to use the optimal feature data set to train and fuse multiple classification algorithm models to obtain a trained tobacco ecological zone fusion classification model;

[0180] The identification module 1004 is used to obtain the optimal features of the tobacco leaves to be classified, and input them into the trained tobacco ecological zone fusion classification model to obtain ecological zone prediction results.

[0181] In one embodiment, the feature screening module 1002 is further configured to calculate the variance of the sample feature set, and select a number of first optimal features with the largest variance in the calculation results; calculate the linear correlation degree between each sample feature in the sample feature set through the Pearson correlation coefficient, and select a number of second optimal features; calculate the feature importance of the sample feature set through the LightGBM algorithm, and select a number of third optimal features with the highest importance; take the union of the first optimal features, the second optimal features, and the third optimal features to form the optimal feature data set.

[0182] In one embodiment, the training module 1003 is further configured to select multiple classification algorithm models, and set the hyperparameters of each of the classification algorithm models; use the optimal feature data set to train the multiple classification algorithm models respectively; use the model fusion method to fuse the trained multiple classification algorithm models to obtain a fused tobacco leaf ecological area fusion classification model.

[0183] In one embodiment, the training module 1003 is further configured to obtain the class prediction results of the multiple classification algorithm models for the optimal feature data set; count the number of votes for each class prediction result to be predicted as the final class; determine the class prediction result with the highest number of votes as the final prediction class of the optimal feature data set.

[0184] In one embodiment, the training module 1003 is further configured to predict the tobacco leaf ecological area classification fusion model based on the test data set to generate an ecological area prediction result; calculate an evaluation index between the ecological area prediction result and the true ecological area identifier, where the evaluation index includes at least one of accuracy, precision, recall, or F1 score; when the evaluation index does not reach the preset value, adjust the hyperparameters of the multiple classification algorithm models, and retrain the multiple classification algorithm models; repeat the above steps until the evaluation index reaches the preset value or the maximum number of iterations is reached.

[0185] Each module in the above tobacco leaf ecological area classification device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0186] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the tobacco leaf ecological area classification method in the above embodiment are implemented.

[0187] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0188] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0189] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for classifying tobacco leaf ecological regions based on metabolomics, characterized in that, The method includes: Obtaining a metabolite dataset of a tobacco leaf sample set; the tobacco leaf sample set includes multiple tobacco leaf samples from each ecological region and corresponding ecological region identifiers; Performing data preprocessing on the metabolite dataset to obtain a sample feature set; Using multiple feature selection methods to separately select several optimal features from the sample feature set to form an optimal feature dataset; Using the optimal feature dataset to separately train and fuse multiple classification algorithm models to obtain a trained tobacco leaf ecological region fusion classification model; Obtaining the several optimal features for the tobacco leaf to be classified and inputting them into the trained tobacco leaf ecological region fusion classification model to obtain an ecological region prediction result.

2. The method according to claim 1, characterized in that, The using the optimal feature dataset to separately train and fuse multiple classification algorithm models to obtain a trained tobacco leaf ecological region fusion classification model includes: Selecting multiple classification algorithm models and setting hyperparameters of each classification algorithm model; Using the optimal feature dataset to separately train the multiple classification algorithm models; Using a model fusion method to fuse the trained multiple classification algorithm models to obtain a fused tobacco leaf ecological region fusion classification model.

3. The method according to claim 2, wherein The using a model fusion method to fuse the trained multiple classification algorithm models to obtain a fused tobacco leaf ecological region fusion classification model includes: Obtaining the class prediction results of the multiple classification algorithm models for the optimal feature dataset; Counting the number of votes for each class prediction result to be predicted as the final class; Determining the class prediction result with the highest number of votes as the final prediction class of the optimal feature dataset.

4. The method according to claim 2, characterized in that, The multiple classification algorithm models include at least two of a support vector machine (SVM) model, a random forest model, and a multi-layer perceptron (MLP) model.

5. The method according to claim 1, wherein The using multiple feature selection methods to separately select several optimal features from the sample feature set to form an optimal feature dataset includes: Calculating the variance of the sample feature set and selecting several first optimal features with the largest variance in the calculation results; Calculating the linear correlation degree between each sample feature in the sample feature set through the Pearson correlation coefficient and selecting several second optimal features; Calculating the feature importance of the sample feature set through the LightGBM algorithm and selecting several third optimal features with the highest importance; Taking the union of the first optimal features, the second optimal features, and the third optimal features to form the optimal feature dataset.

6. The method according to claim 1, characterized in that The evaluating the tobacco leaf ecological region classification fusion model includes: Performing prediction on the tobacco leaf ecological region classification fusion model based on a test dataset to generate an ecological region prediction result; Calculating an evaluation index between the ecological region prediction result and the true ecological region identifier, where the evaluation index includes at least one of accuracy, precision, recall, or F1 score; When the evaluation index does not reach a preset value, adjusting the hyperparameters of the multiple classification algorithm models and retraining the multiple classification algorithm models; Repeating the above steps until the evaluation index reaches the preset value or reaches the maximum number of iterations.

7. The method according to claim 1, characterized in that, The performing data preprocessing on the metabolite dataset to obtain a sample feature set includes: Clean the metabolite dataset; Perform data normalization on the data after cleaning to obtain the sample feature set.

8. The method according to claim 1, characterized in that, The metabolite dataset of the tobacco leaf sample set is obtained by analyzing the tobacco leaf sample set through GC-MS and LC-MS analysis platforms.

9. An ecological area classification device for tobacco leaves, characterized in that, The device includes: An acquisition module for acquiring a metabolite dataset of tobacco leaf samples; A feature screening module for selecting a number of optimal features to form an optimal feature dataset; A training module for training and fusing multiple classification algorithm models respectively using the optimal feature dataset to obtain a trained tobacco leaf ecological area fusion classification model; An identification module for obtaining the number of optimal features for the tobacco leaf to be classified and inputting them into the trained tobacco leaf ecological area fusion classification model to obtain an ecological area prediction result.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the tobacco leaf ecological area classification method according to any one of claims 1 to 8 are implemented.