A method and system for classifying pear varieties based on minimal feature set
By selecting a pear variety classification method based on the smallest feature component set, and utilizing principal component analysis and a random forest model, efficient, low-cost, and accurate classification of pear varieties is achieved, solving the problems of high cost and low efficiency caused by multiple detection indicators in existing technologies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF QUALITY STANDARD & TESTING TECH FOR AGRO PROD OF CAAS
- Filing Date
- 2026-02-11
- Publication Date
- 2026-05-26
AI Technical Summary
Existing pear variety classification methods rely on a large number of physicochemical indicators, resulting in high testing costs, long testing times, and susceptibility to human factors, making it difficult to achieve efficient and low-cost accurate identification.
By employing a method based on the minimum feature set, key differential components are screened out through principal component analysis and orthogonal partial least squares discriminant analysis. A classification model is then constructed by combining this with a random forest model, achieving 100% classification accuracy with only 4 feature components to be detected.
It significantly shortens testing time and reduces costs, while improving the accuracy and efficiency of pear variety classification, providing technical reference for the classification of other agricultural products.
Smart Images

Figure CN122087604A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural product classification technology, specifically to a method and system for classifying pear varieties based on a minimum set of feature components. Background Technology
[0002] Pears are one of the most widely cultivated fruits in my country, and different varieties exhibit significant differences in nutritional value, taste, and market value. In recent years, with the improvement of living standards, pear production and sales have been steadily increasing. Therefore, establishing accurate and efficient methods for identifying pear varieties is of great significance for ensuring product quality, maintaining market order, preventing the sale of inferior products, and achieving varietal-specific deep processing.
[0003] In existing technologies, especially chemometric methods, most still rely on a large number (usually more than 10) of physicochemical indicators to construct pear variety classification models, making it difficult to determine the minimum set of key indicators required while ensuring extremely high accuracy. Therefore, there is an urgent need for a method that can accurately screen out the minimum set of feature components to achieve high-efficiency and low-cost accurate identification.
[0004] Currently, the methods for identifying pear varieties can be mainly divided into the following categories: The first category is traditional morphological methods, which rely on professional technicians to subjectively identify the fruit's appearance, shape, color, size, and other phenotypic characteristics. While this method is simple and direct, it is easily affected by subjective human factors, fruit maturity, and environmental conditions, making it difficult to guarantee the accuracy and consistency of the identification results. Furthermore, it cannot meet the identification needs of processed products.
[0005] The second category is molecular biology methods, such as DNA molecular marker technology. This method identifies species by detecting variety-specific genetic information and has high accuracy. However, its operation is complex, requiring specialized molecular laboratories, expensive reagents and consumables, and a long testing cycle, making it difficult to promote and apply on a large scale in grassroots quality inspection departments or enterprises.
[0006] The third category comprises methods combining instrumental detection and chemometrics, which is currently a hot research topic. This type of method typically utilizes techniques such as high-performance liquid chromatography (HPLC) and gas chromatography-mass spectrometry (GC-MS) to determine the content of multiple chemical components in a sample, and then combines these with pattern recognition algorithms (such as principal component analysis (PCA) and partial least squares discriminant analysis (OPLS-DA)) to establish a classification model. While this method is more objective and scientific than the previous two, it still has significant limitations: most studies tend to measure as many indicators as possible to pursue high precision, resulting in high data dimensionality, expensive detection costs, and cumbersome analytical procedures. Many models use all or a large number (usually more than 10) of feature indicators, which inevitably contain a large amount of redundant or irrelevant information. This not only increases unnecessary detection costs and time but may also lead to model overfitting, affecting its generalization ability and robustness in practical applications. It is worth noting that the chemical component determinations involved in the current third category of methods (such as soluble solids, titratable acids, vitamin C, etc.) are mostly based on national standard methods, which have good standardization and operability, providing a reliable data foundation for this study.
[0007] Therefore, there is an urgent need in this field for a pear variety identification method that can significantly reduce the number of required testing indicators while maintaining the reliability of national standard methods, and simultaneously achieve extremely high classification accuracy. If this method can be successfully implemented, by identifying a set of minimal but highly discriminative core feature components and constructing a simpler and more efficient classification model based on this method, not only can the economy and applicability of agricultural product classification and identification be improved, but the accuracy of classification results can also be guaranteed, greatly promoting the development of the industry. Summary of the Invention
[0008] In view of this, the present invention provides a pear variety classification method and system based on a minimum feature component set. By reducing the number of indicator data to be detected for pear variety classification and identification, the core indicator data is accurately screened from a large amount of initial indicator data. This achieves the same 100% classification accuracy as detecting 14 indicators by only detecting the content of 4 feature components of the pear to be classified, which greatly improves sorting efficiency and reduces costs, providing an efficient and reliable method for pear variety classification.
[0009] To achieve the above objectives, the present invention adopts the following technical solution: A pear variety classification method based on a minimal feature set includes: S1. Data acquisition steps: Obtain raw physicochemical index data of multiple characteristic components of pear samples from multiple varieties; S2. Data preprocessing step: Standardize the raw physicochemical index data obtained in S1 to obtain a standardized dataset; S3. Feature selection steps: Perform principal component analysis (PCA) on the standardized dataset. Verify the separability of different pear varieties based on the clustering results of the PCA score graph. After confirming separability, use the orthogonal partial least squares discriminant analysis (OPLS-DA) model to perform multiple pairwise comparisons on the standardized dataset. Select the components whose variable importance projection (VIP) value is greater than 1 in all comparisons and take their union. After summing and sorting, select the components whose VIP sum is greater than 2 to form the key difference component subset. S4. Model building steps: Based on the random forest model, the key difference component sets are divided into components and sorted by importance, and then input into the random forest model for training. Cross-validation is used to determine the optimal number of features required to achieve the highest classification accuracy, thus obtaining the minimum feature component set; the final random forest model is built based on the minimum feature component set. S5. Classification Application Steps: Input the content of all components in the minimum feature component set of the pear sample to be tested into the final random forest model obtained in S4, and output the variety of the pear sample to be tested.
[0010] In the above method, optionally, the number of pear samples in S1 is four, and the characteristic components include: soluble solids, titratable acid, soluble sugar, solid-acid ratio, vitamin C, vitamin B1, vitamin B2, vitamin B6, niacin, niacinamide, total phenols in peel, total phenols in pulp, flavonoids in peel, and flavonoids in pulp.
[0011] In the above method, optionally, the formula for standardization in S2 is: ; in, Z The standardized value. X These are the original values of the original physicochemical index data. μ This represents the mean value corresponding to the original physicochemical index data. σ This represents the standard deviation of the original physicochemical index data.
[0012] Optionally, in the above method, step S3, verifying the separability of different pear varieties, includes: Principal component analysis (PCA) was performed on the standardized dataset using the statistical analysis software SIMCA to obtain the PCA score map. The separability of the original physicochemical index data for the pear samples was verified based on the PCA score map.
[0013] Optionally, in the above method, the random forest model training process in S4 includes: S41. Importance ranking steps: Rank the key difference component subsets by importance to obtain the ranked key difference component subsets {1,2,...,m,...N}, arranged from high to low importance, where m represents the m-th important feature component, N≥m, and m and N are both positive integers; S42. Validation and evaluation steps: Input the first n feature components of the key difference component subset into the random forest model for training in turn. Evaluate the accuracy of the current random forest model based on leave-one-out cross-validation to obtain the accuracy of the nth random forest model, where n≤N and n is a positive integer. S43. Accuracy Judgment Steps: Judge the current random forest model based on the accuracy of the nth random forest model and the preset accuracy: If the accuracy of the nth random forest model is equal to the preset accuracy, then take the first difference component to the nth difference component as the minimum feature component set, determine that the current random forest model is a trained random forest model, and stop training; if the accuracy of the current random forest model is less than the preset accuracy, then let n=n+1, return to S42, and evaluate the accuracy of the current random forest model; the preset accuracy is 100%.
[0014] Of the methods described above, the nth random forest model in S42 has the highest accuracy of 100%.
[0015] In the above method, optionally, the minimum feature component set contains 4 feature components.
[0016] In the above method, the four optional characteristic components are: solid-acid ratio, titratable acid, vitamin B6, and vitamin B2.
[0017] Based on the above methods, the present invention also discloses a pear variety classification system based on a minimum feature component set, used to implement a pear variety classification method based on a minimum feature component set as described in any of the above methods, including a data acquisition module, a data preprocessing module, a feature screening module, a model building module, and a classification application module connected in sequence; The data acquisition module is used to acquire raw physicochemical index data models of multiple characteristic components of pear samples from multiple varieties. The data preprocessing module is used to standardize the raw physicochemical index data obtained from the data acquisition module to obtain a standardized dataset; The feature selection module is used to perform principal component analysis (PCA) on the standardized dataset. The clustering results of the PCA score map are used to verify the separability of different pear varieties. After confirming separability, the orthogonal partial least squares discriminant analysis (OPLS-DA) model is used to perform multiple pairwise comparisons on the standardized dataset. Components with variable importance projection (VIP) values greater than 1 in all comparisons are selected and their union is taken. After summing and sorting, the components with a total VIP value greater than 2 are selected to form the key difference component subset. The model building module is used to sort the subsets by importance based on the random forest model and input them into the random forest model for training. Leave-one-out cross-validation is used to determine the optimal number of features required to achieve the highest classification accuracy, thus obtaining the minimum feature set. The final random forest model is then built based on the minimum feature set. The classification application module is used to input the content of all components in the minimum feature component set of the pear sample to be tested into the final random forest model obtained in the model building module, and output the variety of the pear sample to be tested.
[0018] In the above system, optionally, the model building module includes an importance ranking module, an accuracy evaluation module, and a judgment module connected in sequence; The importance ranking module is used to rank the subsets of key difference components by importance, resulting in a ranked subset of key difference components {1,2,...,m,...N}, arranged from high to low importance, where m represents the m-th important feature component, N≥m, and m and N are both positive integers; The accuracy evaluation module is used to sequentially input the first n feature components of the key difference component subset into the random forest model for training, evaluate the accuracy of the current random forest model based on leave-one-out cross-validation, and obtain the accuracy of the nth random forest model, where n≤N and n is a positive integer; The judgment module is used to judge the current random forest model based on the accuracy of the nth random forest model and the preset accuracy: If the accuracy of the nth random forest model is equal to the preset accuracy, then take the first difference component to the nth difference component as the minimum feature component set, determine that the current random forest model is a trained random forest model, and stop training; if the accuracy of the current random forest model is less than the preset accuracy, then let n=n+1, return to the accuracy evaluation module, and evaluate the accuracy of the current random forest model; the preset accuracy is 100%.
[0019] As can be seen from the above technical solution, compared with the prior art, the present invention provides a pear variety classification method and system based on the minimum feature component set, which has the following beneficial effects: (1) High accuracy: Through rigorous chemometric screening and machine learning modeling, this invention can achieve a 100% accuracy rate in classifying pear varieties; (2) High efficiency: Compared with traditional pear variety detection methods that require the detection of 14 or more components, this invention only requires the detection of 4 characteristic components, which greatly shortens the detection time and improves the detection efficiency of pear varieties; (3) Low cost: By reducing the number of indicators required for pear variety testing, the cost of reagents, consumables and instruments is directly reduced; (4) The technical approach provided by this invention, which is to extract the feature components of the pear to be tested based on the OPLS-DA model and determine the type of pear to be tested through the random forest model, can be used as a general feature screening and model optimization method. It can be widely applied to classification problems such as variety and origin traceability of other agricultural products, promote industry development, and provide practitioners with richer technical references. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0021] Figure 1 This is a flowchart of a pear variety classification method based on a minimum feature component set disclosed in this invention; Figure 2 This is a PCA score graph output by the statistical analysis software SIMCA disclosed in this embodiment of the invention; Figure 3 This is a graph showing the quality parameters of the OPLS-DA model disclosed in an embodiment of the present invention; Figure 4 This is a graph showing the relationship between the number of features and the accuracy of leave-one-out cross-validation as disclosed in the embodiments of the present invention. Figure 5 This is a graph showing the relationship between cumulative feature importance and cumulative importance as disclosed in an embodiment of the present invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] In this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0024] See Figure 1 As shown, this invention discloses a pear variety classification method based on a minimal feature component set, comprising the following steps: S1. Data acquisition steps: Obtain raw physicochemical index data of multiple characteristic components of pear samples from multiple varieties.
[0025] S2. Data preprocessing steps: Standardize the raw physicochemical index data obtained in S1 to obtain a standardized dataset.
[0026] S3. Feature selection steps: Perform principal component analysis (PCA) on the standardized dataset. Verify the separability of different pear varieties based on the clustering results of the PCA score graph. After confirming separability, use the orthogonal partial least squares discriminant analysis (OPLS-DA) model to perform multiple pairwise comparisons on the standardized dataset. Select the components whose variable importance projection (VIP) value is greater than 1 in all comparisons and take their union. After summing and sorting, select the components whose VIP sum is greater than 2 to form the key difference component subset.
[0027] S4. Model building steps: Based on the random forest model, the key difference component sets are divided into components and sorted by importance, and then input into the random forest model for training. Cross-validation is used to determine the optimal number of features required to achieve the highest classification accuracy, thus obtaining the minimum feature component set; the final random forest model is built based on the minimum feature component set.
[0028] S5. Classification Application Steps: Input the content of all components in the minimum feature component set of the pear sample to be tested into the final random forest model obtained in S4, and output the variety of the pear sample to be tested.
[0029] Optionally, the number of pear samples in S1 is four, and the characteristic components include: soluble solids, titratable acid, soluble sugar, solid-acid ratio, vitamin C, vitamin B1, vitamin B2, vitamin B6, niacin, niacinamide, total phenols in peel, total phenols in pulp, flavonoids in peel, and flavonoids in pulp.
[0030] Furthermore, the pear samples include: Snowflake Pear, Akizuki Pear, Crown Pear, and Ya Pear. The present invention will now use these four common and market-representative varieties—Snowflake Pear, Akizuki Pear, Crown Pear, and Ya Pear—as examples to elaborate in detail on the pear variety classification method disclosed in this invention.
[0031] First, the experimental materials and sample preparation were conducted. Snowflake pears, Akizuki pears, Crown pears, and Ya pears of uniform maturity and free from mechanical damage and pests were collected from major production areas. For each variety, 18 individual fruit samples were collected. Then, the 18 fruits within the same variety were grouped for sample preparation: the 18 fruits were randomly divided into 6 groups, each containing 3 fruits; the 3 fruits in each group were cut into pieces, crushed, and homogenized. After thorough mixing, each piece was used to prepare an analytical sample. Using this method, 6 biological replicates were obtained for each variety, resulting in a total of 24 analytical samples for the 4 varieties.
[0032] Secondly, the characteristic component indicators were determined. A total of 14 characteristic component physicochemical indicators were measured for each mixed sample, including: titratable acid, soluble sugar, soluble solids, solid-acid ratio, vitamin C, vitamin B1, vitamin B2, vitamin B6, niacin, niacinamide, total phenols (peel), total phenols (pulp), flavonoids (peel), and flavonoids (pulp) content. Since relevant professionals are familiar with the specific measurement methods for each indicator, and there are no special improvements to the methods, this invention measures each indicator using conventional methods, which will not be elaborated upon here.
[0033] Next comes the standardization of the data. Before conducting multivariate analysis, the multivariate data, namely the 14 characteristic component physicochemical index data, need to be standardized using Z-scores so that the mean of each variable is 0 and the standard deviation is 1. The standardized data is then used for subsequent analysis.
[0034] Optionally, S3 specifically includes: the Z-score normalization formula is: ; in, Z These are the standardized values, i.e., the standard data of the characteristic components. X These are the original values of the physicochemical index data of the characteristic components. μ The mean value corresponding to the physicochemical index data of the characteristic components. σ This represents the standard deviation of the physicochemical index data corresponding to the characteristic components. In subsequent practical applications of this classification model, when standardizing the index data of unknown samples, the mean previously calculated based on the 14 physicochemical index data should be used. μ ) and standard deviation ( σ ) to process.
[0035] Finally, a feasibility analysis was conducted. Principal Component Analysis (PCA) was performed on the standardized physicochemical index data of the 14 characteristic components. The obtained standardized data of 24 samples × 14 indicators was imported into the statistical analysis software SIMCA. The SIMCA software used PCA to analyze the standardized data and output PCA score plots, as shown in the figure below. Figure 2 As shown, H represents Crown Pear, Q represents Autumn Moon Pear, X represents Snowflake Pear, and Y represents Duck Pear. The horizontal and vertical axes in the figure represent the comprehensive feature variables after dimensionality reduction: the horizontal axis represents the first principal component (t[1]), which is a linear combination of the original 14 physicochemical indicators, explaining 32.9% of the feature information of the original dataset; the vertical axis represents the second principal component (t[2]), which explains 26.9% of the feature information of the original dataset. Through these two comprehensive components, the four types of pears can be effectively distinguished on a two-dimensional plane.
[0036] Based on the pear variety classification method based on the minimum feature component set disclosed in this invention, it can be seen that when classifying pear varieties, this invention is based on the premise that the physicochemical index data of the 14 feature components can achieve accurate classification of pear varieties. Therefore, the possibility of classifying pear varieties using the physicochemical index data of the 14 feature components has been verified.
[0037] Optionally, in S3, the separability of different pear varieties is verified, including: Principal component analysis (PCA) was performed on the standardized dataset using the statistical analysis software SIMCA to obtain the PCA score map. The separability of the original physicochemical index data for the pear samples was verified based on the PCA score map.
[0038] according to Figure 2 It can be clearly observed that the sample points of the four varieties—Snowflake Pear, Autumn Moon Pear, Crown Pear, and Duck Pear—are clustered separately, making it relatively easy to distinguish them. Thus, through the above examples, it is fully demonstrated that these 14 physicochemical indicators can effectively distinguish the four pear varieties, demonstrating the feasibility of subsequent pattern recognition and classification.
[0039] After confirming feasibility, to further screen out the marker components that contribute most to distinguishing different varieties, this embodiment of the invention adopts a pairwise comparison strategy. Sample data from the four varieties are combined in pairs to construct a total of six OPLS-DA (Orthogonal Partial Least Squares-Discriminant Analysis) classification models (i.e., models are constructed using the complete sample dataset of Snowflake Pear and the complete sample datasets of Crown Pear, Akizuki Pear, and Ya Pear, as well as Crown Pear vs. Akizuki Pear, Crown Pear vs. Ya Pear, and Akizuki Pear vs. Ya Pear). Through this pairwise discriminant analysis, the characteristic differences between varieties are explored.
[0040] like Figure 3 As shown, in the horizontal header, A represents the number of principal components of the model, N represents the number of samples participating in the modeling, and R²Y and Q² both represent model quality parameters. Theoretically, the closer these two metrics are to 1, the higher the accuracy and reliability of the model. However, generally, a value greater than 0.5 and a difference between them less than 0.3 are acceptable. Figure 3 In the study, six OPLS-DA models (M2 to M7) were constructed and arranged sequentially, with each model undergoing 200 cross-validations. According to... Figure 3 The data shows that the explanatory power parameter R²Y of any OPLS-DA model is greater than 0.9, and the predictive power parameter Q² is greater than 0.8, which is far higher than the threshold of 0.5, and the difference is less than 0.1. This indicates that the constructed OPLS-DA model has a high good fit and good predictive ability.
[0041] In each OPLS-DA model, the Variable Importance in Projection (VIP) value is extracted. In this embodiment of the invention, the VIP value is automatically calculated based on the data processing software SIMCA for each OPLS-DA model, and each model has a corresponding VIP value for its 14 feature data. The calculation principle of the VIP value is as follows: .
[0042] in, p This represents the total number of variables (i.e., 14). A Represents the total number of components in the model. SS a Represents the sum of squares explained by the a-th component. w ajThis represents the weight of the j-th variable in the a-th component. The VIP value essentially reflects the combined contribution of each variable in explaining the variance of the model's classification (Y variable) and the variance of the independent variables (X variable). In this embodiment of the invention, for each OPLS-DA model, the VIP value corresponding to all 14 physicochemical indicators is calculated based on the above formula.
[0043] The Importance Projection (VIP) value is used to measure the explanatory and discriminative power of each variable in variety classification. According to general statistical rules in chemometrics, the sum of squares of VIP values is averaged at 1. Therefore, a VIP value greater than 1 is generally considered an important indicator of a variable's significant contribution, indicating that the variable's explanatory power is above average and is an important condition for significance. Based on this, the names of feature components with a VIP value greater than 1 are extracted from each OPLS-DA model. The union of the feature components selected from the six models is used to obtain an initial pool of differential components containing all potential indicators. To further condense key information, the sum of the VIP values of each variable in the pool across all six OPLS-DA models is calculated. A threshold of 2.0 is set for the total VIP value. This threshold is based on the fact that this method constructs six pairwise comparison models, and in a single model, a VIP > 1 only indicates that the indicator can distinguish a specific pair of varieties. Setting the threshold to 2.0 requires that the selected feature components must meet the conditions of 'significant in at least two models (1+1=2)' or 'extremely high discriminative contribution in a single model (VIP>>2)'. This effectively filters out "marginal indicators" that only play a minor role in a very few groups, retaining the core indicators with the greatest universality and discriminative power. This method eliminates accidental features with weak discriminative power in only a single dimension, ensuring that the selected core physicochemical indicators have stronger broad applicability and classification robustness in complex classification systems composed of multiple varieties. Thus, from the above set, 12 core physicochemical indicators that contribute most significantly to distinguishing these four varieties are ultimately selected, forming the key differential component subset for subsequent modeling. The correspondence between these 12 indicators and their feature numbers in the random forest model is as follows: 1. ss_ta_ratio (solid-acid ratio): Importance 0.1451; 2. Titratable acidity: Importance 0.1297; 3. Vitamin B6: Importance 0.1294; 4. Vitamin B2: Importance 0.1000; 5. Vitamin C: Importance 0.0969; 6. Nicotinamide: Importance 0.0927; 7. Vitamin B1: Importance 0.0820; 8. flavonoid skin: Importance 0.0616; 9. flavonoid_flesh (Flavones from Fruits): Importance 0.0473; 10. soluble_solid_content: Importance 0.0461; 11. phenolic skin (total phenols in fruit peel): Importance 0.0382; 12. Niacin: Importance 0.0310.
[0044] Optionally, the training process for the random forest model in S4 includes: S41. Importance ranking steps: Rank the key difference component subsets by importance to obtain the ranked key difference component subsets {1,2,...,m,...N}, arranged from high to low importance, where m represents the m-th important feature component, N≥m, and m and N are both positive integers.
[0045] S42. Validation and Evaluation Steps: Input the first n feature components of the key difference component subset into the random forest model for training. Evaluate the accuracy of the current random forest model based on leave-one-out cross-validation to obtain the accuracy of the nth random forest model, where n ≤ N, and n is a positive integer. Furthermore, the highest accuracy of the nth random forest model is 100%.
[0046] S43. Accuracy Judgment Steps: Judge the current random forest model based on the accuracy of the nth random forest model and the preset accuracy: If the accuracy of the nth random forest model is equal to the preset accuracy, then take the first difference component to the nth difference component as the minimum feature component set, determine that the current random forest model is a trained random forest model, and stop training; if the accuracy of the current random forest model is less than the preset accuracy, then let n=n+1, return to S42, and evaluate the accuracy of the current random forest model; the preset accuracy is 100%.
[0047] Optionally, the minimum set of characteristic components includes four characteristic components: solid-acid ratio, titratable acid, vitamin B6, and vitamin B2.
[0048] Based on the above embodiments, to find the minimum number of features required to achieve optimal classification performance, the present invention conducts the following systematic tests: Random Forest and Sequence Forward Selection: The 12 key components are sorted from highest to lowest importance. A random forest classification model is constructed using Python's scikit-learn library.
[0049] Leave-One-Out Cross-Validation (LOOCV) was used to evaluate model performance to maximize the use of the limited sample size and ensure unbiased evaluation. A consistent evaluation process was followed regardless of whether the input consisted of a single feature or a combination of features. For comparison, the LOOCV evaluation process was repeated using all 14 original features to verify the screening effect.
[0050] Taking the 12 key components in this embodiment of the invention as examples, the specific verification steps of LOOCV are as follows: 1) Constructing a feature subset: Let n be the number of features in the current test (n=1,2,...,12). Using a sequential accumulation of features, select the first n sorted features to form the current feature subset matrix.
[0051] 2) Perform the LOOCV loop: For a dataset containing M samples (M=24 in this example), perform M training and testing loops. In the i-th loop (i=1 to M), the i-th sample is used as the validation set (test set), and the remaining M-1 samples are used as the training set. The training set (n feature data of M-1 samples) is input into the random forest model for training. The trained model is then used to perform classification prediction on the validation set (n feature data of the i-th sample), and the accuracy of the prediction results is recorded.
[0052] 3) Calculate accuracy: After the loop ends, count the number of correct classifications in the M predictions, and calculate the overall accuracy under the current n feature combinations.
[0053] 4) Determine the optimal feature set: Increment the value of n sequentially (from 1 to 12), repeating the LOOCV process above. Starting with the first-ranked feature (solid-acid ratio), sequentially increase the number of features (using the first 1, first 2, first 3... up to the first 12 features) as input to train the model, and record the average accuracy and standard deviation of LOOCV for each iteration, as shown below. Figure 4 The graph shown illustrates the relationship between the number of features and the accuracy of leave-one-out cross-validation. Figure 4 Clearly, when using all 14 original features, the random forest model achieves 100% accuracy in classifying pear varieties.
[0054] At the same time, according to Figure 4 and Figure 5It can be seen that when the number of features increases to 4 (i.e., solid-acid ratio, titratable acid, vitamin B6, and vitamin B2), the model's classification accuracy reaches 100% for the first time, with a standard deviation of 0, demonstrating stable performance. Further increasing the number of features does not improve accuracy and may even cause fluctuations due to the introduction of noise. Therefore, the first 4 components are determined as the minimum feature set required to achieve 100% classification accuracy. Detecting only these 4 components achieves the same perfect classification effect as detecting all 14 components.
[0055] The present invention also discloses a pear variety classification system based on a minimum feature set, which is used to implement the pear variety classification method based on a minimum feature set described in any of the above embodiments.
[0056] It includes a data acquisition module, a data preprocessing module, a feature selection module, a model building module, and a classification application module, which are connected in sequence.
[0057] The data acquisition module is used to acquire raw physicochemical index data models of various characteristic components of pear samples from multiple varieties.
[0058] The data preprocessing module is used to standardize the raw physicochemical index data obtained from the data acquisition module to obtain a standardized dataset.
[0059] The feature selection module is used to perform principal component analysis (PCA) on the standardized dataset. The clustering results of the PCA score map are used to verify the separability of different pear varieties. After confirming separability, the orthogonal partial least squares discriminant analysis (OPLS-DA) model is used to perform multiple pairwise comparisons on the standardized dataset. Components with variable importance projection (VIP) values greater than 1 in all comparisons are selected and their union is taken. After summing and sorting, the components with a total VIP value greater than 2 are selected to form the key difference component subset.
[0060] The model building module is used to sort the subsets by importance based on the random forest model and input them into the random forest model for training. Leave-one-out cross-validation is used to determine the optimal number of features required to achieve the highest classification accuracy, thus obtaining the minimum feature set. The final random forest model is then built based on the minimum feature set.
[0061] The classification application module is used to input the content of all components in the minimum feature component set of the pear sample to be tested into the final random forest model obtained in the model building module, and output the variety of the pear sample to be tested.
[0062] Corresponding to the pear variety classification based on the minimum feature component set disclosed above, the model building module in the pear variety classification system includes an importance ranking module, an accuracy evaluation module, and a judgment module connected in sequence.
[0063] The importance ranking module is used to rank the subsets of key difference components by importance, resulting in a ranked subset of key difference components {1,2,...,m,...N}, arranged from highest to lowest importance, where m represents the m-th important feature component, N≥m, and m and N are both positive integers.
[0064] The accuracy evaluation module is used to sequentially input the first n feature components of the key difference component subset into the random forest model for training, evaluate the accuracy of the current random forest model based on leave-one-out cross-validation, and obtain the accuracy of the nth random forest model, where n≤N and n is a positive integer.
[0065] The judgment module is used to judge the current random forest model based on the accuracy of the nth random forest model and the preset accuracy: If the accuracy of the nth random forest model is equal to the preset accuracy, then take the first difference component to the nth difference component as the minimum feature component set, determine that the current random forest model is a trained random forest model, and stop training; if the accuracy of the current random forest model is less than the preset accuracy, then let n=n+1, return to the accuracy evaluation module, and evaluate the accuracy of the current random forest model; the preset accuracy is 100%.
[0066] In practical applications, for pear samples of unknown varieties, it is only necessary to measure the content of the above four core characteristic components (solid acid ratio, titratable acid, vitamin B6 and vitamin B2), and then perform Z-score standardization on the raw data obtained by measuring the mean (μ) and standard deviation (σ) calculated from the previous training set. Then, the standardized data is input into the pre-trained final random forest model, which can quickly and accurately output the variety (snowflake pear, autumn moon pear, crown pear or duck pear).
[0067] By inputting the four measured indicators into software that carries a trained random forest model, the species can be predicted.
[0068] In summary, this invention, through a rigorous feature selection strategy combining chemometrics initial screening and machine learning optimization, successfully reduced the number of detection indicators required to distinguish the four pear varieties from 14 to 4 core indicators. This method, while ensuring 100% classification accuracy, establishes the optimal minimum feature set, significantly reducing detection costs and time, and providing an efficient and convenient solution for the rapid and accurate identification of pear varieties.
[0069] The proposed automatic evaluation method for web interface visual complexity establishes a complete technical chain, encompassing stable data collection, rigorous data cleaning, multi-dimensional quantification, in-depth evaluation, and interpretive output, while ensuring both engineering feasibility and cognitive consistency. This method can be used as an offline review tool or deployed as a service in a production environment, demonstrating significant value for web design and user experience evaluation.
[0070] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for classifying pear varieties based on a minimal feature component set, characterized in that, include: S1. Data acquisition steps: Obtain raw physicochemical index data of multiple characteristic components of pear samples from multiple varieties; S2. Data preprocessing step: Standardize the raw physicochemical index data obtained in S1 to obtain a standardized dataset; S3. Feature selection steps: Perform principal component analysis (PCA) on the standardized dataset. Verify the separability of different pear varieties based on the clustering results of the PCA score graph. After confirming separability, use the orthogonal partial least squares discriminant analysis (OPLS-DA) model to perform multiple pairwise comparisons on the standardized dataset. Select the components whose variable importance projection (VIP) value is greater than 1 in all comparisons and take their union. After summing and sorting, select the components whose VIP sum is greater than 2 to form the key difference component subset. S4. Model building steps: Based on the random forest model, the key difference component sets are divided into components and sorted by importance, and then input into the random forest model for training. Cross-validation is used to determine the optimal number of features required to achieve the highest classification accuracy, thus obtaining the minimum feature component set; the final random forest model is built based on the minimum feature component set. S5. Classification Application Steps: Input the content of all components in the minimum feature component set of the pear sample to be tested into the final random forest model obtained in S4, and output the variety of the pear sample to be tested.
2. The pear variety classification method based on the minimum feature component set according to claim 1, characterized in that, The number of pear samples in S1 is four, and the characteristic components include: soluble solids, titratable acid, soluble sugar, solid-acid ratio, vitamin C, vitamin B1, vitamin B2, vitamin B6, niacin, niacinamide, total phenols in peel, total phenols in pulp, flavonoids in peel, and flavonoids in pulp.
3. The pear variety classification method based on the minimum feature component set according to claim 1, characterized in that, The formula for standardization in S2 is: ; in, Z The standardized value. X These are the original values of the original physicochemical index data. μ This represents the mean value corresponding to the original physicochemical index data. σ This represents the standard deviation of the original physicochemical index data.
4. The pear variety classification method based on the minimum feature component set according to claim 1, characterized in that, S3 verifies the separability of different pear varieties, including: Principal component analysis (PCA) was performed on the standardized dataset using the statistical analysis software SIMCA to obtain the PCA score map. The separability of the original physicochemical index data for the pear samples was verified based on the PCA score map.
5. The pear variety classification method based on the minimum feature component set according to claim 1, characterized in that, The training process for the random forest model in S4 includes: S41. Importance ranking steps: Rank the key difference component subsets by importance to obtain the ranked key difference component subsets {1,2,...,m,...N}, arranged from high to low importance, where m represents the m-th important feature component, N≥m, and m and N are both positive integers; S42. Validation and evaluation steps: Input the first n feature components of the key difference component subset into the random forest model for training in turn. Evaluate the accuracy of the current random forest model based on leave-one-out cross-validation to obtain the accuracy of the nth random forest model, where n≤N and n is a positive integer. S43. Accuracy Judgment Steps: Judge the current random forest model based on the accuracy of the nth random forest model and the preset accuracy: If the accuracy of the nth random forest model is equal to the preset accuracy, then take the first difference component to the nth difference component as the minimum feature component set, determine that the current random forest model is a trained random forest model, and stop training; if the accuracy of the current random forest model is less than the preset accuracy, then let n=n+1, return to S42, and evaluate the accuracy of the current random forest model; the preset accuracy is 100%.
6. The pear variety classification method based on the minimum feature component set according to claim 5, characterized in that, The nth random forest model in S42 has the highest accuracy of 100%.
7. The pear variety classification method based on the minimum feature component set according to claim 2, characterized in that, The minimum feature set contains four feature components.
8. The pear variety classification method based on the minimum feature component set according to claim 7, characterized in that, The four characteristic components are: acid-solid ratio, titratable acid, vitamin B6, and vitamin B2.
9. A pear variety classification system based on a minimal feature component set, characterized in that, A method for classifying pear varieties based on a minimum feature component set as described in any one of claims 1-8 includes a data acquisition module, a data preprocessing module, a feature selection module, a model building module, and a classification application module connected in sequence. The data acquisition module is used to acquire raw physicochemical index data models of multiple characteristic components of pear samples from multiple varieties. The data preprocessing module is used to standardize the raw physicochemical index data obtained from the data acquisition module to obtain a standardized dataset; The feature selection module is used to perform principal component analysis (PCA) on the standardized dataset. The clustering results of the PCA score map are used to verify the separability of different pear varieties. After confirming separability, the orthogonal partial least squares discriminant analysis (OPLS-DA) model is used to perform multiple pairwise comparisons on the standardized dataset. Components with variable importance projection (VIP) values greater than 1 in all comparisons are selected and their union is taken. After summing and sorting, the components with a total VIP value greater than 2 are selected to form the key difference component subset. The model building module is used to sort the subsets by importance based on the random forest model and input them into the random forest model for training. Leave-one-out cross-validation is used to determine the optimal number of features required to achieve the highest classification accuracy, thus obtaining the minimum feature set. The final random forest model is then built based on the minimum feature set. The classification application module is used to input the content of all components in the minimum feature component set of the pear sample to be tested into the final random forest model obtained in the model building module, and output the variety of the pear sample to be tested.
10. A pear variety classification system based on a minimal feature component set according to claim 9, characterized in that, The model building module consists of an importance ranking module, an accuracy evaluation module, and a judgment module, which are connected in sequence. The importance ranking module is used to rank the subsets of key difference components by importance, resulting in a ranked subset of key difference components {1,2,...,m,...N}, arranged from high to low importance, where m represents the m-th important feature component, N≥m, and m and N are both positive integers; The accuracy evaluation module is used to sequentially input the first n feature components of the key difference component subset into the random forest model for training, evaluate the accuracy of the current random forest model based on leave-one-out cross-validation, and obtain the accuracy of the nth random forest model, where n≤N and n is a positive integer; The judgment module is used to judge the current random forest model based on the accuracy of the nth random forest model and the preset accuracy: If the accuracy of the nth random forest model is equal to the preset accuracy, then take the first difference component to the nth difference component as the minimum feature component set, determine that the current random forest model is a trained random forest model, and stop training; if the accuracy of the current random forest model is less than the preset accuracy, then let n=n+1, return to the accuracy evaluation module, and evaluate the accuracy of the current random forest model; the preset accuracy is 100%.