Multi-index fusion leafy vegetable freshness grading method and spectral detection model
By using a multi-index fusion grading method and a spectral detection model, the subjectivity and labeling bias issues in the freshness grading of leafy vegetables have been resolved, achieving rapid, accurate, and non-destructive grading that is applicable to the evaluation of the freshness of leafy vegetables under different varieties and storage conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JILIN AGRICULTURAL UNIV
- Filing Date
- 2026-05-13
- Publication Date
- 2026-06-12
AI Technical Summary
Existing methods for grading the freshness of leafy vegetables are highly subjective, suffer from labeling bias, and lack objective basis for determining the number of clusters. This results in insufficient stability and accuracy of grading results, making it difficult to meet the needs for rapid and non-destructive testing.
A multi-index fusion grading method was adopted, combined with visible-near infrared reflectance spectroscopy. By measuring indicators such as weight loss rate, SPAD value and Fv/Fm, the number of clusters was determined using K-means clustering and CRITIC method. Evaluation was carried out by combining elbow method, profile coefficient method, CH index and DBI index. Particle swarm optimization BP neural network and support vector machine models were constructed for spectral detection.
It enables rapid, accurate, and non-destructive grading of leafy vegetables based on their freshness, improves the uniqueness and stability of cluster number selection results, enhances the scientific rigor and reliability of grading results, and is highly adaptable to different varieties and storage conditions.
Smart Images

Figure CN122193165A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of leafy vegetable freshness grading and detection technology, specifically to a leafy vegetable freshness grading method and spectral detection model based on multi-index fusion. Background Technology
[0002] Leafy vegetables, as an important component of the daily diet, are characterized by high water content, fragile tissue structure, and vigorous post-harvest metabolism. They are highly susceptible to quality deterioration during harvesting, storage, and transportation, such as wilting, dehydration, and leaf senescence. Therefore, rapid and accurate evaluation of the freshness of leafy vegetables is crucial for ensuring their quality and safety, reducing losses, and optimizing supply chain management.
[0003] Currently, the evaluation methods for the freshness of leafy vegetables mainly fall into two categories: sensory evaluation and physicochemical index testing. While sensory evaluation is simple to perform, it is highly subjective, and the results are easily influenced by the evaluator's experience, lacking objectivity and consistency. Physicochemical index testing methods, although able to accurately reflect quality changes, typically rely on laboratory conditions and destructive sampling, making the testing process time-consuming and inefficient, and unable to meet the rapid testing needs of actual production and distribution. Furthermore, because leafy vegetables undergo coordinated changes in multiple physiological and biochemical processes during post-harvest decay, such as water metabolism, pigment degradation, and damage to the photosynthetic system, a single index often only reflects one aspect of the information, making it difficult to comprehensively depict the overall freshness status, resulting in incomplete information and biased evaluation.
[0004] With the development of spectroscopic technology, non-destructive testing methods based on visible-near-infrared spectroscopy are gradually being applied to the field of agricultural product quality testing due to their advantages such as speed, non-destructiveness, and rich information. By acquiring the spectral response information of samples at different wavelengths, it is possible to reflect the differences in pigments, moisture, and tissue structure of leafy vegetables during storage, thereby achieving rapid detection of their quality status.
[0005] However, in existing technologies, spectral models typically classify samples based on storage time, using post-harvest time as a freshness label for model training. Since the physiological changes of leafy vegetables are influenced by various factors such as varietal differences, initial state, and environmental conditions, different samples may be at different freshness levels even after the same storage time. Therefore, storage time cannot accurately represent their true physiological state; that is, there is no strict correspondence between time and freshness. Using time as a classification criterion can easily lead to sample label bias, thus affecting the model's classification accuracy and generalization ability.
[0006] To avoid the uncertainties introduced by manual grading or time labeling, some studies have attempted to use cluster analysis to automatically grade samples and uncover the inherent structural features of the data. However, clustering methods typically require pre-setting the number of clusters, and the choice of the number of clusters has a significant impact on the clustering results. Existing methods often rely on a single evaluation indicator or empirical judgment when determining the number of clusters, lacking a unified and objective basis for judgment. Differences may exist between the results obtained from different evaluation methods, leading to insufficient stability and poor repeatability of the grading results. Furthermore, under multi-indicator data conditions, the contribution of each evaluation indicator to the grading results varies. If these indicators are not properly integrated, they can easily cause grading bias, affecting the accuracy and reliability of the final grading results. Summary of the Invention
[0007] To address the problems of existing methods for grading the freshness of leafy vegetables, such as strong subjectivity, label bias, and a lack of objective basis for determining cluster numbers, leading to insufficient stability and accuracy of freshness grading results, this invention aims to propose a multi-index fusion-based method for grading the freshness of leafy vegetables and a spectral detection model. A multi-index fusion-based freshness grading method is constructed, and based on this, it is combined with visible-near-infrared reflectance spectroscopy to build a freshness detection model, achieving rapid, accurate, and non-destructive grading of freshness.
[0008] The method for grading the freshness of leafy vegetables based on multi-indicator fusion includes the following steps: S11. Determine the weight loss rate, SPAD value, and Fv / Fm of leafy vegetables; S12. Use the K-means clustering method to cluster the data obtained in step S1 to obtain a number of clusters; S13. Using several evaluation methods, evaluate several cluster numbers respectively, and obtain several evaluation results corresponding to each cluster number; S14. Using the CRITIC method, weights are assigned to several evaluation results corresponding to each cluster number, and the comprehensive score of each cluster number is calculated. S15. Select the cluster number with the highest comprehensive score as the freshness grade number.
[0009] Furthermore, the evaluation methods include: elbow method, contour coefficient method, CH index, and DBI index.
[0010] Furthermore, the elbow method is identified using the Kneedle algorithm, and the corresponding evaluation results are obtained.
[0011] Furthermore, in step S4, before weighting using the CRITIC method, the evaluation results of the elbow method, contour coefficient method, and CH index are set as positive indicators; the evaluation results of the DBI index are set as negative indicators.
[0012] The spectral detection model for the freshness of leafy vegetables based on multi-index fusion includes the following steps in its construction: S21. Obtain the raw spectral data of leafy vegetables; S22. Several preprocessing methods are used to preprocess the original spectral data to obtain several preprocessed spectral data. S23. Using partial least squares discriminant analysis, several preprocessed spectral data are screened to obtain the optimal preprocessed spectral data. S24. Construct several freshness detection models. After dimensionality reduction of the optimal preprocessed spectral data, use it as the input of several freshness detection models. Use the grade number obtained by the above-mentioned multi-index fusion-based leafy vegetable freshness grading method as the output of several freshness detection models. Train several freshness detection models. S25. Evaluate the accuracy of several trained freshness detection models respectively, and select the freshness detection model with the best accuracy as the leafy vegetable freshness spectral detection model.
[0013] Furthermore, several preprocessing methods include: smoothing, standard normal transformation, multivariate scattering correction, baseline shift, average normalization, first derivative, and second derivative.
[0014] Furthermore, in step S24, the optimal preprocessed spectral data is subjected to dimensionality reduction using principal component analysis.
[0015] Furthermore, several freshness detection models include: Particle Swarm Optimized Backpropagation Neural Network Model, Particle Swarm Optimized Support Vector Machine Model, Gray Wolf Optimized Random Forest Model, and Gray Wolf Optimized Gradient Boosting Decision Tree Model.
[0016] An electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the above-described method for grading the freshness of leafy vegetables based on multi-index fusion.
[0017] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the steps of the above-described method for grading the freshness of leafy vegetables based on multi-index fusion.
[0018] The beneficial effects of the method described in this invention are as follows: (1) This invention selects four clustering effectiveness evaluation indicators: elbow method, silhouette coefficient, CH index, and DBI index. These indicators evaluate clustering quality from four different dimensions: intra-cluster compactness, comprehensive balance between cohesion and separation, variance analysis perspective, and worst-case robustness, forming a structured functionally complementary relationship. Based on this, the CRITIC method is used to objectively allocate weights according to the comparative strength of each indicator under different cluster numbers and the conflict between indicators. This effectively solves the problems of one-sided evaluation perspectives of single indicators and lack of objective adjudication basis when multiple indicator conclusions contradict each other in existing technologies. It ensures the uniqueness, stability, and reproducibility of the cluster number selection results, providing a reliable technical guarantee for the accurate grading of leafy vegetable freshness.
[0019] (2) This invention automatically identifies the elbow inflection point of the SSE curve in the elbow method using the Kneedle algorithm, and constructs a multi-index evaluation system by combining the contour coefficient, CH index, and DBI index. The CRITIC objective weighting method is used to comprehensively integrate the evaluation indicators, and finally the optimal number of clusters is determined based on the principle of maximizing the comprehensive score. Compared with the existing technology that relies on manual observation of the elbow, judgment based on a single indicator, or storage time as a label, this invention realizes the objectivity and automation of the entire process of determining the number of clusters, eliminates the subjective bias of human interpretation, and avoids the labeling bias problem caused by the discrepancy between storage time and true freshness, thus ensuring the scientificity and reliability of the grading basis from the source.
[0020] (3) This invention addresses the characteristics of leafy vegetables, where post-harvest quality degradation is continuous and gradual, and there are no natural discrete boundaries between different freshness grades. It utilizes a clustering analysis method based on multiple indicators to automatically classify spectral data. This method does not rely on preset grade thresholds or manual labeling and can be directly applied to freshness evaluation scenarios for leafy vegetables of different varieties, origins, and storage conditions, demonstrating strong adaptability and generalization ability. Furthermore, the technical framework of this invention can be extended to other agricultural product quality evaluation fields facing similar needs of "continuous quality degradation and unsupervised grading," providing a promising technical solution for rapid quality detection and intelligent grading in the fresh agricultural product supply chain. Attached Figure Description
[0021] Figure 1 This is a flowchart of the grading method and the freshness spectral detection model construction method described in this invention; Figure 2 This is a schematic diagram showing the weight distribution of each evaluation index described in this invention under CRITIC weighting; Figure 3 This is a schematic diagram illustrating the evaluation results of several clustering numbers described in this invention; Figure 4 This is a schematic diagram of the spectral data acquisition device described in this invention. Detailed Implementation
[0022] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] Example 1 This embodiment provides a method for grading the freshness of leafy vegetables based on multi-index fusion. The flowchart of the method is as follows: Figure 1 As shown, the method includes the following steps: S11. Determine the weight loss rate, SPAD value, and Fv / Fm of leafy vegetables; The relevant operations in step S11 will be introduced with specific examples: In this embodiment, 30 commercially available lettuce plants of relatively uniform size were selected as leafy vegetable samples. To avoid interference from aging, damage, and moisture evaporation of the outer leaves on the freshness test results, and to ensure the accuracy and consistency of the experiment, the outermost 2-3 leaves were removed from each sample, and the freshness was tested at room temperature. The entire experimental period was 72 hours, with index measurements taken every 12 hours, for a total of 7 data collections. To ensure that the leaves measured each time were the same leaves, the 30 lettuce plants were numbered.
[0024] The process for determining the SPAD value of chlorophyll content in leaves in this embodiment is as follows: The SPAD value of lettuce leaves is measured using a handheld chlorophyll meter (SPAD-502 Plus, Japan). The main vein is avoided during the measurement, and each leaf is measured 3 times. The average value is taken as the SPAD value of that leaf.
[0025] The procedure for determining the maximum photochemical efficiency Fv / Fm in this embodiment is as follows: The maximum photochemical efficiency Fv / Fm of lettuce leaves was measured using a Pocket PEA (Hansatech, UK) plant efficiency analyzer. Each leaf was measured three times, and the average value was taken as the maximum photochemical efficiency Fv / Fm of that leaf.
[0026] The formula for calculating the weight loss rate in this embodiment is: In the formula The percentage of weightlessness is expressed as % This is the initial weight of the lettuce, in grams. for Weight at any given time, expressed in grams (g).
[0027] The initial weight of the whole lettuce head was determined using an electronic balance with an accuracy of 0.01 g. Calculate the weight loss rate, where the weight loss rate is the whole-head weight loss rate of the corresponding numbered lettuce.
[0028] S12. Use the K-means clustering method to cluster the data obtained in step S1 to obtain a number of clusters; The relevant operations in step S12 will be introduced with specific examples: K-means clustering is computationally efficient and converges quickly, making it particularly suitable for handling large-scale, multi-dimensional data. Furthermore, the clustering results are highly interpretable; the centroid of each cluster can serve as a representative of that cluster, providing a reliable source for subsequent data analysis and applications. The K-means clustering method can be used to achieve objective classification of multi-indicator data for leafy vegetables.
[0029] In this embodiment, the K-means clustering method is used to cluster the data obtained in step S1. The number of clusters (K value) obtained includes: 2, 3, 4, 5, 6 and 7 (since a total of 7 data were collected, the maximum number of clusters is 7). S13. Using several evaluation methods, evaluate several cluster numbers respectively, and obtain several evaluation results corresponding to each cluster number; The relevant operations in step S13 will be described with specific examples: In practical applications, the K-means algorithm typically requires a pre-defined value for the number of clusters, K. However, in most cases, this value cannot be determined beforehand. From a clustering principle perspective, if K is too small, the algorithm will force samples with significant differences in the feature space into the same cluster, leading to overlapping of samples at different freshness levels, excessive dispersion within clusters, and a loss of grading significance. Conversely, if K is too large, it will over-segment samples belonging to the same freshness level, generating numerous redundant micro-clusters, fragmenting the grading results, reducing interpretability, and decreasing the model's generalization ability. Furthermore, the quality decay of leafy vegetables is a continuous, gradual process, with samples exhibiting a continuous manifold distribution in the feature space without significant discontinuities. This further exacerbates the ambiguity and uncertainty in determining the true number of clusters, making it difficult to determine the optimal number of clusters in practical applications. An incorrect choice of the number of clusters can lead to unsatisfactory clustering results.
[0030] This embodiment employs several evaluation methods, namely the elbow method, silhouette coefficient, Calinski-Harabasz (CH) index, and Davies-Bouldin (DBI) index, to comprehensively determine the number of cluster categories. The selection is based on the fact that these indicators characterize clustering effects from different evaluation perspectives, forming a structured functional complementarity and synergistically achieving a comprehensive measurement of clustering quality. Specifically: The elbow method measures the compactness of samples within a cluster by calculating the sum of squared distances (SSE) between each sample point and the centroid of its cluster, reflecting the basic cohesion level of the cluster structure. The silhouette coefficient examines both intra-cluster cohesion and inter-cluster separation, providing a comprehensive and balanced perspective for evaluating cluster goodness. The CH index is based on the idea of analysis of variance. It judges the rationality of clustering by the ratio of inter-class dispersion to intra-class aggregation, and focuses on the statistical significance of the overall structure. The DBI index employs a worst-case analysis strategy, characterizing the ratio of the sum of the average intra-class distances between any two categories to the corresponding inter-class center distance, reflecting the separation performance and robustness of the clustering results under the most unfavorable conditions.
[0031] The four indicators mentioned above evaluate clustering quality from four different and complementary dimensions: "intra-cluster compactness," "comprehensive balance between cohesion and separation," "structural rationality from the perspective of analysis of variance," and "worst-case robustness." This effectively overcomes the limitation that a single evaluation indicator can only reflect one aspect of the clustering effect and cannot comprehensively measure the quality of the cluster structure. By introducing the above multi-perspective evaluation indicator system for comprehensive analysis, the objectivity and reliability of determining the number of clusters can be significantly improved.
[0032] Among the above indicators, SSE continuously decreases as K increases. The smaller the SSE, the denser the sample points within a cluster, and the higher the degree of aggregation. However, if K is too large, the meaning of clustering is easily lost, and the clustering effect is not good. As the number of clusters increases, SSE will first decrease rapidly, then slow down at a certain critical position, forming an elbow (inflection point). The number of clusters corresponding to this elbow can be considered the ideal number. The closer the silhouette coefficient is to 1, the better the sample points fit the current cluster; the closer the silhouette coefficient is to -1, the better the sample points fit other clusters. The advantage of the silhouette coefficient method is that it allows for a direct comparison of the effects under different numbers of clusters; the number of clusters with the largest silhouette coefficient is the optimal number. A larger CH value indicates a better clustering effect. A smaller DBI value indicates smaller intra-cluster distances and larger inter-cluster distances, resulting in better performance.
[0033] S14. Using the CRITIC method, weights are assigned to several evaluation results corresponding to each cluster number, and the comprehensive score of each cluster number is calculated. The relevant operations in step S14 will be described with specific examples: This invention employs the CRITIC method to comprehensively evaluate the clustering effect (several evaluation results obtained in step S13). The CRITIC method is an objective weighting method based on data characteristics, whose weight allocation comprehensively considers the contrast strength and conflict of evaluation indicators. Contrast strength reflects the dispersion of the same indicator across different evaluation objects, typically characterized by standard deviation. A larger standard deviation indicates stronger discriminative ability of the indicator, and its corresponding weight is higher. Conflict is characterized by the correlation between indicators. When the correlation between indicators is high, it indicates some information redundancy and weak conflict, and the corresponding weight should be reduced. By simultaneously incorporating the dispersion and correlation information of indicators, the CRITIC method achieves objective determination of the weights of each evaluation indicator, thereby improving the rationality and scientific nature of weight allocation.
[0034] The core idea of the CRITIC method is to consider both the comparative strength and conflict of indicators simultaneously, thereby objectively assigning weights to a set of comparable variables.
[0035] In this invention, the application of the CRITIC method is expanded from traditional original feature variables to cluster effectiveness evaluation indicators (such as silhouette coefficient, CH index, and DBI index). Its role is to "weight the evaluation indicator system," selecting the most representative cluster evaluation criteria to determine the optimal number of clusters. Unlike the conventional CRITIC method, which assigns weights to original feature variables, the CRITIC method in this invention achieves a functional migration from the data feature layer to the evaluation decision layer, representing a substantial expansion in the method's operational level rather than a simple replacement of the application scenario. This method effectively avoids the problem of a single evaluation indicator dominating the clustering result determination, improves the stability and consistency of the optimal number of clusters selection, and thus enhances the objectivity and robustness of the cluster structure evaluation.
[0036] Before comprehensively evaluating the clustering effect using the CRITIC method, this invention employs the Kneedle algorithm to automatically identify the elbow points of the SSE curve. The Kneedle algorithm automatically determines the optimal values of parameters (such as the number of clusters K, threshold, etc.) by finding the point with the maximum curvature or the most significant deviation from the linear trend in a monotonically decreasing (or increasing) curve. The SSE index, silhouette coefficient, and Calinski-Harabasz (CH) index are set as positive indicators, while the Davies-Bouldin (DBI) index is set as a negative indicator.
[0037] In this embodiment, the CRITIC method assigns weights to several evaluation results corresponding to each cluster number, such as... Figure 2 As shown, according to Figure 2 The weighting results shown are used to calculate the comprehensive score for each cluster number. The specific steps are as follows: First, each indicator is positively oriented and standardized to eliminate the influence of dimensions; then, the variability (standard deviation) of each indicator is calculated. Based on this, the correlation between the indicators is quantified by calculating the Pearson correlation coefficient matrix, and the conflict index value of each indicator is calculated based on the correlation coefficient. The variability and conflict index values of the indicators are then multiplied to obtain the information content; finally, the information content of each indicator is normalized to obtain the weights. Figure 2 (The weighting results are shown).
[0038] In the comprehensive evaluation process, the standardized indicator data are multiplied by their corresponding weights and then summed using a weighted average to obtain the comprehensive score for each sample. Since the range of scores after the weighted sum is not fixed, to improve the comparability between different samples, the comprehensive score is normalized again to map it to a uniform numerical range, thus obtaining, for example... Figure 3 The final evaluation results shown are the combined scores for each cluster number.
[0039] S15. Select the cluster number with the highest comprehensive score as the freshness grade number.
[0040] The relevant operations in step S15 will be described with specific examples: like Figure 3 As shown, the overall score reaches its maximum value when the number of clusters (K value) is 3, indicating that the number of clusters performs optimally under the weighting of multiple indicators. Therefore, the optimal number of clusters in this embodiment is determined to be 3, that is, the freshness is divided into 3 levels. In this embodiment, combined with freshness-related indicators, these three categories are divided into fresh, slightly fresh and not fresh.
[0041] Example 2 This embodiment further defines Embodiment 1. This embodiment provides a multi-index fusion-based model for grading the freshness of leafy vegetables and for spectral detection. The method for constructing the model is as follows: Figure 1 As shown, the method includes the following steps: S21. Obtain the raw spectral data of leafy vegetables; The relevant operations in step S21 will be introduced with specific examples: like Figure 4 As shown, this embodiment uses a fiber optic spectrometer to collect the visible-near-infrared reflectance spectrum of lettuce leaves. During data acquisition, the fiber optic cable is first connected to the spectrometer and the light source. Then, the fiber optic cable is fixed using a reflectance probe holder, ensuring a 45° angle between the cable and the sample. The spectrometer is preheated for 10 minutes before acquisition, followed by white board calibration. Calibration is considered successful when the white board's reflectance reaches 100%. White board calibration is performed every 30 minutes throughout the acquisition process. Three spectral curves are collected from each leaf, avoiding the main vein, and the average value is taken as the visible-near-infrared reflectance spectrum data for that leaf.
[0042] These indicators were chosen for detection in this embodiment because they can reflect the physiological state and quality changes of leafy vegetables from multiple perspectives, including moisture status, pigment content, and photosynthetic system function. Furthermore, these indicators can all be obtained through non-destructive measurement methods, allowing data collection without damaging leaf tissue. This ensures the integrity of the vegetables, provides a reliable data foundation for subsequent spectral detection, and meets the practical needs of rapid, online monitoring.
[0043] S22. Several preprocessing methods are used to preprocess the original spectral data to obtain several preprocessed spectral data. The relevant operations in step S22 will be described with specific examples: To reduce the impact of instrument noise, light scattering effects, baseline drift, and other factors on the raw spectral data during acquisition, and to improve the signal-to-noise ratio and stability of the spectral data, this invention preprocesses the raw spectral data using methods such as smoothing (SG), standard normal transformation (SNV), multivariate scattering correction (MSC), baseline shift (B), mean normalization (MN), first derivative (FD), and second derivative (SD).
[0044] The above preprocessing methods are designed for different interference factors and information expression characteristics in spectral data: SG is used to reduce high-frequency noise and improve the signal-to-noise ratio, and can effectively reduce noise and spurious signals in data; SNV processing and MSC are used to eliminate the influence of scattering effects on the spectrum, which can reduce the scattering effect on the spectral curve; B is used to correct spectral baseline drift, which can reduce the impact of baseline drift and systematic errors; MN is used to eliminate differences in measurement conditions, which helps to unify the intensity of spectral curves; FD and SD are used to enhance spectral detail features, reduce interference from overlapping peaks, resolve overlapping peaks in spectral curves, and enhance differences between spectra. The aforementioned preprocessing methods process spectral data from different technical paths, such as noise suppression, scattering correction, baseline correction, scale unification, and feature enhancement, forming a structured complementary relationship in terms of function.
[0045] This invention does not simply superimpose multiple preprocessing methods, but rather constructs a candidate set of preprocessing methods. By processing spectral data independently and evaluating and comparing the processing effects of each method in conjunction with model performance indicators, the optimal preprocessing strategy is selected. This method can fully leverage the differentiated advantages of different preprocessing methods in terms of noise suppression, scattering correction, and feature enhancement, thereby achieving the optimization of spectral data preprocessing results.
[0046] S23. Using partial least squares discriminant analysis, several preprocessed spectral data are screened to obtain the optimal preprocessed spectral data. The relevant operations in step S23 will be introduced with specific examples: In this embodiment, to optimize different spectral preprocessing methods, partial least squares discriminant analysis (PLS-DA) is used to establish a classification model to evaluate each preprocessed data and select the best preprocessing method. During PLS-DA modeling, the number of latent variables needs to be determined to optimize model performance and improve prediction stability. Too few latent variables will lead to insufficient extraction of key information and a decrease in the model's discriminative ability, while too many will easily introduce noise and reduce the model's generalization performance. Therefore, this embodiment uses 5-fold cross-validation on the training set to optimize the selection of the number of latent variables, using accuracy as the evaluation index to determine the optimal number of latent variables. The training and test sets are randomly divided in a 4:1 ratio to establish a freshness detection model, and the results are shown in Table 1. Among the established models, the model established by SG preprocessing has the highest accuracy. Therefore, further analysis of the spectral data after SG preprocessing will be conducted.
[0047] Table 1
[0048] Because SG-processed spectra are characterized by high dimensionality and strong correlation (strong collinearity) among variables, direct modeling can easily lead to complex models, low computational efficiency, and poor stability. Therefore, this embodiment uses principal component analysis (PCA) to perform dimensionality reduction analysis on SG-processed spectra. PCA, a widely used multivariate statistical method, transforms correlated original variables into principal components using coordinate transformation, significantly reducing the number of variables while ensuring that the reduced features reflect the internal structure of the original variables as much as possible. The choice of the number of principal components directly affects the accuracy of the model. Based on the principle of eigenvalue ≥1 and cumulative contribution rate ≥85%, the number of principal components was determined. In this embodiment, the first 7 principal components were selected for further analysis. The analysis results are shown in Table 2. The highest cumulative contribution rate of the first 7 principal components reached 97.848%, which can reflect most of the information in the original data, effectively reducing the data dimensionality while retaining the main characteristic information of the original data.
[0049] Table 2
[0050] S24. Construct several freshness detection models. After dimensionality reduction of the optimal preprocessed spectral data, use it as the input of several freshness detection models. Use the grading number obtained by the grading method described in Example 1 as the output of several freshness detection models to train several freshness detection models. The relevant operations in step S24 will be described with specific examples: The freshness detection models mentioned are Particle Swarm Optimized Backpropagation Neural Network (PSO-BPNN), Particle Swarm Optimized Support Vector Machine (PSO-SVM), Grey Wolf Optimized Random Forest (GWO-RF), and Grey Wolf Optimized Gradient Boosting Decision Tree (GWO-XGBoost) models.
[0051] Backpropagation (BP) neural networks possess strong nonlinear mapping capabilities, as well as good self-learning, adaptive abilities, and high fault tolerance. However, they still have certain limitations, such as slow convergence speed, susceptibility to local optima, and sensitivity to initial weights and thresholds. SVM (Structured Variable Machine) is a supervised machine learning method based on statistical theory, commonly used for classification and regression. By searching for the minimum structural risk, the generalization ability of the SVM classifier can be improved and the empirical risk minimized. Therefore, it can exhibit good classification performance even with a small sample size. However, traditional SVMs typically require manual specification of two parameters (c and g). c is the penalty coefficient, representing the tolerance for error; an excessively high c value can easily lead to overfitting. The parameter g in the kernel function plays a decisive role in the distribution of the data after mapping to the new feature space. Randomization (RF) classification is a machine learning algorithm composed of decision trees. This algorithm is suitable for processing high-dimensional data and runs relatively fast, exhibiting high accuracy and robustness in predicting multi-factor, nonlinear variables. However, RF performance depends on hyperparameters such as the number of decision trees and the maximum tree depth; inappropriate parameters can easily lead to decreased accuracy or overfitting. XGBoost is an improvement on the GBDT algorithm, achieving better-than-expected results in few-shot classification problems. Its idea is to integrate simple weak classifiers and perform a second-order Taylor expansion of the objective function to retain more target information, thereby improving model accuracy. However, conventional XGBoost models are characterized by excessive parameters and computational complexity, making parameter optimization crucial.
[0052] To eliminate the impact of manual parameter tuning on model performance, Particle Swarm Optimization (PSO) was used to optimize BPNN and SVM. PSO can quickly find approximate optimal solutions to problems and effectively avoid getting trapped in local optima, thus getting closer to the global optimum. Grey Wolf Optimization (GWO) was used to optimize RF and XGBoost. GWO, a swarm intelligence algorithm derived from the social hierarchy and cooperative hunting behavior of grey wolves, has advantages such as simple structure, few parameters, and ease of implementation.
[0053] To further explore the prediction effect of the SG preprocessed spectra after principal component dimensionality reduction, the optimal preprocessed spectral data was used as input after dimensionality reduction, and the clustered labels were used as categories (the number of levels obtained by the grading method described in Example 1 was used as output). The training set and test set were divided in a ratio of 4:1, and the PSO-BPNN, PSO-SVM, GWO-RF and GWO-XGBoost models were trained and tested.
[0054] S25. Evaluate the accuracy of several trained freshness detection models respectively, and select the freshness detection model with the best accuracy as the leafy vegetable freshness spectral detection model.
[0055] The relevant operations in step S25 will be described with specific examples: Table 3 shows the accuracy results of the trained PSO-BPNN, PSO-SVM, GWO-RF, and GWO-XGBoost models on the training and test sets. Among them, the PSO-SVM model performed best, with accuracy rates of 100% on the training set and 97.62% on the test set. This indicates that visible-near-infrared reflectance spectroscopy can achieve rapid, non-destructive, and accurate detection of leafy vegetable freshness. Therefore, in this embodiment, the trained PSO-SVM model was selected as the spectral detection model for leafy vegetable freshness.
[0056] Table 3
[0057] In this embodiment, the accuracy is calculated as follows:
[0058] Where A represents accuracy, TP represents the number of true positive samples, TN represents the number of true negative samples, FP represents the number of false positive samples, and FN represents the number of false negative samples.
Claims
1. A multi-indicator fusion method for grading the freshness of leafy vegetables, characterized in that, The method includes the following steps: S11. Determine the weight loss rate, SPAD value, and Fv / Fm of leafy vegetables; S12. Use the K-means clustering method to cluster the data obtained in step S1 to obtain a number of clusters; S13. Using several evaluation methods, evaluate several cluster numbers respectively, and obtain several evaluation results corresponding to each cluster number; S14. Using the CRITIC method, weights are assigned to several evaluation results corresponding to each cluster number, and the comprehensive score of each cluster number is calculated. S15. Select the cluster number with the highest comprehensive score as the freshness grade number.
2. The multi-index fusion method for grading the freshness of leafy vegetables according to claim 1, characterized in that, The evaluation methods include: elbow method, contour coefficient method, CH index, and DBI index.
3. The multi-index fusion method for grading the freshness of leafy vegetables according to claim 2, characterized in that, The elbow method is identified using the Kneedle algorithm, and the corresponding evaluation results are shown.
4. The multi-index fusion method for grading the freshness of leafy vegetables according to claim 3, characterized in that, In step S4, before weighting using the CRITIC method, the evaluation results of the elbow method, contour coefficient method, and CH index are set as positive indicators; the evaluation results of the DBI index are set as negative indicators.
5. A multi-index fusion spectral detection model for the freshness of leafy vegetables, characterized in that, The method for constructing the model includes the following steps: S21. Obtain the raw spectral data of leafy vegetables; S22. Several preprocessing methods are used to preprocess the original spectral data to obtain several preprocessed spectral data. S23. Using partial least squares discriminant analysis, several preprocessed spectral data are screened to obtain the optimal preprocessed spectral data. S24. Construct several freshness detection models, and use the dimensionality reduction of the optimal preprocessed spectral data as the input of several freshness detection models. Use the grading number obtained by the grading method described in any one of claims 1 to 4 as the output of several freshness detection models, and train several freshness detection models. S25. Evaluate the accuracy of several trained freshness detection models respectively, and select the freshness detection model with the best accuracy as the leafy vegetable freshness spectral detection model.
6. The multi-index fusion spectral detection model for leafy vegetable freshness according to claim 5, characterized in that, If the intervention methods include: smoothing, standard normal transformation, multivariate scattering correction, baseline shift, mean normalization, first derivative, and second derivative.
7. The multi-index fusion spectral detection model for leafy vegetable freshness according to claim 6, characterized in that, In step S24, the optimal preprocessed spectral data is subjected to dimensionality reduction using principal component analysis.
8. The multi-index fusion spectral detection model for leafy vegetable freshness according to claim 7, characterized in that, Several freshness detection models include: Particle Swarm Optimized Backpropagation Neural Network Model, Particle Swarm Optimized Support Vector Machine Model, Gray Wolf Optimized Random Forest Model, and Gray Wolf Optimized Gradient Boosting Decision Tree Model.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-4.
10. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-4.