A method and system for identifying the authenticity of leaves of anoectochilus roxburghii and classifying strains
Through multi-view spectral data combined with SVM and CNN models, the accuracy of leaf identification and strain classification of nigra plants was solved, high-precision authenticity identification and strain classification were achieved, and the application of hyperspectral imaging in medicinal plant certification was promoted, ensuring consumer health and sustainable practices in traditional pharmaceutical industries.
Patent Information
- Application Number
- CN202510694179.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The prior art has low accuracy in the identification and strain classification of authenticity of leaves of genus filament plants, and relies on strong subjective chemical analysis methods, making it difficult to achieve high-precision identification and classification.
Multi-view spectral data combined with support vector machine (SVM) and convolutional neural network (CNN) models are used to obtain the front and back spectral data of the leaves of the genus milita plant, and then input the trained SVM and CNN models after preprocessing to achieve authenticity and false identification and strain classification.
The accuracy of authenticity identification and strain classification of leaves of nigra plant has been improved, non-destructive and high-precision identification and classification have been achieved, and the application of hyperspectral imaging in medicinal plant certification has been promoted, ensuring consumer health and sustainable practices in the traditional pharmaceutical industry.
Smart Images

Figure CN120213831B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of plant leaf identification and classification, and in particular to a method and system for authenticity identification and strain classification of roxburghii plant leaves. Background Art
[0002] Currently, the quality identification of Anoectochilus roxburghii usually relies on chemical analysis methods, including visual inspection, microscopic identification, high-performance liquid chromatography, DNA molecular identification, and near-infrared spectroscopy. However, these methods are generally subjective and have low accuracy. Therefore, how to improve the accuracy of authenticity identification and strain classification of Anoectochilus roxburghii plant leaves has become a technical problem that needs to be solved urgently in this field. Summary of the Invention
[0003] The purpose of this application is to provide a method and system for authenticity identification and strain classification of leaves of anoectochilus roxburghii plants, which can effectively improve the accuracy of authenticity identification and strain classification of leaves of anoectochilus roxburghii plants.
[0004] To achieve the above objectives, this application provides the following solutions.
[0005] In a first aspect, the present application provides a method for identifying the authenticity of leaves of a roxburghii plant and classifying its strains, which specifically includes the following steps.
[0006] Acquire multi-view spectral data of the plant sample to be tested; the multi-view spectral data includes front spectral data and back spectral data of each leaf of the plant sample to be tested.
[0007] The multi-view spectral data is preprocessed to obtain preprocessed multi-view spectral data.
[0008] The preprocessed multi-view spectral data is input into the pre-trained SVM model to output the authenticity identification result; the pre-trained SVM model refers to the preprocessed multi-view spectral data based on the golden thread lotus sample and the counterfeit sample, and the SVM model and the pre-processing model are jointly adjusted to find the optimal parameters and then trained to obtain the model, and the pre-processing model refers to the model corresponding to the filtering algorithm.
[0009] When the authenticity identification result is that the plant sample to be tested is a roxburghii sample, the preprocessed multi-view spectral data of the plant sample to be tested is input into the pre-trained CNN model, and the variety classification result is output; the pre-trained CNN model refers to a model obtained by jointly adjusting the parameters of the CNN model and the preprocessing model based on the preprocessed multi-view spectral data of the roxburghii sample, and finding the optimal parameters for training.
[0010] Optionally, obtaining multi-view spectral data of the plant sample to be tested specifically includes the following steps.
[0011] The plant sample to be tested is cleaned to obtain a cleaned plant sample to be tested.
[0012] The cleaned plant sample to be tested is air-dried to obtain an air-dried plant sample to be tested.
[0013] The front spectral data and the back spectral data of each leaf of the air-dried plant sample to be tested are collected to obtain multi-view spectral data.
[0014] Optionally, collecting the front spectral data and the back spectral data of each leaf of the air-dried plant sample to be tested to obtain multi-view spectral data specifically includes the following steps.
[0015] A GaiaField-N17E hyperspectral imaging system was used to scan the front and back of each leaf of the air-dried plant sample to be tested, respectively, to collect spectral data of the front and back of each leaf, and obtain multi-view spectral data.
[0016] Optionally, preprocessing the multi-view spectral data to obtain preprocessed multi-view spectral data specifically includes the following steps.
[0017] Black-white correction is performed on the multi-viewing angle spectral data to obtain black-white corrected multi-viewing angle spectral data.
[0018] The regions of interest are extracted from the black-and-white corrected multi-view spectral data to obtain pre-processed multi-view spectral data.
[0019] Optionally, before training the pre-trained SVM model, the pre-processed multi-view spectral data is standardized using a Standard Scaler; when training the pre-trained SVM model, grid search and five-fold cross-validation methods are used to optimize hyperparameters, including penalty parameters, kernel type, gamma, and polynomial degree; and accuracy, precision, recall rate, F1 score, and confusion matrix are used as evaluation indicators to evaluate model performance.
[0020] Optionally, the filtering algorithm is at least one of a median filtering algorithm, an average filtering algorithm, a Gaussian filtering algorithm, a Savitzky-Golay filtering algorithm and a principal component analysis method.
[0021] Optionally, the pre-trained CNN model is a 1D-CNN model.
[0022] In the second aspect, the present application provides a system for identifying the authenticity of leaves of a golden thread vine plant and classifying its strains, comprising a memory, a processor, and a computer program stored on the memory and runnable on the processor, wherein the processor executes the computer program to implement the above-mentioned method for identifying the authenticity of leaves of a golden thread vine plant and classifying its strains.
[0023] According to the specific embodiments provided in this application, this application has the following technical effects:
[0024] The present application provides a method and system for identifying the authenticity of leaves of a golden thread vine plant and classifying strains, which combines multi-view spectral data, SVM model and CNN model, and is applied to the scene of authenticity identification and strain classification of leaves of a golden thread vine plant, realizing the effective combination of hyperspectral imaging technology and machine learning technology. Among them, by obtaining the front spectral data and back spectral data of each leaf of the plant sample to be tested, multi-view spectral data is formed. Based on this, in terms of the authenticity identification of golden thread vine and counterfeit varieties, by using a pre-trained SVM model and utilizing the spectral data of the front and back leaves, it is possible to accurately distinguish between golden thread vine and counterfeit varieties, and realize accurate identification of the authenticity of golden thread vine plant leaves. In terms of the classification of different varieties of golden thread vine, by introducing a pre-trained CNN model and combining the spectral data of the front and back leaves, the classification accuracy of the model can be effectively improved, and the accuracy of strain classification can be improved. Therefore, the present application can effectively improve the accuracy of authenticity identification and strain classification of golden thread vine plant leaves by combining multi-view spectral data, SVM model and CNN model. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0026] Figure 1 A flowchart of a method for identifying the authenticity of leaves of a roxburghii plant and classifying its strains is provided in one embodiment of the present application.
[0027] Figure 2 A schematic diagram of the structure of a hyperspectral imaging system provided in one embodiment of the present application.
[0028] Figure 3 A schematic diagram of the structure of the 1D-CNN model provided in one embodiment of the present application.
[0029] Figure 4 A graph showing how the loss value of a loss function changes with the number of training times during the training process provided in one embodiment of the present application.
[0030] Figure 5 This is a graph showing how the accuracy changes with the number of training times during the training process provided in one embodiment of the present application.
[0031] Figure 6 A schematic diagram of the confusion matrix of the trained 1D-CNN model on the test set provided in one embodiment of the present application. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0033] At present, methods for identifying medicinal plants such as golden thread vine, such as visual inspection, microscopic identification, high-performance liquid chromatography, and DNA molecular identification, need to be carried out under the guidance of professionals. They have defects such as strong subjectivity, time-consuming and low-precision processes, complex operations and high costs. They are not only labor-intensive, but also damage samples. Most importantly, they often cannot distinguish subtle interspecies variations.
[0034] Recent advances in spectroscopic technology have shown promise in addressing these challenges. For example, a related study combined near-infrared spectroscopy with chemometrics to develop a partial least squares discriminant analysis (PLS-DA) model accurately distinguished between authentic Anoectochilus roxburghii powder and two counterfeit products. Another related study used near-infrared spectroscopy to obtain spectral data of Anoectochilus roxburghii and its adulterated counterparts, and designed an improved one-dimensional convolutional neural network (1D-CNN) model to process the near-infrared spectral data and distinguish between authentic and counterfeit Anoectochilus roxburghii. Furthermore, a model for rapid and accurate classification of Anoectochilus roxburghii varieties based on a handheld near-infrared spectrometer and AdaBoost (Adaptive Boosting) ensemble learning achieved a classification accuracy of 95.6%. These studies demonstrate the high accuracy of near-infrared spectroscopy for detecting Anoectochilus roxburghii, but its application is generally limited to the powdered form of the plant, which requires grinding, compression, or other sample preparation methods.
[0035] Hyperspectral imaging (HSI) can capture spatial and spectral information across hundreds of continuous bands while maintaining sample integrity. It captures a wide range of spectral bands across the electromagnetic spectrum, providing detailed spectral information that can be used for precise material identification and quality assessment. One related technology, a hyperspectral model, utilizes spectral data analysis techniques to detect the flavonoid and polysaccharide content in Anoectochilus roxburghii. However, its potential for authenticity identification and multi-variety classification remains unexplored. Furthermore, existing studies typically rely on single-view spectral data (e.g., the front of the leaf) while ignoring the complementary information embedded in multiple viewpoints (e.g., the front and back of the leaf).
[0036] Machine learning (ML) and deep learning (DL) have revolutionized pattern recognition in high-dimensional datasets, making them ideal for analyzing hyperspectral data. While traditional machine learning models are widely used, their performance depends heavily on manual feature selection and preprocessing. In contrast, deep learning architectures such as convolutional neural networks (CNNs) can autonomously extract discriminative features, enabling end-to-end classification. However, there is currently no precedent for combining multi-view hyperspectral imaging with hybrid machine learning / deep learning frameworks to address the dual challenges of authenticity verification and intraspecies classification of Anoectochilus roxburghii.
[0037] This embodiment aims to propose a method and system for authenticating and classifying the leaves of anoectochilus roxburghii plants. By combining hyperspectral imaging with machine learning, the system can be used to identify anoectochilus roxburghii and its counterfeit varieties in a non-destructive and high-precision manner, thereby overcoming the shortcomings of the above-mentioned existing methods. This not only promotes the application of hyperspectral technology in the authentication of medicinal plants, but also provides a scalable non-destructive solution for quality assurance in the herbal medicine market, helping to protect consumer health and promote sustainable practices in the traditional medicine industry.
[0038] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0039] like Figure 1 As shown, this embodiment provides a method for identifying the authenticity of leaves of a roxburghii plant and classifying its strains, which specifically includes the following steps.
[0040] Step S1: Acquire multi-view spectral data of a plant sample to be tested, wherein the multi-view spectral data includes front spectral data and back spectral data of each leaf of the plant sample to be tested.
[0041] In this embodiment, step S1 acquires multi-view spectral data of the plant sample to be tested, which specifically includes the following steps.
[0042] Step S11: cleaning the plant sample to be tested to obtain a cleaned plant sample to be tested.
[0043] Step S12: air-drying the cleaned plant sample to be tested to obtain an air-dried plant sample to be tested.
[0044] Step S13: collecting the front spectral data and the back spectral data of each leaf of the air-dried plant sample to be tested to obtain multi-view spectral data.
[0045] In this embodiment, step S13 collects the front spectral data and the back spectral data of each leaf of the air-dried plant sample to be tested to obtain multi-view spectral data, which specifically includes the following steps.
[0046] A GaiaField-N17E hyperspectral imaging system was used to scan the front and back of each leaf of the air-dried plant sample to be tested, respectively, to collect spectral data of the front and back of each leaf, and obtain multi-view spectral data.
[0047] Step S2: preprocess the multi-view spectral data to obtain preprocessed multi-view spectral data.
[0048] In this embodiment, step S2 pre-processes the multi-view spectral data to obtain pre-processed multi-view spectral data, which specifically includes the following steps.
[0049] Step S21 : performing black-white correction on the multi-view spectral data to obtain black-white corrected multi-view spectral data.
[0050] Step S22 : extracting a region of interest (ROI) from the black-and-white corrected multi-view spectral data to obtain pre-processed multi-view spectral data.
[0051] Step S3: Input the preprocessed multi-view spectral data into a pretrained SVM (support vector machine) model to output an authenticity verification result. The pretrained SVM model is a model obtained by jointly adjusting the parameters of the SVM model and the preprocessed model based on the preprocessed multi-view spectral data of the roxburghii sample and the counterfeit sample, finding the optimal parameters, and then training the model. Therefore, the pretrained SVM model is essentially a pre-trained SVM model.
[0052] In this embodiment, before training the pre-trained SVM model, the pre-processed multi-view spectral data is standardized using a Standard Scaler (standardization tool); when training the pre-trained SVM model, grid search and five-fold cross-validation methods are used to optimize hyperparameters, including penalty parameters, kernel type, gamma, and polynomial degree; and accuracy, precision, recall, F1-score, and confusion matrix are used as evaluation indicators to evaluate model performance.
[0053] In this embodiment, the preprocessing model refers to a model corresponding to a preprocessing method, specifically, a model corresponding to various filtering algorithms. The filtering algorithm is at least one of a median filter (MF) algorithm, an average filter (AF) algorithm, a Gaussian filter (GF) algorithm, a Savitzky-Golay (SG) filter algorithm, and a principal component analysis (PCA) algorithm.
[0054] Step S4: When the authenticity identification result is that the plant sample to be tested is a roxburghii sample, the pre-processed multi-view spectral data of the plant sample to be tested is input into the pre-trained CNN model, and the strain classification result is output. Among them, the pre-trained CNN model refers to a model obtained by jointly adjusting the parameters of the CNN model and the pre-processing model based on the pre-processed multi-view spectral data of the roxburghii sample, finding the optimal parameters, and then training. Therefore, the pre-trained CNN model is essentially a trained CNN model. In this embodiment, the 1D-CNN model is specifically used.
[0055] In order to make the technical solution of this embodiment clearer, the specific implementation process of the technical solution of this embodiment is described in detail below in the form of examples from the aspects of sample preparation, multi-view spectral data acquisition, data preprocessing, authenticity identification, strain classification and result output, which specifically includes the following implementation steps.
[0056] Step 1: Sample preparation.
[0057] The samples prepared in this embodiment include samples of Anoectochilus roxburghii and counterfeit samples. First, samples of Anoectochilus roxburghii and counterfeit samples were obtained, including nine different varieties of Anoectochilus roxburghii, including small round leaf (1), pointed leaf (2), red cloud (3), J6 male (4), colorful cloud (5), large round leaf (6), red cloud large leaf (7), gold vein 1 (8), and Taihong (9), as well as two counterfeit varieties, blood leaf orchid and spotted leaf orchid. Each variety includes 10 leaves of similar size and maturity to minimize biological variation.
[0058] In this embodiment, before collecting multi-view spectral data, the sample is cleaned with distilled water to remove surface contaminants and air-dried under laboratory conditions of 25° C. and 60% humidity before starting the multi-view spectral data collection in step 2.
[0059] Step 2: Multi-view spectral data acquisition.
[0060] This embodiment scans the front and back of each leaf, wherein the front is the adaxial surface and the back is the abaxial surface, to capture multi-view spectral information.
[0061] This embodiment uses the GaiaField-N17E model hyperspectral imaging system, such as Figure 2 As shown, the hyperspectral imaging system covers a spectral range of 900nm to 1700nm, with a spectral resolution of 5nm and a spatial resolution of 640 pixels across 512 bands. The system consists of an indoor test chamber, primarily comprising a hyperspectral camera, a support platform, a lifting platform, a base, and four halogen lamps. The lifting platform is mounted above the base, with support rods arranged around the base. The lifting platform is located directly below the top of the support frame. The hyperspectral camera and four halogen lamps are mounted on top of the support frame. The lifting platform controls the height of the Anoectochilus roxburghii samples and counterfeit samples. The support frame supports the hyperspectral camera and four halogen lamps. The hyperspectral camera collects multi-view spectral data, and four 50W halogen lamps provide stable illumination. The hyperspectral camera uses an array detector located perpendicular to the camera's direction of motion to scan the multi-view spectral data. This allows the camera to scan a two-dimensional space as the scanning platform of the hyperspectral imaging system advances. A conveyor belt, set at a speed of 0.8cm / s, can also be installed on the lifting platform to ensure a stable and uniform speed for samples as they pass through the scanning area. The hyperspectral camera's exposure time was set to automatic, and the gain factor was set to 1. The vertical distance between the sample and the lens was fixed at 42 cm. During scanning, the sample was placed on a black substrate to enhance image contrast and visibility while minimizing background diffuse reflection from interfering with the sample's spectral data. During data acquisition, the hyperspectral camera scanned each sample twice to improve data reliability and repeatability.
[0062] Step 3: Data preprocessing.
[0063] After the scanning is completed in this embodiment, the original hyperspectral image is subjected to black and white correction. The data captured by the hyperspectral imaging system mainly represents the signal intensity, and to accurately obtain the spectral reflectance, it must undergo precise calibration and black and white correction. The key to the calibration process is to match the image data with the actual spectral response to ensure that each pixel in the image accurately reflects the spectral characteristics of the target substance. This usually requires detailed measurement of the spectral response of the equipment and correction of systematic errors. The purpose of black and white correction is to eliminate background interference introduced by system noise and uneven lighting. By analyzing images acquired under no-light conditions (black field) behind the lens and using a standard reflector (white field), the image data can be corrected to ensure its accuracy and consistency. Calibrate the image Calculated by the following formula.
[0064]
[0065] in, represents the calibration image, represents the original hyperspectral image, represents the whiteboard image used for calibration, Represents the blackboard calibration image.
[0066] During the spectral data extraction process in this example, the black background in the rectified hyperspectral image is first removed to accurately extract the region of interest (ROI) for each sample. This step is achieved by selecting a grayscale image at a specific wavelength and applying the OTSU (Optimal Between-Class Variance) method to determine an optimal threshold. Regions exceeding the threshold are labeled "1," while background regions below the threshold are labeled "0." This method generates a binary mask for each sample and applies it to the entire hyperspectral image cube, effectively removing the black background. To improve prediction accuracy and reproducibility, the reflectance values of the rice / polished rice pixels in each hyperspectral image are averaged to extract the average spectrum for each sample. The average spectra extracted from the two hyperspectral images scanned for the same sample are then averaged again. The resulting result serves as the spectral data corresponding to that sample, providing the basis for subsequent analysis. Subsequently, spectral extraction of the ROI for each sample is performed using ENVI (Remote Sensing Image Processing Software). In the software ENVI, each sample is divided into four equal parts using a rectangular frame, and one quarter of each sample is randomly selected as the region of interest. Each pixel in the region of interest contains a different set of spectral information. The final spectral value of the sample can be obtained by taking the average of the spectral reflectance of all pixels in the region.
[0067] This embodiment can use a variety of filtering techniques such as median filtering algorithm, average filtering algorithm, Gaussian filtering algorithm, Savitzky-Golay filtering algorithm, principal component analysis method, etc. to preprocess the hyperspectral data to reduce noise and improve data quality, ensuring the accuracy and robustness of subsequent analysis.
[0068] The median filter algorithm is a nonlinear filtering technique that replaces each pixel value with the median of its neighborhood. This method is particularly effective at removing salt and pepper noise. The filtering process involves specifying a window size to determine the neighborhood of each pixel, then calculating the median value within that neighborhood and replacing the original pixel value. Median filtering preserves edge detail and is particularly suitable for removing impulse noise without excessive image blurring. The core parameter of the median filter is kernel_size, which determines the range of the neighborhood used during filtering. Kernel_size is included in the hyperparameter search range (with values of 3, 5, 7, and 9) and is automatically tuned using Grid Search CV. Therefore, during training, Grid Search CV iterates over different kernel_sizes and, in combination with other classifier hyperparameters, selects the optimal parameter combination through cross-validation.
[0069] The average filter algorithm works by replacing each pixel value with the average of its neighboring pixels. It effectively reduces random noise, but may blur edges, making it less effective at preserving fine detail. This method is computationally simple and efficient, making it a popular choice for initial denoising of hyperspectral data. In the model, the average filter is encapsulated as a custom transformer and integrated into a pipeline, where it is used in conjunction with a classifier. Pipeline is a tool provided by the scikit-learn machine learning library that combines multiple data preprocessing and model training steps into a unified workflow, ensuring that data passes through each step in a predetermined order. It can also be combined with Grid Search CV (Grid Search Cross Validation) to automatically tune hyperparameters. Tunable parameters include window_size, which controls the neighborhood range for calculating the mean for each data point. In Grid Search CV, the window_size range for automatic parameter tuning is set to {2, 3, 4, 5}, meaning that a grid search will test these different window sizes and select the optimal value. Throughout the entire process, Grid Search CV cross-validates different window_size combinations and combines them with other classifier hyperparameters to find the optimal parameter configuration. The best model ultimately trained uses the optimal window_size, ensuring that the filtered spectral data is better suited to the classification task and improving classification accuracy.
[0070] The Gaussian filter algorithm is a common smoothing technique that uses a Gaussian function to weight adjacent pixel values, assigning higher weights to pixels closer to the center. Compared to the average filter algorithm, the Gaussian filter algorithm smoothes the image while preserving more detail in the central region. The Gaussian filter algorithm is widely used in hyperspectral data preprocessing because it effectively reduces noise while minimizing blurring. This example defines a custom Gaussian filter transformer that uses the gaussian_filter function in the scipy.ndimage module to smooth spectral data. The sigma parameter of this Gaussian filter transformer controls the degree of smoothing: smaller sigma values preserve more detail, while larger sigma values produce a stronger smoothing effect. To optimize the impact of the Gaussian filter parameters on the classification task, this example integrates the Gaussian filter transformer into the pipeline as a data preprocessing step and uses Grid Search CV for hyperparameter search. This example selects multiple candidate values for the sigma parameter (0.5, 1, 2, 3, and 4) and performs a grid search on different classifier hyperparameters to find the optimal parameter combination. Finally, this embodiment evaluates the effect of the Gaussian filtering algorithm through the classification results on the test set, and uses accuracy, precision, recall rate and F1-score as measurement indicators.
[0071] The Savitzky-Golay filter algorithm is a smoothing method commonly used in signal processing. It applies polynomial fitting within a sliding window to smooth data, thereby reducing high-frequency noise. In hyperspectral data processing, the Savitzky-Golay filter algorithm helps smooth small random fluctuations while preserving key signal features. Key parameters of the Savitzky-Golay filter algorithm include the window length (window_length) and the polynomial order (polyorder). The window length defines the size of the sliding window and must be an odd number. It determines the number of spectral points considered during the filtering process. Smaller window lengths better preserve detail but have weaker noise immunity, while larger window lengths improve smoothing but may result in loss of detail. In this example, three window sizes of 5, 7, and 9 are selected for automatic parameter adjustment. The polynomial order refers to the order of the polynomial used to fit the data within the sliding window. Lower orders produce smoother filtering but may lose some spectral information. Higher orders more accurately preserve local variations in the spectral curve. In this example, orders of 2, 3, and 4 are selected for experimental comparison. By automatically tuning parameters using Grid Search CV and combining them with the classification model, we evaluated the classification performance under different parameter combinations to determine the optimal SG filter parameters. The parameters ultimately selected were able to reduce noise while preserving the key features of the spectral data, thereby improving the accuracy of the classification model.
[0072] The PCA algorithm is an unsupervised statistical method for data dimensionality reduction. It extracts the most important features, called principal components, by maximizing the variance of the data. In spectral analysis, the PCA algorithm can help identify differences between samples and potential chemical changes. In this example, the PCA algorithm is used for dimensionality reduction to reduce the dimensionality of the spectral data, improve the computational efficiency of the model, and reduce data redundancy and noise. The parameters of the PCA transformer include n_components (the number of principal components), pca.fit(X), and pca.transform(X). n_components defaults to 10 and represents the dimensionality of the data after dimensionality reduction. This parameter determines the number of principal components retained and is typically selected based on the cumulative variance contribution (e.g., above 95%). pca.fit(X) is used to fit the training data and calculate the principal component directions. pca.transform(X) is used to project the original data into the principal component space to achieve dimensionality reduction. In the Grid Search CV process, the PCA algorithm first reduces the dimension of the spectral data, and then processes the data in combination with the Standard Scaler. Combined with other hyperparameters of the classifier, the optimal parameter combination is selected through cross-validation. This not only reduces the complexity of the data but also improves the generalization ability of the model.
[0073] Step 4: Authenticity identification.
[0074] This embodiment uses an SVM model to distinguish nine varieties of golden thread vine from counterfeits based on hyperspectral data. The SVM model can determine the optimal hyperplane that maximizes the margin between different categories and has strong generalization capabilities. The radial basis function (rbf) kernel was selected because it is very effective in processing nonlinear data patterns. Before training, the spectral data was standardized using the Standard Scaler to ensure the consistency of the feature scale. Hyperparameters, including penalty parameters, kernel type, gamma, and polynomial degree, were optimized using grid search and five-fold cross-validation methods. Model performance was evaluated using accuracy, precision, recall, F1 score, and confusion matrix to comprehensively evaluate classification accuracy and error. The specific steps are as follows.
[0075] First, the data set is divided. The data source is the experimental data stored in an Excel file, where the first row is the spectral wavelength feature; the second to the 81st rows are spectral data, with a total of 80 samples, each containing 440-dimensional spectral features. For label construction, true samples (category 0): 40 samples; fakes (category 1): 40 samples. For standardization, since the numerical range of spectral data is large, this embodiment uses Standard Scaler for standardization (mean is 0, standard deviation is 1) to ensure that all features are on the same scale, which is conducive to the stable training of SVM. For the division of training set / test set, 80% training set (64 samples), 20% test set (16 samples). stratify=labels ensures that the proportions of the two types of samples in the training set and test set are consistent, improving the generalization ability of the model.
[0076] The SVM classifier structure maximizes the margin between data classes by finding a hyperplane. Key parameters include the kernel function (kernel), linear (linear kernel), RBF (radial basis function), regularization parameter (C), gamma parameter, and the degree of the polynomial kernel. Linear is suitable for linearly separable data and is computationally simple. RBF is suitable for nonlinear data and can capture complex patterns. The regularization parameter, also known as the penalty parameter, controls the model's tolerance for misclassification. A larger regularization parameter reduces the model's tolerance for misclassification and tightens the decision boundary. This example combines four regularization parameter values from [0.1, 1, 10, 100] with other parameters. The gamma parameter applies only to the radial basis function kernel and controls the extent to which a single training example influences the classification boundary. A larger gamma parameter reduces the extent to which a single training example influences the classification boundary and increases the complexity of the decision boundary. In this example, the kernel function coefficients ('gamma': ['scale', 'auto']) are automatically selected. The degree of the polynomial kernel is only valid when kernel='poly'. It controls the order of the polynomial kernel and affects the complexity of the decision boundary. In this embodiment, three polynomial kernel numbers are selected: [3, 4, 5].
[0077] During the training process, this embodiment defines the SVM classifier as SVM=SVC(). For hyperparameter optimization, different hyperparameter combinations are tested under five-fold cross validation.
[0078] Grid Search CV is used to traverse all possible parameter combinations and select the best parameters based on the cross-validation accuracy score.
[0079] This embodiment uses the following evaluation metrics: Accuracy: The proportion of correctly classified items. Precision: The proportion of positive items predicted by the model that are actually correct. Recall: The proportion of positive items correctly identified by the model. F1-score: The weighted harmonic mean of precision and recall. Confusion matrix: Used to visually display the correct and incorrect classifications predicted by the model. Based on the above evaluation metrics, this embodiment uses the optimal SVM model for prediction and evaluation using the test set.
[0080] The key optimization strategy is to first standardize the data. Since SVM is sensitive to feature scale, StandardScaler is used for standardization to improve classification performance. Hyperparameter optimization is then performed using Grid Search CV cross-validation to automatically select the optimal regularization parameter, gamma, and kernel function to prevent overfitting or underfitting. For kernel function selection, linear kernels are suitable for simple classification problems, while RBF kernels are suitable for more complex nonlinear classification tasks. Five-fold cross-validation is used during cross-validation to ensure model stability on different data subsets and improve generalization.
[0081] This example effectively classifies spectral data from authentic and counterfeit Anoectochilus roxburghii leaves using an SVM model combined with grid search CV-based hyperparameter optimization. Standardized preprocessing improved model stability, and cross-validation was used to select the optimal hyperparameters, ultimately achieving 100% classification accuracy.
[0082] Step 5: Strain classification.
[0083] After verifying the authenticity of roxburghii leaves, this example then classified the authentic roxburghii leaves into different strains. This example compared the CNN model's performance in roxburghii strain classification with multiple models, including the SVM model, the KNN (K-Nearest Neighbors) model, and the LDA (Linear Discriminant Analysis) model. Furthermore, derivative models derived from the SVM, KNN, and LDA models after filtering using algorithms such as the MF algorithm, AF algorithm, GF algorithm, SG algorithm, and PCA algorithm were incorporated into the strain classification performance comparison experiment. Table 1 shows the results of the strain classification performance comparison between the CNN model and these various models.
[0084] Table 1 Comparison of strain classification performance of different models
[0085]
[0086] It can be seen intuitively from Table 1 that when the spectra of the front and back of the leaves are merged, the accuracy, precision, recall rate and F1 score of the CNN model are all 1, and the values of each indicator are all the highest. This proves that the CNN model has great advantages in the variety classification task of the front and back hyperspectral fusion of the leaves of the golden thread lotus, and can accurately realize the variety classification of the golden thread lotus. Therefore, this embodiment uses the CNN model to classify nine golden thread lotus varieties. In order to achieve the classification task of hyperspectral data, this embodiment constructs a 1D-CNN model, which includes multiple convolutional layers, pooling layers and fully connected layers, and uses L1 regularization to improve generalization ability.
[0087] This example first reads hyperspectral data from an Excel file and applies the various preprocessing methods used in step 3, including median filtering, mean filtering, Gaussian filtering, Savitzky-Golay smoothing filtering, and PCA preprocessing, to improve data quality. The preprocessed dataset contains 360 samples from nine categories, each containing 40 samples, for a total of 900 samples. The dataset is partitioned into 80% training and 20% test sets and then normalized.
[0088] like Figure 3As shown in the figure, the input layer of the 1D-CNN model has an input data shape of (n, 1), where n is the length of the spectral data for each sample. The first layer of the 1D-CNN model is a convolutional layer consisting of 64 7×1 convolutional filters, activated by the ReLU activation function with L1 regularization and a regularization parameter of λ=0.001. The second layer is a max pooling layer with a pooling window size of 2. The third layer is a convolutional layer consisting of 128 7×1 convolutional filters, activated by the ReLU activation function with L1 regularization. This is followed by a flattening layer, which flattens the data into a one-dimensional vector. This is followed by a fully connected layer (Dense), consisting of 128 neurons, activated by the ReLU activation function with L1 regularization and a regularization parameter of λ=0.001. Finally, the output layer is also a Dense layer, consisting of 10 neurons (representing the nine categories) and using the softmax activation function.
[0089] During the optimization and training process of this embodiment, sparse categorical crossentropy was used as the loss function. The Adam optimizer (learning rate = 0.0001) was used as the optimizer. The loss function was sparse_categorical_crossentropy, and the metric was accuracy. The training strategy was a batch size of 32, 500 epochs of training, and 20% of the training data was used as a validation set. Standard Scaler was used for data preprocessing to standardize the data. For data augmentation and denoising, Gaussian filtering (σ = 7) and singular value decomposition (SVD, k = 4) were used for data denoising. The data was divided into 80% for training and 20% for testing. During model evaluation, the Accuracy_Score function was used to calculate the classification accuracy of the model on the test set. During training visualization, loss and accuracy curves were plotted to observe training convergence.
[0090] This example uses a 1D-CNN model to effectively extract features from spectral data, combining regularization and optimization algorithms to improve classification accuracy. By plotting training loss and accuracy curves, this example observed that the 1D-CNN model gradually converged during training and achieved high accuracy on the validation set, demonstrating its good generalization capabilities. Figure 4 This is a graph showing how the loss value of the loss function changes with the number of training times during the training process. Figure 5 This is a curve graph showing the change of accuracy with the number of training times during the training process. Figure 6 This is a diagram of the confusion matrix of the trained 1D-CNN model on the test set. Figure 4 The loss function in uses sparse categorical cross entropy loss, which is suitable for multi-class classification problems, especially when the labels are encoded as integers. The loss function formula is as follows:
[0091]
[0092] in, is the loss value, is the total number of categories, For the The true label of the class (if the sample belongs to class, then it is 1, otherwise it is 0), For the model The predicted probability of the class.
[0093] according to Figure 4 It can be seen that with the increase in the number of training times, the training set loss value (Training Loss) and the validation set loss value (Validation Loss) both decrease significantly and tend to be stable, indicating that the learning process of the 1D-CNN model is effective. Figure 5 This shows the accuracy trend during training. Figure 5 It can be seen that as the training progresses, the training set accuracy (Training Accuracy) and the validation set accuracy (Validation Accuracy) gradually approach 1, and the fluctuation becomes smaller and smaller, further proving the stability and reliability of the 1D-CNN model. Figure 6 The confusion matrix in Figure 2 shows that the 1D-CNN model correctly classified all samples in each category, with no misclassifications. The true label refers to the sample's true category, while the predicted label refers to the sample's category predicted by the 1D-CNN model. This result demonstrates the superior performance of the multi-view spectral fusion model, which successfully leverages spectral data from both the front and back of the leaf to improve classification accuracy.
[0094] like Figure 6As shown, the 100% accuracy can be attributed to factors such as complementarity, data augmentation, and feature diversity in model architecture and optimization. By using spectral data from both the front and back sides of leaves, the model leverages multi-view information and provides a richer feature representation. The spectral responses of the front and back sides of leaves may differ in certain physical properties, such as light reflection and scattering. These differences help capture the diverse structures and compositions of leaves, ultimately improving classification accuracy. The complementary nature of spectral data from the front and back sides enhances the robustness and accuracy of the model. Furthermore, different leaf species exhibit significant spectral differences, particularly in specific wavelength regions. The front and back side spectra provide information from different angles, helping to better capture these subtle spectral differences. By incorporating this additional feature dimension (i.e., front and back side spectra), the model gains more information, thereby improving its generalization ability. The introduction of feature diversity enables the model to identify more underlying patterns during training, thereby improving classification accuracy.
[0095] Compared with traditional manual feature extraction methods, deep learning models such as CNN models can automatically learn the best feature combination through end-to-end training, thereby achieving more accurate classification. As the number of layers and neurons in the network increases, the model can process more complex spectral data and extract deeper features. Overall, the qualitative model used in this embodiment showed excellent training performance and generalization ability. These results highlight the potential of CNN-based models in hyperspectral data classification, especially when using multi-view spectral fusion. The model can effectively capture the subtle spectral differences between Anoectochilus roxburghii varieties, highlighting the effectiveness of deep learning technology in solving complex classification tasks in the field of plant species identification.
[0096] Step 6: Output the results.
[0097] In this embodiment, the final output result is the strain classification result of the real golden thread lotus leaf, corresponding to the varieties (1) to (9) of the golden thread lotus sample in step one, thereby judging whether the golden thread lotus leaf is small round leaf (1), pointed leaf (2), Hongxia (3), J6 male (4), Caixia (5), large round leaf (6), Hongxia large leaf (7), Jinmai No. 1 (8) or Taihong (9).
[0098] This example combines hyperspectral imaging and machine learning technologies for high-precision classification and authenticity verification of roxburghii and its counterfeits. Hyperspectral data were collected from the front and back surfaces of nine roxburghii varieties and two counterfeits (blood leaf orchid and spotted leaf orchid). The SVM model was then used for authenticity verification. Experimental results show that the SVM model achieved 100% classification accuracy in distinguishing roxburghii from its counterfeits, effectively capturing the spectral differences between the front and back surfaces. In the authenticity verification process, the SVM model, leveraging spectral data from both front and back leaf surfaces, achieved 100% accuracy in distinguishing roxburghii from its counterfeits. In the classification of different roxburghii varieties, the introduction of a multi-view spectral data fusion CNN model significantly improved classification performance and robustness by combining spectral data from both leaf surfaces. This model leverages complementary information to achieve classification, highlighting the potential of hyperspectral imaging and machine learning technologies for plant leaf authenticity verification and strain classification, providing a new perspective for plant authenticity verification and strain classification.
[0099] In an exemplary embodiment, a system for identifying the authenticity of leaves of a roxburghii plant and classifying its strains is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method for identifying the authenticity of leaves of a roxburghii plant and classifying its strains.
[0100] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0101] The present embodiment uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above examples is only intended to help understand the method and core concept of the present application. At the same time, for those skilled in the art, according to the concept of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present application.
Claims
1. A method for identifying the authenticity of leaves of aroectochilus roxburghii plant and classifying its strains, characterized in that: include: Acquiring multi-view spectral data of the plant sample to be tested; The multi-view spectral data includes the front spectral data and the back spectral data of each leaf of the plant sample to be tested; Preprocessing the multi-view spectral data to obtain preprocessed multi-view spectral data; performing black-white correction on the multi-view spectral data to obtain black-white corrected multi-view spectral data; extracting a region of interest from the black-white corrected multi-view spectral data to obtain preprocessed multi-view spectral data; The pre-processed multi-view spectral data is input into the pre-trained SVM model, and the authenticity identification result is output; the pre-trained SVM model refers to the multi-view spectral data after pre-processing based on the golden thread lotus sample and the counterfeit sample, the SVM model and the pre-processing model are jointly adjusted, and the model is obtained after training after finding the optimal parameters, and the pre-processing model refers to the model corresponding to the filtering algorithm; the filtering algorithm is at least one of the median filtering algorithm, the average filtering algorithm, the Gaussian filtering algorithm, the Savitzky-Golay filtering algorithm and the principal component analysis method; before training the pre-trained SVM model, the pre-processed multi-view spectral data is standardized by using Standard Scaler; when training the pre-trained SVM model, the hyperparameters are optimized by grid search and five-fold cross validation methods, including penalty parameters, kernel type, gamma and polynomial degree; and the model performance is evaluated by using accuracy, precision, recall rate, F1 score and confusion matrix as evaluation indicators; When the authenticity identification result is that the plant sample to be tested is a roxburghii sample, the preprocessed multi-view spectral data of the plant sample to be tested is input into the pre-trained CNN model, and the variety classification result is output; the pre-trained CNN model refers to a model obtained by jointly adjusting the parameters of the CNN model and the preprocessing model based on the preprocessed multi-view spectral data of the roxburghii sample, and training the model after finding the optimal parameters; the pre-trained CNN model is a 1D-CNN model.
2. The method for identifying the authenticity of leaves of aroectochilus roxburghii plant and classifying strains according to claim 1, characterized in that: Obtain multi-view spectral data of the plant sample to be tested, including: Cleaning the plant sample to be tested to obtain a cleaned plant sample to be tested; air-drying the cleaned plant sample to be tested to obtain an air-dried plant sample to be tested; The front spectral data and the back spectral data of each leaf of the air-dried plant sample to be tested are collected to obtain multi-view spectral data.
3. The method for identifying the authenticity of leaves of anoectochilus roxburghii plant and classifying strains according to claim 2, characterized in that: Collecting the front spectral data and the back spectral data of each leaf of the air-dried plant sample to be tested to obtain multi-view spectral data, specifically including: A GaiaField-N17E hyperspectral imaging system was used to scan the front and back of each leaf of the air-dried plant sample to be tested, respectively, to collect spectral data of the front and back of each leaf, and obtain multi-view spectral data.
4. A system for identifying the authenticity of leaves of aroectochilus roxburghii and classifying strains, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and runnable on the processor, wherein the processor executes the computer program to implement the method for authenticity identification and variety classification of leaves of anoectochilus plant according to any one of claims 1 to 3.