Anoectochilus formosanus plant leaf authenticity identification and strain classification method and system

Through the combination of multi-view spectral data with SVM and CNN models, the problem of low accuracy in authenticity identification and strain classification of leaves of genus Morifume plants is solved, and high-precision and non-destructive identification and classification are achieved, reducing operational complexity and cost.

CN120213831AActive Publication Date: 2025-06-27ZHEJIANG FORESTRY UNIVERSITY

Patent Information

Application Number
CN202510694179.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The prior art has low accuracy, strong subjectivity, complex operation and expensive in the identification of authenticity and strain classification of leaves of genus filament plants.

Method used

Multi-view spectral data combined with support vector machine (SVM) and convolutional neural network (CNN) models are used to obtain the front and back spectral data of each leaf, and perform preprocessing and feature extraction to achieve authenticity identification and strain classification of leaves of the nigra plant.

Benefits of technology

The accuracy of authenticity identification and strain classification of leaves of nigra plant is improved, and non-destructive and high-precision identification and classification are achieved, reducing operational complexity and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120213831A_ABST
    Figure CN120213831A_ABST
Patent Text Reader

Abstract

The invention discloses an anoectochilus formosanus plant leaf authenticity identification and strain classification method and system, and relates to the technical field of plant leaf identification and classification, and the method comprises the following steps: obtaining multi-view spectral data of a to-be-detected plant sample; preprocessing the multi-view spectral data; inputting the preprocessed multi-view spectral data into a pre-trained SVM model, and outputting an authenticity identification result; and when the authenticity identification result is that the plant sample to be detected is the anoectochilus formosanus sample, inputting the preprocessed multi-view spectral data of the plant sample to be detected into the pre-trained CNN model, and outputting a strain classification result. According to the method, the multi-view spectral data, the SVM model and the CNN model are combined and applied to scenes of true and false identification and strain classification of the anoectochilus formosanus plant leaves, so that the accuracy of true and false identification and strain classification of the anoectochilus formosanus plant leaves can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of plant leaf identification and classification, and in particular to a method and system for authenticity identification and strain classification of leaves of a golden thread vine plant. Background Art

[0002] At present, the quality identification of Anoectochilus roxburghii usually relies on chemical analysis methods, mainly including visual inspection, microscopic identification, high performance liquid chromatography, DNA molecular identification and near infrared spectroscopy detection technology, etc. However, these methods are generally subjective and have low accuracy. Therefore, how to improve the accuracy of authenticity identification and strain classification of Anoectochilus roxburghii plant leaves has become a technical problem that needs to be solved in this field. Summary of the invention

[0003] The purpose of this application is to provide a method and system for authenticity identification and strain classification of leaves of a golden thread vine plant, which can effectively improve the accuracy of authenticity identification and strain classification of leaves of a golden thread vine plant.

[0004] To achieve the above objectives, this application provides the following solutions.

[0005] In a first aspect, the present application provides a method for authenticity identification and strain classification of leaves of a roxburghii plant, which specifically includes the following steps.

[0006] Acquire multi-viewing angle spectral data of the plant sample to be tested; the multi-viewing angle spectral data includes the front spectral data and the back spectral data of each leaf of the plant sample to be tested.

[0007] The multi-viewing angle spectral data is preprocessed to obtain preprocessed multi-viewing angle spectral data.

[0008] The preprocessed multi-view spectral data is input into the pre-trained SVM model to output the authenticity identification result; the pre-trained SVM model refers to the preprocessed multi-view spectral data based on the golden thread lotus sample and the counterfeit sample, and the SVM model and the pre-processing model are jointly adjusted to find the optimal parameters and then trained to obtain the model, and the pre-processing model refers to the model corresponding to the filtering algorithm.

[0009] When the authenticity identification result is that the plant sample to be tested is a golden thread vine sample, the preprocessed multi-view spectral data of the plant sample to be tested is input into the pre-trained CNN model, and the variety classification result is output; the pre-trained CNN model refers to a model obtained by jointly adjusting the parameters of the CNN model and the preprocessing model based on the preprocessed multi-view spectral data of the golden thread vine sample, and then training the model after finding the optimal parameters.

[0010] Optionally, obtaining multi-view spectral data of the plant sample to be tested specifically includes the following steps.

[0011] The sample of the plant to be measured is cleaned to obtain a cleaned sample of the plant to be measured.

[0012] The cleaned sample of the plant to be measured is air-dried to obtain an air-dried sample of the plant to be measured.

[0013] The front spectral data and back spectral data of each leaf of the air-dried sample of the plant to be measured are collected to obtain multi-view spectral data.

[0014] Optionally, collecting the front spectral data and back spectral data of each leaf of the air-dried sample of the plant to be measured to obtain multi-view spectral data specifically includes the following steps.

[0015] Using a hyperspectral imaging system of the GaiaField-N17E model, the front and back of each leaf of the air-dried sample of the plant to be measured are scanned respectively to collect the front spectral data and back spectral data of each leaf, and multi-view spectral data is obtained.

[0016] Optionally, the multi-view spectral data is preprocessed to obtain preprocessed multi-view spectral data, which specifically includes the following steps.

[0017] The multi-view spectral data is corrected for black and white to obtain black-and-white corrected multi-view spectral data.

[0018] The region of interest of the black-and-white corrected multi-view spectral data is extracted to obtain preprocessed multi-view spectral data.

[0019] Optionally, before training the pre-trained SVM model, the preprocessed multi-view spectral data is standardized using Standard Scaler; when training the pre-trained SVM model, the hyperparameters, including the penalty parameter, kernel type, gamma, and polynomial degree, are optimized using the methods of grid search and five-fold cross-validation; and the accuracy, precision, recall, F1 score, and confusion matrix are used as evaluation indicators for model performance evaluation.

[0020] Optionally, the filtering algorithm is at least one of the median filtering algorithm, average filtering algorithm, Gaussian filtering algorithm, Savitzky-Golay filtering algorithm, and principal component analysis method.

[0021] Optionally, the pre-trained CNN model is a 1D-CNN model.

[0022] In a second aspect, the present application provides a system for authenticating the genuineness and classifying the strains of Anoectochilus roxburghii plant leaves, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the method for authenticating the genuineness and classifying the strains of Anoectochilus roxburghii plant leaves described above.

[0023] According to the specific embodiments provided by the present application, the present application has the following technical effects: The present application provides a method and a system for authenticating the genuineness and classifying the strains of Anoectochilus roxburghii plant leaves. By combining multi-perspective spectral data, an SVM model, and a CNN model, and applying them to the scenario of authenticating the genuineness and classifying the strains of Anoectochilus roxburghii plant leaves, an effective combination of hyperspectral imaging technology and machine learning technology is achieved. Among them, by obtaining the front spectral data and the back spectral data of each leaf of the plant sample to be tested, multi-perspective spectral data is formed. On this basis, in terms of authenticating the genuineness of Anoectochilus roxburghii and counterfeit varieties, by using a pre-trained SVM model and utilizing the spectral data of the front and back leaves, Anoectochilus roxburghii and counterfeit varieties can be accurately distinguished, and the accurate authentication of the genuineness of Anoectochilus roxburghii plant leaves can be realized. In terms of classifying different varieties of Anoectochilus roxburghii, by introducing a pre-trained CNN model and combining the spectral data of the front and back leaves, the classification accuracy of the model can be effectively improved, and the accuracy of strain classification can be enhanced. Therefore, by combining multi-perspective spectral data, an SVM model, and a CNN model, the present application can effectively improve the accuracy of authenticating the genuineness and classifying the strains of Anoectochilus roxburghii plant leaves. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0025] Figure 1 It is a schematic flowchart of a method for authenticating the genuineness and classifying the strains of Anoectochilus roxburghii plant leaves provided by an embodiment of the present application.

[0026] Figure 2 It is a schematic structural diagram of a hyperspectral imaging system provided by an embodiment of the present application.

[0027] Figure 3 It is a schematic structural diagram of a 1D-CNN model provided by an embodiment of the present application.

[0028] Figure 4 It is a curve graph showing the change of the loss value of the loss function with the number of training times during the training process provided by an embodiment of the present application.

[0029] Figure 5 A graph showing the change of accuracy with the number of training times during the training process provided by an embodiment of the present application.

[0030] Figure 6 A schematic diagram of the confusion matrix of the 1D-CNN model after training provided by an embodiment of the present application on the test set. Detailed implementation manners

[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0032] Currently, methods for identifying medicinal plants such as Anoectochilus roxburghii, such as visual inspection, microscopic identification, high-performance liquid chromatography, DNA molecular identification, etc., need to be carried out under the guidance of professionals, and have defects such as strong subjectivity, time-consuming process, low accuracy, complex operation process and high cost. Not only is the labor intensity high, but also the sample will be damaged. Most importantly, it is often impossible to distinguish subtle interspecific variations.

[0033] In recent years, the progress of spectral technology is expected to solve the above problems. For example, in related technologies, the partial least squares discriminant analysis (PLS-DA) model is established by using near-infrared spectroscopy combined with chemometrics, realizing the accurate distinction between genuine Anoectochilus roxburghii powder and its two counterfeits. In related technologies, there is also the use of near-infrared spectroscopy to obtain the spectral data of Anoectochilus roxburghii and its adulterated products, and an improved one-dimensional convolutional neural network (1D-CNN) model is designed to process the near-infrared spectral data to distinguish true and false Anoectochilus roxburghii. In addition, there is also a rapid and accurate classification model for Anoectochilus roxburghii varieties based on a handheld near-infrared spectrometer and AdaBoost (Adaptive Boosting) ensemble learning in related technologies, and the classification accuracy rate is as high as 95.6%. These studies show that near-infrared spectroscopy has high accuracy in detecting Anoectochilus roxburghii, but its application is usually limited to the powder form of plants, and requires grinding, compression or other sample preparation methods.

[0034] Hyperspectral Imaging (HSI) can capture spatial and spectral information in hundreds of continuous bands, maintain the integrity of the sample, capture a wide range of spectral bands in the electromagnetic spectrum, provide detailed spectral information, and can be used for precise material identification and quality assessment. In related technologies, there is a hyperspectral model that uses spectral data analysis technology to detect the flavonoid and polysaccharide contents in Anoectochilus roxburghii. However, its potential in authenticity identification and multi-variety classification has not been explored yet. In addition, existing studies usually rely on single-view spectral data (e.g., the front side of the leaf), while ignoring the complementary information embedded in the multi-view perspective (e.g., the front and back sides of the leaf).

[0035] Machine Learning (ML) and Deep Learning (DL) have revolutionized pattern recognition in high-dimensional datasets, making them ideal for analyzing hyperspectral data. Although traditional machine learning models are widely used, their performance depends to a large extent on manual feature selection and preprocessing. In contrast, deep learning architectures like Convolutional Neural Networks (CNNs) can autonomously extract discriminative features, enabling end-to-end classification. However, there is currently no precedent in the existing technology to combine multi-view hyperspectral imaging with a hybrid machine learning / deep learning framework to address the dual challenges of authenticity verification and intra-species classification of Anoectochilus roxburghii.

[0036] This embodiment aims to propose a method and system for authenticating and classifying the leaves of Anoectochilus roxburghii plants, which combines hyperspectral imaging with machine learning to identify Anoectochilus roxburghii and its counterfeit varieties in a non-destructive and high-precision manner, thus making up for the deficiencies of the above existing methods, not only promoting the application of hyperspectral in the authentication of medicinal plants, but also providing a scalable non-destructive solution for quality assurance in the herbal medicine market, helping to ensure consumer health, and promoting sustainable practices in the traditional medicine industry.

[0037] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0038] As Figure 1 shown, this embodiment provides a method for authenticating and classifying the leaves of Anoectochilus roxburghii plants, which specifically includes the following steps.

[0039] Step S1: Obtain multi-view spectral data of the plant sample to be measured. Among them, the multi-view spectral data includes the front spectral data and the back spectral data of each leaf of the plant sample to be measured.

[0040] In this embodiment, step S1 for obtaining multi-view spectral data of the plant sample to be measured specifically includes the following steps.

[0041] Step S11: Clean the sample of the plant to be measured to obtain the cleaned sample of the plant to be measured.

[0042] Step S12: Air-dry the cleaned sample of the plant to be measured to obtain the air-dried sample of the plant to be measured.

[0043] Step S13: Collect the front spectral data and the back spectral data of each leaf of the air-dried sample of the plant to be measured to obtain multi-view spectral data.

[0044] In this embodiment, step S13 collects the front spectral data and the back spectral data of each leaf of the air-dried sample of the plant to be measured to obtain multi-view spectral data, which specifically includes the following steps.

[0045] Use a hyperspectral imaging system of the GaiaField-N17E model to scan the front and back of each leaf of the air-dried sample of the plant to be measured respectively, so as to collect the front spectral data and the back spectral data of each leaf and obtain multi-view spectral data.

[0046] Step S2: Preprocess the multi-view spectral data to obtain the preprocessed multi-view spectral data.

[0047] In this embodiment, step S2 preprocesses the multi-view spectral data to obtain the preprocessed multi-view spectral data, which specifically includes the following steps.

[0048] Step S21: Perform black and white correction on the multi-view spectral data to obtain the multi-view spectral data after black and white correction.

[0049] Step S22: Extract the region of interest (ROI) from the multi-view spectral data after black and white correction to obtain the preprocessed multi-view spectral data.

[0050] Step S3: Input the preprocessed multi-view spectral data into a pre-trained SVM (support vector machine) model and output the authenticity discrimination result. Among them, the pre-trained SVM model refers to a model obtained by jointly tuning the parameters of the SVM model and the preprocessing model based on the preprocessed multi-view spectral data of the samples of Anoectochilus roxburghii and counterfeit samples, and then training after finding the optimal parameters. Therefore, the pre-trained SVM model is essentially a trained SVM model.

[0051] In this embodiment, before training the pre-trained SVM model, the pre-processed multi-view spectral data is standardized using Standard Scaler; when training the pre-trained SVM model, the hyperparameters are optimized by grid search and five-fold cross-validation, including the penalty parameter, kernel type, gamma, and polynomial degree; and the accuracy, precision, recall, F1-score, and confusion matrix are used as evaluation metrics to evaluate the model performance.

[0052] In this embodiment, the preprocessing model refers to the model corresponding to the preprocessing method, specifically the models corresponding to various filtering algorithms. Among them, the filtering algorithms are at least one of the median filtering (MF) algorithm, average filtering (AF) algorithm, Gaussian filtering (GF) algorithm, Savitzky-Golay (SG) filtering algorithm, and principal component analysis (PCA).

[0053] Step S4: When the authenticity identification result is that the tested plant sample is a Anoectochilus roxburghii sample, the pre-processed multi-view spectral data of the tested plant sample is input into the pre-trained CNN model, and the strain classification result is output. Among them, the pre-trained CNN model refers to the model obtained by jointly tuning the parameters of the CNN model and the preprocessing model based on the pre-processed multi-view spectral data of the Anoectochilus roxburghii sample, and then training after finding the optimal parameters. Therefore, the pre-trained CNN model is essentially a trained CNN model. In this embodiment, the 1D-CNN model is specifically used.

[0054] To make the technical solution of this embodiment clearer, the following takes the form of examples to detail the specific implementation process of the technical solution of this embodiment from aspects such as sample preparation, multi-view spectral data acquisition, data preprocessing, authenticity identification, strain classification, and result output, specifically including the following implementation steps.

[0055] Step 1: Sample preparation.

[0056] The samples prepared in this embodiment include Anoectochilus roxburghii samples and counterfeit samples. First, Anoectochilus roxburghii samples and counterfeit samples are obtained, including 9 different varieties of Anoectochilus roxburghii, namely small round leaf (1), pointed leaf (2), Hongxia (3), J6 male (4), Caixia (5), large round leaf (6), large Hongxia leaf (7), Jinmai No. 1 (8), and Taihong (9), as well as 2 counterfeit varieties, Ludisia discolor and Goodyera schlechtendaliana. Among them, each variety includes 10 leaves of similar size and maturity to minimize biological variation.

[0057] Before collecting multi-view spectral data in this embodiment, the sample is cleaned with distilled water to remove surface contaminants and air-dried under laboratory conditions of 25 °C and 60% humidity, and then the collection of multi-view spectral data in Step 2 is started.

[0058] Step 2: Collection of multi-view spectral data.

[0059] In this embodiment, the front and back of each leaf are scanned, where the front is the adaxial surface and the back is the abaxial surface, to capture multi-view spectral information.

[0060] In this embodiment, a GaiaField-N17E hyperspectral imaging system is used, as Figure 2 shown. The spectral range of the hyperspectral imaging system covers 900 nm to 1700 nm, the spectral resolution is 5 nm, and the spatial resolution is 640 pixels in 512 bands. The hyperspectral imaging system consists of an indoor test chamber, mainly including a hyperspectral camera, a support platform, a lifting platform, a base, and four halogen lamps. The lifting platform is arranged above the base, the support rods of the support frame are arranged around the base, the lifting platform is directly below the top of the support frame, the hyperspectral camera and the four halogen lamps are arranged on the top of the support frame. The lifting platform is used to control the height of the Dendrobium officinale samples and counterfeit samples, the support frame is used to support the hyperspectral camera and the four halogen lamps, the hyperspectral camera is used to collect multi-view spectral data, and by equipping four 50W halogen lamps, it is used to provide stable illumination. When scanning multi-view spectral data, the hyperspectral camera uses an array detector inside the hyperspectral camera and perpendicular to the moving direction, enabling it to scan a two-dimensional space when the scanning platform of the hyperspectral imaging system advances. A conveyor belt can also be set on the lifting platform, and the speed of the conveyor belt is set to 0.8 cm / s to ensure that the sample passes through the scanning area at a stable and uniform speed. The exposure time of the hyperspectral camera is adjusted automatically, and the gain factor is set to 1. The vertical distance between the sample and the lens is fixed at 42 cm. During scanning, the sample is placed on a black bottom plate to enhance image contrast and recognition, and at the same time, minimize the interference of background diffuse reflection on the sample spectral data. During the data collection process, the hyperspectral camera scans each sample twice to improve the reliability and repeatability of the data.

[0061] Step 3: Data preprocessing.

[0062] After the scanning of this embodiment is completed, black and white calibration is performed on the original hyperspectral image. The data captured by the hyperspectral imaging system mainly represents signal intensity. To accurately obtain spectral reflectance, precise calibration and black and white calibration must be performed. The key to the calibration process lies in matching the image data with the actual spectral response to ensure that each pixel in the image accurately reflects the spectral characteristics of the target substance. This usually requires detailed measurement of the spectral response of the device and correction of systematic errors. The purpose of black and white calibration is to eliminate background interference introduced by system noise and uneven illumination. By analyzing the images obtained under lightless conditions (black field) after covering the lens and using a standard reflectance plate (white field), the image data can be calibrated to ensure its accuracy and consistency. Calibrated image It is calculated by the following formula.

[0063]

[0064] Wherein, represents the calibrated image, represents the original hyperspectral image, represents the whiteboard image for calibration, represents the blackboard calibration image.

[0065] In the process of spectral data extraction in this embodiment, first, the black background in the calibrated hyperspectral image is removed to accurately extract the region of interest for each sample. This step is achieved by selecting a grayscale image at a specific wavelength and using the Otsu method to determine an optimal threshold. Among them, the region exceeding the threshold is marked as "1", while the background region below the threshold is marked as "0". By this method, a binary mask is generated for each sample and applied to the entire hyperspectral image cube, effectively removing the black background. To improve the accuracy and reproducibility of prediction, the reflectance values of the rice / polished rice pixels in each hyperspectral image are averaged to extract the average spectrum of each sample. Then, the average spectra extracted from two hyperspectral images scanned for the same sample are averaged again, and the final result obtained is used as the spectral data corresponding to the sample, providing a basis for subsequent analysis. Subsequently, the software ENVI (remote sensing image processing software) is used to extract the spectrum of the region of interest for each sample. In the software ENVI, each sample is equally divided into four parts with a rectangular frame, and one-fourth of each sample is randomly selected as the region of interest. Each pixel in the region of interest contains a set of different spectral information. By averaging the spectral reflectance of all pixels in this region, the final spectral value of the sample can be obtained.

[0066] In this embodiment, various filtering techniques such as median filtering algorithm, average filtering algorithm, Gaussian filtering algorithm, Savitzky-Golay filtering algorithm, and principal component analysis method can be used to preprocess hyperspectral data to reduce noise and improve data quality, ensuring the accuracy and robustness of subsequent analysis.

[0067] The median filtering algorithm is a non-linear filtering technique that replaces each pixel value with the median of its neighborhood. This method is particularly effective in removing salt-and-pepper noise. The filtering process involves specifying a window size to determine the neighborhood of each pixel, then calculating the median within that neighborhood and replacing the original pixel value. Median filtering preserves edge details and is especially suitable for removing impulse noise without overly blurring the image. The core parameter of median filtering is kernel_size (kernel size), which determines the neighborhood range used during filtering. kernel_size incorporates the range of hyperparameter search (taking values 3, 5, 7, 9) and is automatically tuned through Grid Search CV (Grid Search). Therefore, during the training process, Grid Search CV will iterate through different kernel_sizes, combined with other hyperparameters of the classifier, and select the best parameter combination through cross-validation.

[0068] The working principle of the average filtering algorithm is to replace each pixel value with the average of its adjacent pixels. It can effectively reduce random noise but may blur edges, making it less effective in preserving fine details. This method is simple and efficient to calculate and is a popular choice for initial noise reduction of hyperspectral data. In the model, average filtering is encapsulated as a custom transformer and integrated into the Pipeline, used in combination with the classifier. Pipeline is a tool provided by scikit-learn (machine learning library) for integrating multiple data preprocessing and model training steps into a unified process, ensuring that data passes through each step in a preset order and can be combined with Grid Search CV (Grid Search Cross-Validation) to achieve automatic hyperparameter tuning. The adjustable parameter specifically includes window_size (window size), which is used to control the neighborhood range when calculating the mean for each data point. The range of automatic hyperparameter tuning is set to {2, 3, 4, 5} for window_size in Grid Search CV, that is, the grid search will test these different window sizes and select the optimal value. Throughout the process, Grid Search CV cross-validates different window_size combinations and combines with other hyperparameters of the classifier to find the optimal parameter configuration. The finally trained best model will adopt the optimal window_size to ensure that the filtered spectral data can better adapt to the classification task and improve the classification accuracy.

[0069] The Gaussian filtering algorithm is a common smoothing technique that uses a Gaussian function to weight adjacent pixel values, giving higher weights to pixels closer to the center. Compared with the average filtering algorithm, the Gaussian filtering algorithm can smooth the image while retaining more details in the central region. The Gaussian filtering algorithm is widely used in hyperspectral data preprocessing because it can effectively reduce noise while minimizing the blurring effect. In this embodiment, a custom Gaussian filtering transformer is defined to smooth the spectral data using the gaussian_filter function in the scipy.ndimage module. The sigma parameter of this Gaussian filtering transformer controls the smoothness of the filtering. A smaller sigma value retains more details, while a larger sigma value produces a stronger smoothing effect. To optimize the impact of Gaussian filtering parameters on the classification task, this embodiment integrates the Gaussian filtering transformer as a data preprocessing step in the Pipeline and uses Grid Search CV for hyperparameter search. In this embodiment, multiple candidate values (0.5, 1, 2, 3, 4) are selected for the sigma parameter, and grid search is performed in combination with different hyperparameters of the classifier to find the best parameter combination. Finally, this embodiment evaluates the effect of the Gaussian filtering algorithm through the classification results on the test set, and uses accuracy, precision, recall, and F1-score as measurement metrics.

[0070] The Savitzky-Golay filtering algorithm is a smoothing method commonly used in signal processing. It applies polynomial fitting within a sliding window to smooth the data, thereby reducing high-frequency noise. In hyperspectral data processing, the Savitzky-Golay filtering algorithm helps to smooth small random fluctuations while preserving the main signal features. The key parameters of the Savitzky-Golay filtering algorithm include the window length (window_length) and the polynomial order (polyorder). Among them, the window length is used to define the size of the sliding window and must be an odd number, which determines the number of spectral points considered during the filtering process. A smaller window length can better preserve details but has weaker noise resistance, while a larger window length can improve the smoothing effect but may cause detail loss. In this embodiment, three window sizes of 5, 7, and 9 are selected for automatic parameter tuning. The polynomial order refers to the order of the polynomial used to fit the data within the sliding window. When the order is low, the filtering effect is smoother, but some spectral information may be lost; when the order is high, the local change characteristics of the spectral curve can be more accurately retained. In this embodiment, three orders of 2, 3, and 4 are selected for experimental comparison. Automatic parameter tuning is performed through Grid Search CV, combined with a classification model, to evaluate the classification performance under different parameter combinations to determine the optimal SG filtering parameters. The finally selected parameters can maintain the key features of the spectral data while reducing noise, thereby improving the accuracy of the classification model.

[0071] The PCA algorithm is an unsupervised statistical method for data dimensionality reduction. It extracts the most important features, called principal components, by maximizing the variance of the data. In spectral analysis, the PCA algorithm can help identify differences between samples and potential chemical changes. In this embodiment, the PCA algorithm is used for dimensionality reduction to reduce the dimensionality of spectral data, improve the computational efficiency of the model, and reduce data redundancy and noise interference. The parameters of the PCA transformer include n_components (the number of principal components), pca.fit(X), and pca.transform(X). Among them, n_components defaults to 10 and is used to represent the dimensionality of the data after dimensionality reduction. This parameter determines the number of principal components to be retained and usually selects a suitable value according to the cumulative variance contribution rate (such as more than 95%). pca.fit(X) is used to fit the training data and calculate the principal component directions. pca.transform(X) is used to project the original data into the principal component space to achieve dimensionality reduction. During the Grid Search CV process, the PCA algorithm first reduces the dimensionality of the spectral data, then combines the Standard Scaler to process the data, and combines other hyperparameters of the classifier to select the best parameter combination through cross-validation. This can not only reduce the complexity of the data but also improve the generalization ability of the model.

[0072] Step 4: Authenticity identification.

[0073] In this embodiment, an SVM model is used. Based on hyperspectral data, the SVM model is used to distinguish between 9 varieties of Anoectochilus roxburghii and counterfeits. The SVM model can determine the optimal hyperplane that maximizes the margin between different classes and has strong generalization ability. The radial basis function (rbf) kernel is selected because it is very effective in dealing with non-linear data patterns. Before training, the spectral data is standardized using Standard Scaler to ensure the consistency of the feature scale. The hyperparameters, including the penalty parameter, kernel type, gamma, and polynomial degree, are optimized using the methods of grid search and five-fold cross-validation. The model performance is evaluated using accuracy, precision, recall, F1-score, and confusion matrix, so as to comprehensively evaluate the classification accuracy and error. The specific steps are as follows.

[0074] First, divide the dataset. The data source is the experimental data stored in an Excel file, where the first row is the spectral wavelength feature; the 2nd to 81st rows are the spectral data, with a total of 80 samples, and each sample contains 440-dimensional spectral features. For label construction, real samples (category 0): 40 samples; counterfeits (category 1): 40 samples. For standardization, since the numerical range of the spectral data is large, this embodiment uses Standard Scaler for standardization (mean is 0, standard deviation is 1) to ensure that all features are on the same scale, which helps the stable training of the SVM. For the division of the training set / test set, 80% training set (64 samples), 20% test set (16 samples). stratify=labels ensures that the proportions of the two types of samples in the training set and the test set are consistent, improving the generalization ability of the model.

[0075] For the structure of the SVM classifier, the SVM maximizes the margin between data classes by finding a hyperplane. The core parameters include the kernel function (kernel), linear (linear kernel), rbf, regularization parameter (C), gamma parameter, and the degree of the polynomial kernel, etc. Among them, linear is applicable to linearly separable data and is computationally simple. Rbf is applicable to non-linear data and can capture complex patterns. The regularization parameter, also known as the penalty parameter, controls the model's tolerance for misclassification. The larger the regularization parameter, the lower the model's tolerance for misclassification, and the more stringent the decision boundary. In this embodiment, combinations are made with other parameters from four regularization parameter values of [0.1, 1, 10, 100]. The gamma parameter is only applicable to the radial basis function kernel and is used to control the influence range of a single training sample on the classification boundary. The larger the gamma parameter, the smaller the influence range, and the more complex the decision boundary. In this embodiment, the coefficient of the kernel function is automatically selected ('gamma': ['scale', 'auto']). The degree of the polynomial kernel only takes effect when kernel = 'poly' and controls the order of the polynomial kernel, affecting the complexity of the decision boundary. In this embodiment, it is selected from three polynomial kernel numbers of [3, 4, 5].

[0076] During the training process, in this embodiment, the SVM classifier is defined in the form of SVM = SVC(). For hyperparameter optimization, different hyperparameter combinations are tested under five-fold cross-validation.

[0077] All possible parameter combinations are traversed through Grid Search CV, and the best parameters are selected based on the accuracy score of cross-validation.

[0078] This embodiment sets the following evaluation metrics: Accuracy: The proportion of correctly classified cases. Precision: The proportion of actually correct cases among the positive classes predicted by the model. Recall: The proportion of actually positive classes correctly identified by the model. F1-score: The weighted harmonic mean of precision and recall. Confusion matrix: Used to visually display the correct and incorrect classification situations predicted by the model. Based on the above evaluation metrics, when this embodiment uses the test set for prediction and evaluation, the optimal SVM model is used for prediction.

[0079] For the key optimization strategies, first, standardize the data. Since the SVM is sensitive to the feature scale, StandardScaler is first used for standardization to improve the classification effect. Then, hyperparameter optimization is performed. Cross-validation is carried out through Grid Search CV to automatically select the optimal regularization parameter, gamma, and kernel function to prevent overfitting or underfitting. For kernel function selection, the linear kernel is applicable to simple classification problems, and the rbf kernel is applicable to more complex non-linear classification tasks. Five-fold cross-validation is used during cross-validation to ensure the stability of the model on different data subsets and improve the generalization ability.

[0080] In this embodiment, the SVM model combined with hyperparameter optimization based on Grid Search CV is used to effectively classify the spectral data of genuine and fake Anoectochilus roxburghii leaves. Standardized preprocessing is adopted to improve the stability of the model, and the optimal hyperparameters are selected through cross-validation. Finally, a classification accuracy of 100% is obtained.

[0081] Step Five: Strain Classification.

[0082] After the authenticity identification of Anoectochilus roxburghii leaves is completed in this embodiment, strain classification is performed on the genuine Anoectochilus roxburghii leaves among them. In this embodiment, the CNN model is compared with various models such as the SVM model, KNN (K-Nearest Neighbors) model, and LDA (Linear Discriminant Analysis) model for the strain classification performance of Anoectochilus roxburghii. At the same time, the corresponding derivative models obtained by filtering algorithms such as the MF algorithm, AF algorithm, GF algorithm, SG algorithm, and PCA algorithm for the SVM model, KNN model, and LDA model are added to the strain classification performance comparison experiment. Table 1 shows the comparison results of the strain classification performance of the CNN model and other various models.

[0083] Table 1 Comparison Results of Strain Classification Performance of Different Models

[0084] It can be directly seen from Table 1 that when the spectra of the front and back of the leaves are merged, the values of accuracy, precision, recall, and F1 score of the CNN model are all 1, and the values of all indicators reach the highest. This proves that the CNN model has great advantages in the strain classification task of hyperspectral fusion of the front and back of Anoectochilus roxburghii leaves and can accurately classify the strains of Anoectochilus roxburghii. Therefore, in this embodiment, the CNN model is used to classify nine Anoectochilus roxburghii varieties. To achieve the classification task of hyperspectral data, a 1D-CNN model is constructed in this embodiment. The 1D-CNN model includes multiple convolutional layers, pooling layers, and fully connected layers, and L1 regularization is adopted to improve the generalization ability.

[0085] In this embodiment, the hyperspectral data in the Excel file is first read, and various preprocessing methods used in Step Three are performed on it, including median filtering, mean filtering, Gaussian filtering, Savitzky-Golay smoothing filtering, and PCA preprocessing, to improve the data quality. The preprocessed dataset contains samples of 9 categories, with 40 samples in each category, for a total of 360 samples. The dataset is divided according to the ratio of 80% training set and 20% test set, and standardized processing is performed on it.

[0086] As Figure 3As shown, the input layer of the 1D-CNN model has an input data shape of (n, 1), where n is the length of the spectral data for each sample. The first layer of the 1D-CNN model is a convolutional layer, which contains 64 convolutional kernels (filters) of 7×1, with the ReLU activation function and L1 regularization adopted, and the regularization parameter λ = 0.001. The second layer is a max pooling layer (MaxPooling) with a pooling window size of 2. The third layer is a convolutional layer, which contains 128 convolutional kernels of 7×1, with the ReLU activation function and L1 regularization adopted. Then comes the Flatten layer, which unfolds the data into a one-dimensional vector. Then there is a fully connected layer (Dense), which contains 128 neurons, with the ReLU activation function and L1 regularization adopted, and the regularization parameter λ = 0.001. Finally, there is the output layer, which is also a Dense layer, containing 10 neurons (representing 9 classifications), and uses the softmax activation function.

[0087] During the optimization and training process of this embodiment, the loss function uses sparse categorical cross-entropy. The optimizer uses the Adam optimizer (learning rate = 0.0001). The loss function is sparse_categorical_crossentropy, and the metric is accuracy. The training strategy is that the batch size is 32, training for 500 epochs, and dividing 20% of the training data as the validation set. Data preprocessing uses Standard Scaler to standardize the data. For data augmentation and noise reduction, Gaussian filtering (σ = 7) and singular value decomposition (SVD, k = 4) are used for data noise reduction. 80% of the data is used for training and 20% for testing. When evaluating the model, for the accuracy of the test set, the Accuracy_Score function is used to calculate the classification accuracy of the model on the test set. During the visualization of the training process, the loss curve and the accuracy curve are plotted to observe the training convergence situation.

[0088] This embodiment effectively extracts features from spectral data through the 1D-CNN model, combines regularization and optimization algorithms, and improves the classification accuracy. By plotting the training loss and accuracy curves, this embodiment observes that the 1D-CNN model gradually converges during the training process, and the accuracy on the validation set is relatively high, indicating that the 1D-CNN model has good generalization ability. Figure 4 It is a graph showing the change of the loss value of the loss function with the number of training times during the training process. Figure 5 It is a graph showing the change of the accuracy with the number of training times during the training process. Figure 6 It is a schematic diagram of the confusion matrix of the trained 1D-CNN model on the test set. Among them,Figure 4 The loss function in Figure 4 uses sparse categorical cross - entropy loss, which is applicable to multi - class classification problems, especially when the labels are encoded as integers. The formula for the loss function is as follows:

[0089] where is the loss value, is the total number of classes, is the true label of the -th class (1 if the sample belongs to the -th class, otherwise 0), is the predicted probability of the model for the -th class.

[0090] According to Figure 4 it can be seen that as the number of training times increases, both the training set loss value (Training Loss) and the validation set loss value (Validation Loss) decrease significantly and tend to be stable, indicating that the learning process of the 1D - CNN model is effective. Figure 5 illustrates the accuracy trend during the training process. From Figure 5 it can be seen that as the training progresses, the training set accuracy (Training Accuracy) and the validation set accuracy (Validation Accuracy) gradually approach 1, and the fluctuations become smaller and smaller, further proving the stability and reliability of the 1D - CNN model. Figure 6 The confusion matrix in Figure 6 shows that the 1D - CNN model correctly classifies all samples of each class without misclassification. Among them, the true label refers to the true class of the sample, and the predicted label refers to the class of the sample predicted by the 1D - CNN model. This result proves the excellent performance of the multi - perspective spectral fusion model. The 1D - CNN model successfully utilizes the spectral data of the front and back of the leaf to improve the classification accuracy.

[0091] As Figure 6As shown, 100% accuracy can be attributed to factors such as complementarity, data augmentation, and feature diversity in model structure and optimization. By using spectral data from the front and back of the leaves, the model utilizes multi-view information to provide a richer feature representation. The spectral responses of the front and back of the leaves may differ in certain physical properties, such as light reflection and scattering characteristics. These differences help capture different structures and compositions of the leaves, ultimately improving the classification accuracy. The spectral data from the front and back are complementary, enhancing the robustness and accuracy of the model. In addition, different leaf species exhibit distinct spectral differences, especially in specific wavelength regions. The front and back spectra provide information from different angles, helping to better capture these subtle spectral differences. By incorporating this additional feature dimension (i.e., the front spectrum and the back spectrum), the model obtains more information, thereby improving its generalization ability. The introduction of feature diversity enables the model to identify more potential patterns during training, thus improving the classification accuracy.

[0092] Compared with traditional manual feature extraction methods, deep learning models such as CNN models can automatically learn the optimal feature combination through end-to-end training, thus achieving more precise classification. As the number of layers and neurons in the network increases, the model can process more complex spectral data and extract deeper features. Overall, the qualitative model used in this embodiment exhibits excellent training performance and generalization ability. These results highlight the potential of CNN-based models in hyperspectral data classification, especially when using multi-view spectral fusion. The model can effectively capture the subtle spectral differences between different varieties of Anoectochilus roxburghii, highlighting the effectiveness of deep learning techniques in solving complex classification tasks in the field of plant species identification.

[0093] Step Six: Result Output.

[0094] In this embodiment, the final output result is the strain classification result of the genuine Anoectochilus roxburghii leaves, corresponding to the (1)-(9) varieties of the Anoectochilus roxburghii samples in Step One, so as to determine that the Anoectochilus roxburghii leaves are small round leaves (1), pointed leaves (2), Hongxia (3), J6 male (4), Caixia (5), large round leaves (6), Hongxia large leaves (7), Jinmai No. 1 (8), or Taihong (9).

[0095] In this embodiment, hyperspectral imaging technology and machine learning technology are combined and applied to the scenario of high-precision classification and authenticity identification of Anoectochilus roxburghii and its counterfeits. Hyperspectral data are collected from the front and back leaf surfaces of 9 varieties of Anoectochilus roxburghii and 2 counterfeit varieties (Ludisia discolor and Goodyera schlechtendaliana), and then an SVM model is used for authenticity identification. The experimental results show that the SVM model achieves a 100% classification accuracy in distinguishing Anoectochilus roxburghii and its counterfeit varieties, effectively capturing the spectral differences between the front and back leaf surfaces. In the process of authenticity identification, by using the SVM model and the spectral data of the front and back leaf blades, a 100% accuracy is achieved in distinguishing Anoectochilus roxburghii and counterfeit varieties. In the classification of different varieties of Anoectochilus roxburghii, by introducing a multi-view spectral data fusion CNN model and combining the spectral data of the front and back leaf blades, the classification performance and robustness are significantly improved, and the classification task is completed by using complementary information, highlighting the potential of hyperspectral imaging technology and machine learning technology in the authenticity identification and strain classification of plant leaves, and providing a new perspective for plant authenticity identification and strain classification.

[0096] In an exemplary embodiment, a system for authenticity identification and strain classification of Anoectochilus roxburghii plant leaves is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the method for authenticity identification and strain classification of Anoectochilus roxburghii plant leaves described above.

[0097] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0098] In this embodiment, specific examples are used to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, based on the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for authenticating the leaves of Anoectochilus roxburghii plants and classifying their strains, characterized in that, Including: Obtaining multi-view spectral data of a plant sample to be tested; The multi-view spectral data includes the front spectral data and the back spectral data of each leaf of the plant sample to be tested; Preprocessing the multi-view spectral data to obtain preprocessed multi-view spectral data; Inputting the preprocessed multi-view spectral data into a pre-trained SVM model to output a true / false discrimination result; The pre-trained SVM model refers to a model obtained by jointly tuning the parameters of the SVM model and the preprocessing model based on the preprocessed multi-view spectral data of the Dendrobium officinale samples and the counterfeit samples, and training after finding the optimal parameters. The preprocessing model refers to the model corresponding to the filtering algorithm; When the true / false discrimination result is that the plant sample to be tested is a Dendrobium officinale sample, inputting the preprocessed multi-view spectral data of the plant sample to be tested into a pre-trained CNN model to output a strain classification result; The pre-trained CNN model refers to a model obtained by jointly tuning the parameters of the CNN model and the preprocessing model based on the preprocessed multi-view spectral data of the Dendrobium officinale samples, and training after finding the optimal parameters.

2. The method for authenticating the true or false and classifying the strains of the leaves of Anoectochilus roxburghii plants according to claim 1, characterized in that, Obtaining multi-view spectral data of a plant sample to be tested specifically includes: Performing a cleaning process on the plant sample to be tested to obtain a cleaned plant sample to be tested; Performing an air-drying process on the cleaned plant sample to be tested to obtain an air-dried plant sample to be tested; Collecting the front spectral data and the back spectral data of each leaf of the air-dried plant sample to be tested to obtain multi-view spectral data.

3. The method for authenticating the authenticity and classifying the strains of the leaves of Anoectochilus roxburghii plants according to claim 2, characterized in that, Collecting the front spectral data and the back spectral data of each leaf of the air-dried plant sample to be tested to obtain multi-view spectral data specifically includes: Using a hyperspectral imaging system of the GaiaField-N17E model to scan the front and back of each leaf of the air-dried plant sample to be tested respectively, so as to collect the front spectral data and the back spectral data of each leaf to obtain multi-view spectral data.

4. The method for authenticating the true and false of the leaves of Anoectochilus roxburghii plants and classifying the strains according to claim 1, characterized in that, Preprocessing the multi-view spectral data to obtain preprocessed multi-view spectral data specifically includes: Performing black-and-white correction on the multi-view spectral data to obtain black-and-white corrected multi-view spectral data; Performing region of interest extraction on the black-and-white corrected multi-view spectral data to obtain preprocessed multi-view spectral data.

5. The method for authenticating the true and false of the leaves of Anoectochilus roxburghii plants and classifying the strains according to claim 1, characterized in that, Before training the pre-trained SVM model, standardizing the preprocessed multi-view spectral data using Standard Scaler; when training the pre-trained SVM model, optimizing the hyperparameters including the penalty parameter, kernel type, gamma, and polynomial degree using the grid search and five-fold cross-validation method; and using accuracy, precision, recall, F1 score, and confusion matrix as evaluation indicators for model performance evaluation.

6. The method for authenticating the authenticity and classifying the strains of the leaves of Anoectochilus roxburghii plants according to claim 1, characterized in that, The filtering algorithm is at least one of the median filtering algorithm, average filtering algorithm, Gaussian filtering algorithm, Savitzky-Golay filtering algorithm, and principal component analysis method.

7. The method for authenticating the true and false of the leaves of Anoectochilus plants and classifying the strains according to claim 1, characterized in that, The pre-trained CNN model is a 1D-CNN model.

8. A system for authenticating the leaves of Anoectochilus plants and classifying their strains, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the method for authenticating the authenticity and classifying the strains of the leaves of Anoectochilus roxburghii plants according to any one of claims 1-7.

Citation Information

Patent Citations

  • NIR (Near Infrared Spectrum) undamaged identification authenticity method for wild ginseng

    CN102636452A

  • Identification method of anoectohilus formosanus and adulterants thereof

    CN108152245A

  • Method for quickly identifying authenticity and quality of semen armeniacae amarae based on hyperspectral imaging technology

    CN113008817A

  • Plant leaf nitrogen content hyperspectral modeling method and device based on textural features

    CN116087108A

  • Mass spectrum data classification method based on support vector machine

    CN117407779A

Cited By

  • Pinellia ternate and processed product authenticity identification method and system based on hyperspectral imaging and transfer learning network

    CN122336547A