Origin Traceability Method of Honeysuckle Based on Near-Infrared Spectral Features and 1D-VD-CNN
Through near-infrared spectral characteristics and improved 1D-VD-CNN model, the problem of difficulty in identifying honeysuckle origin is solved, and fast and high-precision origin traceability is achieved, and detection efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202210637425.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-06-08
AI Technical Summary
The lack of unified chemical testing standards in the existing technology has led to difficulties in identifying the origin of honeysuckle. There is a fraud in the market, and a fast, accurate and low-cost detection method is needed.
Using near-infrared spectral characteristics and improved 1D-VD-CNN model, by acquiring honeysuckle near-infrared spectral data, building feature layers and classification layers, using constraint optimization design and multiple convolutional layers for training, to achieve high-precision traceability of origin.
It realizes rapid and high-precision identification of honeysuckle production areas, improves detection efficiency and accuracy, and reduces detection costs.
Smart Images

Figure CN115184300B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data mining, and particularly to a method for tracing the origin of honeysuckle based on near-infrared spectral features and 1D-VD-CNN. Background Art
[0002] At present, there is no unified chemical detection standard for the origin identification and analysis of honeysuckle. With the rising price of honeysuckle in recent years, Pingyi honeysuckle, as a geographical indication product and a medicinal material from the authentic producing area, often encounters counterfeiting in the market. Therefore, establishing an accurate, reliable, simple, easy-to-implement, low-cost and rapid detection method for honeysuckle has become a new research direction, which is of great significance for promoting the development of honeysuckle identification technology and safeguarding the interests of honeysuckle growers and consumers. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a method for tracing the origin of honeysuckle based on near-infrared spectral features and 1D-VD-CNN, which can quickly and accurately identify the origin of honeysuckle.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions:
[0005] A method for tracing the origin of honeysuckle based on near-infrared spectral features and 1D-VD-CNN, comprising the following steps:
[0006] Step S1: Obtain the near-infrared spectral data of honeysuckle and preprocess it;
[0007] Step S2: Construct an improved 1D-VD-CNN model and train it based on the preprocessed near-infrared spectral data of honeysuckle;
[0008] Step S3: Trace the origin of the honeysuckle to be tested according to the trained 1D-VD-CNN model.
[0009] The specific preprocessing method in step S1 is as follows:
[0010] In the matlab software, use the KS method to divide the sample data set into a training data set and a test data set according to a preset ratio;
[0011] Sort all samples according to the spectral absorption rate value from low to high, divide them into N groups on average, calculate the actual number of good samples and bad samples in each of the N groups for each origin, accumulate the number of good samples and bad samples in each origin, and calculate the proportion and difference of the accumulated number of good samples and bad samples in each origin;
[0012] Among them, the actual number of good and bad samples in a certain production area is the number of good and bad samples in this data, the cumulative number of good and bad samples in the production area is the cumulative number of good and bad samples in this group, the proportion of the cumulative number of good and bad samples is the ratio of the cumulative number of good and bad samples to the total number of good and bad samples, and the difference is the proportion of the cumulative number of bad samples minus the proportion of the cumulative number of good samples; the KS index is the ratio of the absolute value of the difference.
[0013] Randomly select a preset proportion of the sample data as the test set of the model, and the remaining sample data is the training set of the model. Among them, a preset proportion of the validation set is randomly split from the training set as the validation basis for model selection, and the remaining part in the training set is used as the training data of the model.
[0014] Furthermore, the construction of the improved 1D-VD-CNN model is specifically as follows:
[0015] (1) Divide the hidden layer of the traditional 1D-CNN into two parts - the feature layer and the classification layer;
[0016] (2) Convert the design of the feature layer into a constrained optimization design;
[0017] (3) For the design of the classification layer, remove the fully connected layer and replace it with multiple convolutional layers.
[0018] Furthermore, the constrained optimization design includes learning ability design and learning necessity design, specifically as follows:
[0019] (1) Learning ability design
[0020] Introduce a value C:
[0021]
[0022] Among them, nconc is the actual size of the convolutional kernel, nfield is the receptive field size, that is, the size of the feature map, and the C value of each convolutional layer should be greater than or equal to 1 / 6;
[0023] (2) Learning necessity design
[0024] According to the design of the first constraint condition, add several-dimensional convolutional layers and combine them with downsampling, and the receptive field size of the top layer should be equal to the size of the data vector.
[0025] Furthermore, the training of the improved 1D-VD-CNN model is specifically as follows:
[0026] 1) Set the structure of the CNN, given the input, and then initialize the corresponding weights and set the hyperparameters;
[0027] 2) During the training process, the model error is calculated through forward propagation to determine whether the set number of training iterations is reached. If not, the error is backpropagated to calculate and update the parameters of each layer; when the training iteration number is reached, the training ends, the model parameters of each layer of the CNN are solidified, and the model is saved.
[0028] Further, the testing of the improved 1D-VD-CNN model is specifically as follows: First, read the model file saved in the training phase, initialize the model structure, restore the model parameters, input the test data, and finally output the model classification result.
[0029] Further, the evaluation of the improved 1D-VD-CNN model is based on the accuracy ACC multilabel, macro-average recall REC macro , macro-average precision PRE macro and macro-average F1 score F1 macro , specifically as follows:
[0030]
[0031]
[0032]
[0033] The accuracy of the multi-label task is calculated by summing up the TP and TN of each category and then dividing by the total number of predicted samples. The calculation formula is as follows:
[0034]
[0035] Among them, TP and TN respectively refer to the number of positive classes predicted as positive classes and the number of negative classes predicted as negative classes. FP and FN respectively refer to the number of negative classes predicted as positive classes and the number of positive classes predicted as negative classes.
[0036] The present invention has the following beneficial effects compared with the prior art:
[0037] The present invention utilizes near-infrared spectroscopy data and combines the improved 1D-VD-CNN model. Compared with the traditional 1D-CNN model, it can identify honeysuckle from different origins more quickly and efficiently, and has the ability to accurately identify honeysuckle from different origins. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is the step of collecting near-infrared spectroscopy data of honeysuckle in an embodiment of the present invention
[0039] Figure 2 is the 1D-VD-CNN feature layer model in an embodiment of the present invention;
[0040] Figure 3 It is the classification layer optimization scheme in an embodiment of the present invention;
[0041] Figure 4 It is the forward and backward training process of the convolutional neural network in an embodiment of the present invention;
[0042] Figure 5 It is the network model information table of 1D-VD-CNN in an embodiment of the present invention. Specific implementation manners
[0043] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0044] Please refer to Figure 1 , the present invention provides a method for tracing the origin of honeysuckle based on near-infrared spectral features and 1D-VD-CNN, including the following steps:
[0045] Step S1: Obtain the near-infrared spectral data of honeysuckle and perform preprocessing;
[0046] Step S2: Construct an improved 1D-VD-CNN model and train it based on the preprocessed near-infrared spectral data of honeysuckle;
[0047] Step S3: Trace the origin of the honeysuckle to be tested according to the trained 1D-VD-CNN model.
[0048] In this embodiment, step S1 is specifically:
[0049] Step S11: Obtain fresh honeysuckle samples from different origins, and after drying the fresh samples in a dryer (60°C);
[0050] Step S12: Use a traditional Chinese medicine grinder to crush the samples into powders, pass through a 60-mesh sieve, and take an appropriate amount of powder and put it into a finger tube through a near-infrared spectrometer, and finally obtain the near-infrared spectral data of the samples as the sample dataset;
[0051] Step S13: Use the KS method to divide the sample set in matlab software, that is, select the calibration set by calculating the Euclidean distance between the average spectrum and each spectrum, and select the samples with larger spectral differences into the calibration set to ensure that the established model can be sufficiently representative and avoid uneven distribution of samples in the modeling set to a certain extent.
[0052] Preferably, in this embodiment, the specific implementation method of ks:
[0053] Sort all samples in ascending order according to their spectral absorption rate values, and divide them evenly into 20 groups. Calculate the actual number of good samples and bad samples in each of the 20 groups for each origin, accumulate the number of good samples and bad samples from each region, calculate the proportion and difference of the accumulated number of good samples and bad samples from each region to the total number of good and bad samples. Among them, the actual number of good and bad samples for a certain origin is the number of good and bad samples in the data, the accumulated number of good and bad samples for the origin is the accumulated number of good and bad samples in the group, the proportion of the accumulated number of good and bad samples is the ratio of the accumulated number of good and bad samples to the total number of good and bad samples, and the difference is the proportion of the accumulated number of bad samples minus the proportion of the accumulated number of good samples. The KS index is the ratio of the absolute value of the difference.
[0054] Exclude one sample, and take 500 honeysuckle samples from different origins collected in this experiment as the research object for model establishment and evaluation. Randomly select 20% of the sample data as the test set of the model, and the remaining 80% of the sample data as the training set of the model. Among them, the training set is randomly split into 20% as the validation set for model selection, and the remaining 80% in the training set is used as the training data of the model. Therefore, the sample data is divided into 3 data sets, namely the training set, the validation set and the test set. Use the training set data to establish the model, the validation set data for model selection, and the test set for finally verifying the classification performance of the model. The detailed information of the data set splitting of honeysuckle samples from different origins is shown in Table 1.
[0055] Table 1
[0056]
[0057] In this embodiment, an improved 1D-VD-CNN model is constructed, specifically:
[0058] (1) Divide the hidden layer of the traditional 1D-CNN into two parts - the feature layer and the classification layer;
[0059] (2) Convert the design of the feature layer into a constrained optimization design;
[0060] (3) For the design of the classification layer, remove the fully connected layer and replace it with multiple convolutional layers. The specific classification layer optimization scheme is as Figure 3 shown. The steps are as follows: First, reduce the size of the feature vector output by the feature layer to a smaller size (1x6, 1x7 or 1x8) through downsampling; then use two convolutional layers with a size of 1x5 and a pooling layer with Dropout to downsample the data size to a vector size with only one vector; finally, use a fully connected layer for classification to output.
[0061] Preferably, the constrained optimization design includes a learning ability design and a learning necessity design, specifically:
[0062] (1) Learning ability design (the first constraint condition)
[0063] Introduce a value C:
[0064]
[0065] where nconc is the actual convolution kernel size, nfield is the receptive field size, i.e., the feature map size, and the C value of each convolution layer should be greater than or equal to 1 / 6; for example Figure 2 For the 1D-VD-CNN feature layer model, there are a total of four major convolution layers, and downsampling with a stride of 2 is added between every two layers. The receptive field size of each layer is represented by an array. For example, the first major layer uses three 1×5, [5:4:13], and the specific meaning is that the range is from 5 to 13, and a number is taken every stride of 4, that is, the values 5, 9, and 13 are taken. That is, for a convolution kernel of size 1×k, after one downsampling with a stride of 2, its nconc is 1×2k; after another downsampling with a stride of 2, its nconc is 1×4k, that is, nconc doubles to 1×2 n k after n downsamplings with a stride of 2. The explicit constraint condition of the 1D-VD-CNN of the present invention: the C value of each convolution layer should be greater than or equal to 1 / 6. Taking 1 / 6 is an optimal lower limit of the C value.
[0066] (2) Design of the necessity of learning
[0067] According to the design of the first constraint condition, add more-dimensional convolution layers and combine them with downsampling, which will lead to the continuous growth of the receptive field, thus generating new and more complex patterns. However, when the size of the receptive field is greater than the length of the data vector, the neuron can already see the entire data at this time. Therefore, no matter how the depth is increased, no new or more complex patterns will appear in the data. At this time, increasing the network depth will not help the performance at all, but will also increase the risk of overfitting. Therefore, based on the above discussion, the present invention proposes the second specific constraint condition of 1D-VD-CNN: the receptive field size of the top layer should be equal to the size of the data vector.
[0068] In this embodiment, first, the NIRS data of honeysuckle from 4 regions of Shandong, Henan, Hebei, and Chongqing obtained by a near-infrared spectrometer are used as the research object to construct a 1D-VD-CNN model. The specific model is as shown in Figure 5. Forward and backward training and verification are carried out, and then the constructed model is experimentally verified, and finally the results are analyzed.
[0069] Such as Figure 4 are the CNN training flow chart and the test flow chart. Specific steps:
[0070] 1) Set the structure of the CNN, given the input, and then initialize the corresponding weights and set the hyperparameters;
[0071] 2) During the training process, calculate the model error through forward propagation, and determine whether the set number of training iterations is reached. If not, backpropagate the error to calculate and update the parameters of each layer; when the training iteration number is reached, end the training, solidify the model parameters of each layer of the CNN, and save the model.
[0072] 3) Enter the model testing stage. First, read the model file saved in the training stage, initialize the model structure, restore the model parameters, input the test data, and finally output the model classification result. In fact, the model testing stage is equivalent to performing a forward propagation, and the classification result obtained by the forward propagation is used as the final classification result of the model.
[0073] In this embodiment, the performance of the model is mainly evaluated by calculating the following four indicators: accuracy (Accuracy, ACC multilabel), macro-averaged recall (Macro averaging recall, RECmacro), macro-averaged precision (Macro averaging precision, PREmacro), and macro-averaged F1 score (Macro averaging F1 score, F1macro), as follows:
[0074]
[0075]
[0076]
[0077] The accuracy of the multi-label task is calculated by summing the TP and TN of each category and then dividing by the total number of predicted samples. The calculation formula is as follows:
[0078]
[0079] Among them, TP and TN refer to the number of positive classes predicted as positive classes and the number of negative classes predicted as negative classes, respectively. FP and FN refer to the number of negative classes predicted as positive classes and the number of positive classes predicted as negative classes, respectively.
[0080] The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope of the present invention.
Claims
1. A method for tracing the origin of honeysuckle based on near-infrared spectral features and 1D-VD-CNN, characterized in that, It includes the following steps: Step S1: Obtain the near-infrared spectral data of honeysuckle and preprocess it; Step S2: Build an improved 1D-VD-CNN model and train it based on the preprocessed near-infrared spectral data of honeysuckle; Step S3: Trace the origin of the to-be-tested honeysuckle according to the trained 1D-VD-CNN model; The construction of the improved 1D-VD-CNN model is specifically as follows: (1) Divide the hidden layer of the traditional 1D-CNN into two parts - the feature layer and the classification layer; (2) Convert the design of the feature layer into a constrained optimization design; (3) For the design of the classification layer, remove the fully connected layer and replace it with multiple convolutional layers; The constrained optimization design includes learning ability design and learning necessity design, specifically: (1) Learning ability design Introduce a value C: Among them, nconc is the actual convolution kernel size, nfield is the receptive field size, that is, the feature map size, and the C value of each convolutional layer should be greater than or equal to 1 / 6, that is (2) Learning necessity design According to the first constraint condition design, add convolutional layers with several dimensions, combine them with downsampling, and the receptive field size of the top layer should be equal to the size of the data vector.
2. The origin tracing method of honeysuckle based on near-infrared spectral characteristics and 1D-VD-CNN according to claim 1, wherein The preprocessing method in Step S1 is specifically: In the matlab software, use the KS method to divide the sample dataset, and divide it into a training dataset and a test dataset according to a preset ratio; Sort all samples according to the spectral absorption rate value from low to high, divide them into N groups evenly, calculate the actual number of good samples and bad samples in each of the N groups for each origin, accumulate the number of good samples and bad samples in each origin, the proportion and difference of the accumulated number of good samples and bad samples in each origin; Among them, the actual number of good and bad samples in a certain origin is the number of good and bad samples in the data, the accumulated number of good and bad samples in the origin is the accumulated number of good and bad samples in the group, the proportion of the accumulated good and bad samples is the ratio of the accumulated good and bad samples to the total number of good and bad samples, and the difference is the proportion of the accumulated bad samples minus the proportion of the accumulated good samples; the KS index is the ratio of the absolute value of the difference; Randomly select a preset proportion of sample data as the test set of the model, and the remaining sample data is the training set of the model. Among them, a preset proportion of the validation set is randomly split from the training set as the validation basis for model selection, and the remaining in the training set is used as the training data of the model.
3. The honeysuckle origin tracing method based on near-infrared spectral features and 1D-VD-CNN according to claim 1, characterized in that The training of the improved 1D-VD-CNN model is specifically: 1) Set the structure of the CNN, give the input, and then initialize the corresponding weights and set hyperparameters; 2) During the training process, calculate the model error through forward propagation, judge whether the set training iteration times are reached. If not, backpropagate the error, calculate and update the parameters of each layer; When the training iteration times are reached, end the training, solidify the model parameters of each layer of the CNN, and save the model.
4. The origin traceability method of honeysuckle based on near-infrared spectral characteristics and 1D-VD-CNN according to claim 1, wherein The testing of the improved 1D-VD-CNN model is specifically: First, read the model file saved in the training stage, initialize the model structure, restore the model parameters, input the test data, and finally output the model classification result.
Citation Information
Patent Citations
A fast face detection method based on deep cascade convolutional neural network
CN109190442A
Taiping kowkui production place discrimination method and system based on deep learning and near infrared spectroscopy
CN111896495A