Honeysuckle production place identification method and device, server and storage medium
By collecting and pre-processing the hyperspectral and visible light data of honeysuckle samples, combining feature screening and multiple classification models, the top-level input feature vector is formed, which solves the problem of insufficient identification of honeysuckle origin in the existing technology, and achieves a higher identification accuracy.
Patent Information
- Application Number
- CN202510111816.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing technology is not accurate enough when identifying honeysuckle origin, mainly because when the near-infrared spectroscopy is combined with machine vision technology, there are fewer data sources, making it difficult to fully reflect the chemical composition and texture information of honeysuckle samples.
By collecting hyperspectral data and visible light data from honeysuckle samples, pre-processing is performed to extract spectral information, texture information and color information. The arrangement importance algorithm and L1 regularization are used to perform feature screening of spectral information, combined with classification models such as support vector machines and random forests, different data sources are classified to form top-level input feature vectors, and finally, the support vector machines are used for origin identification.
Through the comprehensive utilization of multiple data sources, the various characteristics of honeysuckle samples can be considered from a more comprehensive perspective, which significantly improves the accuracy of honeysuckle origin identification.
Smart Images

Figure CN120106868A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of honeysuckle origin identification, and in particular to a honeysuckle origin identification method, device, server and storage medium. Background Art
[0002] The medicinal material honeysuckle is the dried buds or flowers of honeysuckle and the same genus of the Caprifoliaceae family. It was first recorded in the "Shennong Bencao Jing" and is one of the major traditional Chinese medicinal materials in my country. It has the effects of clearing away heat and detoxifying, and evacuating wind-heat. It is known as the "Chinese medicine penicillin". In addition to being a medicinal material, it can also be used as a raw material for herbal tea, functional food and cosmetics. Due to its strong environmental adaptability, honeysuckle is widely distributed in my country. Among them, Fengqiu, Henan, Pingyi, Shandong and Julu, Hebei are the authentic production areas of honeysuckle. However, due to the influence of factors such as the climate environment, soil and altitude in different geographical regions, there are large differences in the chemical composition of honeysuckle from different production areas, but it is difficult to distinguish the processed honeysuckle by appearance.
[0003] Near-infrared spectroscopy and machine vision technology can be used to analyze the chemical composition of honeysuckle to identify the origin of honeysuckle, and the composition information can be obtained without destroying the sample. Through the near-infrared spectral information of the honeysuckle sample, the texture information of the honeysuckle sample is analyzed by machine vision, so as to classify the origin of the honeysuckle sample. However, since the near-infrared spectrum mainly provides spectral information, combined with texture information, it can only reflect part of the information of the honeysuckle sample, resulting in fewer data sources and insufficient accuracy in identifying the origin of honeysuckle. Summary of the invention
[0004] The embodiments of the present invention provide a method, device, server and storage medium for identifying the origin of honeysuckle, so as to solve the technical problem that the identification of the origin of honeysuckle is not accurate enough.
[0005] In a first aspect, an embodiment of the present invention provides a method for identifying the origin of honeysuckle, comprising:
[0006] Collecting hyperspectral data and visible light data of the honeysuckle sample, and preprocessing the collected hyperspectral data and visible light data, wherein the preprocessed hyperspectral data includes spectral information, hyperspectral texture information and hyperspectral color information, and the preprocessed visible light data includes visible light texture information and visible light color information;
[0007] Using the permutation importance algorithm and L1 regularization to perform feature screening on the spectral information to obtain optimized spectral band information;
[0008] According to the optimized spectral band information and the spectral information, the optimized spectral information is determined, and the optimized spectral information is classified by using a support vector machine to obtain a first bottom-level classification feature vector;
[0009] According to the optimized spectral band information and the hyperspectral texture information, the optimized hyperspectral texture information is determined, the optimized hyperspectral texture information is screened by using the permutation importance algorithm to obtain the screened hyperspectral texture information, and the hyperspectral texture information is classified by using the random forest to obtain the second bottom-level classification feature vector;
[0010] According to the optimized spectral band information and the hyperspectral color information, the optimized hyperspectral color information is determined, the optimized hyperspectral color information is screened by using the permutation importance algorithm to obtain the screened hyperspectral color information, and the hyperspectral color information is classified by using the random forest to obtain the third bottom-level classification feature vector;
[0011] Using a support vector machine to classify the visible light texture information, to obtain a fourth bottom-level classification feature vector;
[0012] Using random forest to classify the visible light color information, to obtain a fifth bottom-level classification feature vector;
[0013] The first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector and the fifth bottom-level classification feature vector are used to perform feature integration to obtain a top-level input feature vector, the top-level input feature vector is classified using a support vector machine, and the classification result is used to identify the origin of honeysuckle.
[0014] Furthermore, the step of integrating the first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector, and the fifth bottom-level classification feature vector to obtain a top-level input feature vector includes:
[0015] Based on stacking generalization, the first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector, and the fifth bottom-level classification feature vector are respectively concatenated to form a plurality of top-level feature vector combinations;
[0016] The support vector machine is trained using each top-level feature vector combination, and the optimal top-level feature vector combination is selected according to the logarithmic loss function;
[0017] A top-level input feature vector is formed according to the optimal top-level feature vector combination and the first, second, third, fourth and fifth bottom-level classification feature vectors.
[0018] Furthermore, the use of the permutation importance algorithm and L1 regularization to perform feature screening on the spectral information to obtain optimized spectral band information includes:
[0019] The permutation importance algorithm is used to extract PI characteristic bands from the spectral information to obtain the filtered spectral band information;
[0020] L1 regularization is used to perform L1 feature band extraction on the screened spectral band information again to obtain the optimized spectral band information.
[0021] Furthermore, the preprocessing of the hyperspectral data and the visible light data includes:
[0022] Extracting regions of interest from the hyperspectral data and the visible light data, respectively, to obtain a hyperspectral grayscale image of interest and a visible light grayscale image of interest, respectively;
[0023] Extracting texture information and color information from the hyperspectral grayscale image of interest to obtain hyperspectral texture information and hyperspectral color information;
[0024] Texture information extraction and color information extraction are performed on the visible light grayscale image of interest to obtain visible light texture information and visible light color information.
[0025] Furthermore, the method further comprises:
[0026] Performing black and white plate correction on the hyperspectral grayscale image of interest to extract reflectance spectrum data;
[0027] Monte Carlo method is used to remove outliers from reflectance spectrum data;
[0028] After removing outlier samples, data enhancement processing is performed on the hyperspectral data to obtain spectral information.
[0029] Furthermore, the extracting of texture information and color information of the grayscale image of interest in the visible light to obtain visible light texture information and visible light color information includes:
[0030] Convert visible light data into HSV images;
[0031] Extract the images of the three color channels H, S, and V in the HSV image respectively;
[0032] The texture information of each color channel is extracted using the grayscale gradient co-occurrence matrix to obtain the visible light texture information;
[0033] According to the pixel coordinates of the region of interest, the color information of the three color channels H, S, and V are extracted respectively, and the color moment is calculated to obtain the visible light color information.
[0034] Furthermore, the collection of hyperspectral data and visible light data of the honeysuckle sample includes:
[0035] Spread the honeysuckle sample flat on a black tray, and place the tray on a conveyor belt moving at a constant speed;
[0036] Use a hyperspectral camera to collect hyperspectral data of honeysuckle samples;
[0037] Use industrial cameras to collect visible light data of honeysuckle samples.
[0038] In a second aspect, an embodiment of the present invention further provides a device for identifying the origin of honeysuckle, comprising:
[0039] Data acquisition module, used to collect hyperspectral data and visible light data of honeysuckle samples;
[0040] A data preprocessing module is used to preprocess the collected hyperspectral data and visible light data;
[0041] A feature screening module is used to screen the spectral information to obtain optimized spectral band information;
[0042] A first bottom-level classification module, used for generating a first bottom-level classification feature vector according to the optimized spectral band information and the spectral information;
[0043] A second bottom-level classification module, used for generating a second bottom-level classification feature vector according to the optimized spectral band information and the hyperspectral texture information;
[0044] A third bottom-level classification module is used to generate a third bottom-level classification feature vector according to the optimized spectral band information and the hyperspectral color information;
[0045] A fourth bottom-level classification module, used to generate a fourth bottom-level classification feature vector according to visible light texture information;
[0046] A fifth bottom-level classification module, used to generate a fifth bottom-level classification feature vector according to visible light color information;
[0047] The top-level classification module is used to generate a top-level input feature vector based on the first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector and the fifth bottom-level classification feature vector, and identify the origin of honeysuckle according to the classification result of the top-level input feature vector.
[0048] In a third aspect, an embodiment of the present invention further provides a server, including:
[0049] one or more processors;
[0050] a storage device for storing one or more programs,
[0051] When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned method for identifying the origin of honeysuckle.
[0052] In a fourth aspect, an embodiment of the present invention provides a storage medium comprising computer executable instructions, which, when executed by a computer processor, are used to execute the above-mentioned method for identifying the origin of honeysuckle.
[0053] The embodiment of the present invention collects hyperspectral data and visible light data of a honeysuckle sample, extracts spectral information, texture information, color information in the hyperspectral data, and texture information and color information in the visible light data, uses five data sources to establish classification models based on a support vector machine or a random forest as a bottom-level classifier, then performs feature splicing on the classification prediction results of the bottom-level classifier based on stacked generalization ensemble learning to form a top-level input feature vector, establishes a top-level classifier based on a support vector machine to perform classification prediction on the top-level input feature vector, and the obtained final classification result is used to identify the origin of the honeysuckle. Multiple data sources are used to classify and identify the origin of the honeysuckle samples from different angles, can consider multiple characteristics of the honeysuckle samples from a more comprehensive angle, can accurately obtain the classification results of the honeysuckle samples, and has a high accuracy rate for identifying the origin of the honeysuckle. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings:
[0055] Figure 1 This is a flow chart of a method for identifying the origin of honeysuckle according to Example 1 of the present invention;
[0056] Figure 2 This is a flow chart of a method for identifying the origin of honeysuckle according to Embodiment 2 of the present invention;
[0057] Figure 3 A result diagram of the comparison between the classification model established according to different data source combinations and the validation set described in the second embodiment of the present invention;
[0058] Figure 4 This is a flow chart of a method for identifying the origin of honeysuckle according to Embodiment 3 of the present invention;
[0059] Figure 5 It is a schematic diagram of the influence of different numbers of bands on the accuracy of the classification model by using L1 regularization screening according to the third embodiment of the present invention;
[0060] Figure 6 This is a schematic diagram of the distribution of the 31 characteristic bands selected in the third embodiment of the present invention on the original spectrum data;
[0061] Figure 7 This is a flow chart of a method for identifying the origin of honeysuckle according to Embodiment 4 of the present invention;
[0062] Figure 8 This is a schematic diagram of the detected outlier sample according to the fourth embodiment of the present invention;
[0063] Fig. 9 It is a schematic diagram of the classification prediction performance of the classification model after deleting different outlier samples according to the fourth embodiment of the present invention;
[0064] Fig.10 A schematic diagram of the classification prediction performance of the classification model after combining different preprocessing methods according to the fourth embodiment of the present invention;
[0065] Fig.11 It is a schematic diagram of the classification prediction performance of the classification model after combining different preprocessing methods with permutation importance screening according to the fourth embodiment of the present invention;
[0066] Fig.12 This is a schematic diagram of the structure of a device for identifying the origin of honeysuckle according to Embodiment 5 of the present invention;
[0067] Fig.13 This is a structural diagram of the server described in Example 6 of the present invention. DETAILED DESCRIPTION
[0068] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.
[0069] Embodiment 1
[0070] Figure 1 This is a flow chart of a method for identifying the origin of honeysuckle according to Embodiment 1 of the present invention. This embodiment is applicable to identifying the origin of honeysuckle from different origins. The method is as follows:
[0071] In my country, honeysuckle is mainly produced in three authentic areas: Fengqiu, Henan, Pingyi, Shandong, and Julu, Hebei. Due to differences in climate, soil, altitude and other factors in different geographical areas, there will be certain differences in the effective ingredients and content of honeysuckle produced. However, the honeysuckle samples obtained after processing and preparation from different producing areas are very similar in shape and appearance, making it difficult to identify the origin of honeysuckle by naked eyes and traditional experience. Hyperspectral imaging has a wider wavelength coverage range and higher band resolution. It can more finely reflect the spectral characteristics of the substance, and can combine spectral information with image information so that each pixel contains rich spectral data, thereby forming a complete spectral curve. Then use an industrial camera to collect visible light data of the honeysuckle sample, combine hyperspectral data with visible light data, classify hyperspectral data and visible light data respectively, and use ensemble learning to integrate the classification results of hyperspectral data and visible light data. It can use multiple data sources to identify the origin of honeysuckle from a more comprehensive perspective, and the identification accuracy is relatively high. Specifically include the following steps:
[0072] S101, collecting hyperspectral data and visible light data of a honeysuckle sample, and preprocessing the collected hyperspectral data and visible light data, wherein the preprocessed hyperspectral data includes spectral information, hyperspectral texture information and hyperspectral color information, and the preprocessed visible light data includes visible light texture information and visible light color information.
[0073] The hyperspectral data and visible light data of the honeysuckle samples after cleaning are collected by hyperspectral imaging equipment and industrial cameras respectively. In order to improve the accuracy of classification, the same batch of honeysuckle samples can be shuffled and tiled and then repeatedly collected to increase the data volume of the samples. At the same time, the collected hyperspectral data and visible light data can be preprocessed to reduce the influence of irrelevant data on the classification results. The preprocessing methods include region of interest extraction, denoising, data smoothing, etc. for hyperspectral data and visible light data to improve the accuracy of classification results. The spectral information, texture information and color information of the honeysuckle samples can be extracted from the hyperspectral data, and the texture information and color information of the honeysuckle samples can be extracted from the visible light data. The hyperspectral texture information can be extracted from the image in the hyperspectral data using the grayscale gradient co-occurrence matrix, and the hyperspectral color information can be extracted from the image in the hyperspectral data using the color moment. The visible light texture information can be extracted by dividing the visible light image into three color channels and using the grayscale gradient co-occurrence matrix, and the visible light color information can be extracted by the color moments of the three channels. The extracted spectral information, hyperspectral texture information, hyperspectral color information, visible light texture information, and visible light color information are used to establish a honeysuckle sample data set for identifying the origin of honeysuckle. It should be noted that before training the classification model, it is necessary to add corresponding origin labels to the honeysuckle samples in the training data. The output result of the classification model is the origin label feature, which reflects the situation of the corresponding origin of the sample obtained after the classification model classifies the input honeysuckle sample.
[0074] S102, using a permutation importance algorithm and L1 regularization to perform feature screening on the spectral information to obtain optimized spectral band information.
[0075] Since the wavelength coverage of hyperspectral data is wide and the band resolution is high, the amount of spectral information extracted from hyperspectral data is large. In order to reduce the amount of calculation during classification, it is necessary to perform feature screening on the spectral information. Under the premise of ensuring the accuracy of classification, the spectrum of the most representative partial bands can be extracted, and the classification results can be quickly obtained based on the spectrum of fewer bands, which improves the efficiency of identification of the origin of honeysuckle. The classification model of honeysuckle samples was established based on support vector machine (SVM) using the original spectral information. According to the prediction accuracy of the trained support vector machine classification model compared with the validation set, the features of the classification model were optimized using the permutation importance algorithm, and the spectral band features that have a more important impact on the prediction accuracy were screened. Then, L1 regularization was used to screen the spectral band features after screening again. The classification model of honeysuckle samples was established based on support vector machine using the spectral band features after screening. L1 regularization was introduced into the classification model to optimize the classification model. The weight coefficients of the spectral information of each band in the classification model after optimization were used to screen the bands with greater contribution, and different numbers were gradually selected from high to low according to the contribution to establish the classification model of honeysuckle samples based on support vector machine. According to the prediction performance of the classification model on the validation set, the optimal number of bands was selected to obtain the optimized spectral band information. Optimizing the spectral band information can select fewer band features with higher importance while ensuring the prediction accuracy of the classification model, which greatly reduces the calculation amount of the classification model.
[0076] S103, determining optimized spectral information according to the optimized spectral band information and the spectral information, and classifying the optimized spectral information using a support vector machine to obtain a first bottom-level classification feature vector.
[0077] The optimized spectral band features with higher importance selected after feature screening are used to form optimized spectral information. A classification model is established based on the support vector machine using the optimized spectral information to classify the origin of the honeysuckle samples. The classification results are used as the first underlying classification feature vector, which can reflect the result of identifying the origin of the honeysuckle based on the spectral information.
[0078] S104, determining the optimized hyperspectral texture information according to the optimized spectral band information and the hyperspectral texture information, screening the optimized hyperspectral texture information using a permutation importance algorithm to obtain screened hyperspectral texture information, and classifying the hyperspectral texture information using a random forest to obtain a second bottom-level classification feature vector.
[0079] According to the optimized spectral band information, the hyperspectral image of the corresponding band in the hyperspectral data is selected, and the texture information is extracted. After the texture information is extracted, the permutation importance algorithm is also used to screen the texture information. The extracted texture information is used to establish a classification model based on random forest (RF). According to the prediction accuracy of the trained classification model on the validation set, the texture information features that have a more important impact on the prediction accuracy are screened to obtain the screened hyperspectral texture information. The classification model is established based on random forest using the screened hyperspectral texture information, and the result of the classification model is used as the second bottom-level classification feature vector. The second bottom-level classification feature vector can reflect the result of identifying the origin of honeysuckle based on the texture information in the hyperspectral image.
[0080] S105, determining the optimized hyperspectral color information according to the optimized spectral band information and the hyperspectral color information, screening the optimized hyperspectral color information using a permutation importance algorithm to obtain screened hyperspectral color information, and classifying the hyperspectral color information using a random forest to obtain a third bottom-level classification feature vector.
[0081] According to the optimized spectral band information, the hyperspectral image of the corresponding band in the hyperspectral data is selected, and the color information is extracted. After the color information is extracted, the classification model is established based on the random forest using the extracted color information. According to the prediction accuracy of the trained classification model relative to the validation set, the permutation importance algorithm is used to screen the color information features that have a more important impact on the prediction accuracy, and the filtered hyperspectral color information is obtained. The random forest classification model is established by filtering the hyperspectral color information, and the result of the classification model is used as the third bottom-level classification feature vector. The third bottom-level classification feature vector can reflect the result of identifying the origin of honeysuckle based on the color information in the hyperspectral image.
[0082] S106: Classify the visible light texture information using a support vector machine to obtain a fourth bottom-level classification feature vector.
[0083] A classification model is established based on a support vector machine using the visible light texture information extracted from the visible light data. The classification result is used as the fourth bottom-level classification feature vector, which can reflect the result of identifying the origin of honeysuckle based on the texture information in the visible light image.
[0084] S107, classify the visible light color information using random forest to obtain a fifth bottom-level classification feature vector.
[0085] A classification model was established based on random forest using the visible light color information extracted from the visible light data. The classification result was used as the fifth bottom-level classification feature vector, which can reflect the result of identifying the origin of honeysuckle based on the color information in the visible light image.
[0086] S108, using the first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector and the fifth bottom-level classification feature vector to perform feature integration to obtain a top-level input feature vector, using a support vector machine to classify the top-level input feature vector, and using the classification result to identify the origin of honeysuckle.
[0087] The first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector and the fifth bottom-level classification feature vector can respectively reflect the classification of honeysuckle samples from five different data sources, and then identify the origin. The classification results of the five different data sources are used to vector splice the classification results to form a top-level input feature vector. The top-level input feature vector contains the classification results of the origin of the honeysuckle samples from multiple angles. The top-level classification model is constructed based on the support vector machine using the top-level input feature vector, and the classification results obtained by the top-level classification model are used as a reference for identifying the origin of honeysuckle.
[0088] This embodiment extracts spectral data in the hyperspectral data and visible light data respectively, and optimizes the extraction of spectral features of some bands under the premise of ensuring prediction accuracy, and uses the optimized spectral information to classify the origin of honeysuckle based on support vector machine; at the same time, the texture information and color information in the hyperspectral image corresponding to the optimized spectral band are used to classify the origin of honeysuckle based on random forest; the texture information and color information in the visible light data are used to classify the origin of the honeysuckle sample based on support vector machine and random forest respectively; the origin prediction classification results of five different data sources after classification by different models are vector spliced, and the top-level input feature vector obtained after splicing is used for classification based on support vector machine, and the classification result is used as a reference for finally identifying the origin of the honeysuckle sample. Based on the characteristics of hyperspectral and visible light, it can collect spectral information with higher band resolution and obtain image data of honeysuckle samples at different levels. The obtained multiple data sources are used to carry out origin identification and prediction respectively, and then the classification results of multiple data sources are integrated through the top-level input feature vector. The final classification result can comprehensively consider the origin of honeysuckle samples from multiple angles for identification. Compared with the classification model that only uses a single data source, it can identify and predict the origin of honeysuckle samples from a more comprehensive perspective, thereby improving the identification accuracy.
[0089] Specifically, when collecting the hyperspectral data and visible light data of the honeysuckle sample, the honeysuckle sample is spread flat in a black tray and the tray is placed on a conveyor belt moving at a constant speed; the hyperspectral data of the honeysuckle sample is collected using a hyperspectral camera; and the visible light data of the honeysuckle sample is collected using an industrial camera. Exemplarily, the hyperspectral camera model in this embodiment is Specim FX17, with a resolution of 8nm and a pixel of 3.5mm, and can collect 224 bands of hyperspectral images within the band range of 750-1700nm; the model of the industrial camera is MVCL042-90GC, with a resolution of 4096×2, a gain of 0dB, an exposure time of 40us, and an aperture size of f / 4. The distance between the hyperspectral camera and the industrial camera and the top surface of the conveyor belt is 50 cm. The cleaned honeysuckle samples are spread flat on a black tray to form a strong contrast, which is conducive to the extraction of image features. The movement speed of the conveyor belt is 4.0 m / min. The tray containing the samples is placed on the conveyor belt and passes through the collection range of the hyperspectral camera and the industrial camera in turn to complete a data collection. After that, the samples can be re-shuffled and spread flat, and data collection can be performed again to obtain the data volume of the current batch of samples and improve the origin identification prediction accuracy of the classification model.
[0090] Embodiment 2
[0091] Figure 2 This is a flow chart of a method for identifying the origin of honeysuckle according to the second embodiment of the present invention. This embodiment is optimized based on the above embodiment. In this embodiment, the first bottom classification feature vector, the second bottom classification feature vector, the third bottom classification feature vector, the fourth bottom classification feature vector and the fifth bottom classification feature vector are used for feature integration to obtain a top-level input feature vector. The specific optimization is as follows:
[0092] Based on stacking generalization, the first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector, and the fifth bottom-level classification feature vector are respectively concatenated to form a plurality of top-level feature vector combinations;
[0093] The support vector machine is trained using each top-level feature vector combination, and the optimal top-level feature vector combination is selected according to the logarithmic loss function;
[0094] A top-level input feature vector is formed according to the optimal top-level feature vector combination and the first, second, third, fourth and fifth bottom-level classification feature vectors.
[0095] Accordingly, the method for identifying the origin of honeysuckle provided in this embodiment specifically includes:
[0096] S201, collecting hyperspectral data and visible light data of a honeysuckle sample, and preprocessing the collected hyperspectral data and visible light data, wherein the preprocessed hyperspectral data includes spectral information, hyperspectral texture information and hyperspectral color information, and the preprocessed visible light data includes visible light texture information and visible light color information.
[0097] S202, using a permutation importance algorithm and L1 regularization to perform feature screening on the spectral information to obtain optimized spectral band information.
[0098] S203, determining optimized spectral information according to the optimized spectral band information and the spectral information, and classifying the optimized spectral information using a support vector machine to obtain a first bottom-level classification feature vector.
[0099] S204, determining the optimized hyperspectral texture information according to the optimized spectral band information and the hyperspectral texture information, screening the optimized hyperspectral texture information using a permutation importance algorithm to obtain screened hyperspectral texture information, and classifying the hyperspectral texture information using a random forest to obtain a second bottom-level classification feature vector.
[0100] S205, determining the optimized hyperspectral color information according to the optimized spectral band information and the hyperspectral color information, screening the optimized hyperspectral color information using a permutation importance algorithm to obtain screened hyperspectral color information, and classifying the hyperspectral color information using a random forest to obtain a third bottom-level classification feature vector.
[0101] S206: Classify the visible light texture information using a support vector machine to obtain a fourth bottom-level classification feature vector.
[0102] S207, classify the visible light color information using random forest to obtain a fifth bottom-level classification feature vector.
[0103] S208 , based on stacking generalization, feature concatenation is performed on the first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector, and the fifth bottom-level classification feature vector to form a plurality of top-level feature vector combinations.
[0104] Based on the stacking generalization (Stacking ensemble learning) thought, by synthesizing the classification results obtained by multiple bottom-level classifiers (prediction classification model), the final classification prediction accuracy can be effectively improved. The first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector and the fifth bottom-level classification feature vector are five different data sources respectively derived from hyperspectral data and visible light data, representing different angles to classify honeysuckle samples, and then identifying the origin of honeysuckle samples, the classification results of different bottom-level classifiers constructed based on different classification models using five data sources, combined by feature splicing to form the top-level input feature vector for inputting the top-level classifier, the classification result obtained by classifying the top-level input feature vector using the top-level classifier can integrate the different angles of multiple data sources, form a comprehensive classification result, can consider the more comprehensive feature information of honeysuckle samples, the classification result obtained is more accurate, can be used as a reference for honeysuckle classification and identification, and the origin of honeysuckle samples is determined according to the classification result. In the present embodiment, the honeysuckle sample origin label prediction result output by the bottom-level classifier is spliced to form a top-level input feature vector. By combining the output results of different underlying classifiers, a variety of different underlying classifier combinations are obtained. Each combination can reflect the identification angle of the origin of honeysuckle from the data source of the corresponding classifier. For example, the first underlying classification feature vector and the fifth underlying classification feature vector are combined. The classification result of the top-level input feature vector after combination can reflect the identification of the origin of honeysuckle from the two angles of spectral information and visible light color information.
[0105] S209, training the support vector machine using each top-level feature vector combination respectively, and selecting the optimal top-level feature vector combination according to the logarithmic loss function.
[0106] The bottom-level classifiers based on five different data sources can form 15 different combinations. The top-level input feature vectors obtained from each combination form are used to establish classification models based on support vector machines. The output classification results are compared to obtain the optimal top-level feature vector combination of the optimal combination. Figure 3 The figure shows the results of the classification models established by different combinations compared with the validation set. Information source 1 in the figure is the first underlying classification feature vector, which represents the classification result using the spectral information extracted from the hyperspectral data; information source 2 is the second underlying classification feature vector, which represents the classification result using the texture information extracted from the hyperspectral data; information source 3 is the third underlying classification feature vector, which represents the classification result using the color information extracted from the hyperspectral data; information source 4 is the fourth underlying classification feature vector, which represents the classification result using the texture information extracted from the visible light data; information source 5 is the fifth underlying classification feature vector, which represents the classification result using the color information extracted from the visible light data. Figure 3 It can be seen that the prediction accuracy of the 15 combinations reached 100%, and then the optimal combination was selected through the logarithmic loss function. The logarithmic loss function of the combination of information sources 1+5 was 0.029, which was the smallest among all the combinations. Therefore, the combination of this bottom-level classifier and the corresponding data source was selected as the top-level classifier combination. The corresponding information source of this method is the spectral information extracted from the hyperspectral data and the color information extracted from the visible light data. The spectral information is classified and predicted by the support vector machine, and the color information is classified and predicted by the random forest. The prediction results of the two models are vectorized and then classified and predicted by the support vector machine of the top-level classifier. Finally, an accurate classification result of the origin of honeysuckle is obtained, and the origin of honeysuckle is identified.
[0107] S210, forming a top-level input feature vector according to the optimal top-level feature vector combination and the first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector and the fifth bottom-level classification feature vector, classifying the top-level input feature vector using a support vector machine, and identifying the origin of honeysuckle using the classification result.
[0108] This embodiment uses stacked generalized ensemble learning to combine classification models for classifying and predicting honeysuckle samples based on a variety of different data sources, integrate the prediction results of each combination through feature splicing, and then use the top classifier for final prediction. The optimal combination of bottom-level classifiers is selected through the validation set, and the top-level classifier can be used to build the top-level classifier. The top-level classifier formed by the optimal combination method is used to perform final classification prediction on the classification prediction results of the integrated bottom-level classifiers. The final classification result obtained can accurately reflect the origin classification result of the honeysuckle sample, and then accurately identify the origin of the honeysuckle sample.
[0109] Embodiment 3
[0110] Figure 4 This is a flow chart of a method for identifying the origin of honeysuckle according to Embodiment 3 of the present invention. This embodiment is optimized based on the above embodiment. In this embodiment, the spectral information is feature screened using the permutation importance algorithm and L1 regularization to obtain optimized spectral band information. The specific optimization is as follows:
[0111] The permutation importance algorithm is used to extract PI characteristic bands from the spectral information to obtain the filtered spectral band information;
[0112] L1 regularization is used to perform L1 feature band extraction on the screened spectral band information again to obtain the optimized spectral band information.
[0113] Accordingly, the method for identifying the origin of honeysuckle provided in this embodiment specifically includes:
[0114] S301, collecting hyperspectral data and visible light data of a honeysuckle sample, and preprocessing the collected hyperspectral data and visible light data, wherein the preprocessed hyperspectral data includes spectral information, hyperspectral texture information and hyperspectral color information, and the preprocessed visible light data includes visible light texture information and visible light color information.
[0115] S302, using a permutation importance algorithm to extract PI characteristic bands from the spectral information to obtain filtered spectral band information.
[0116] Since the hyperspectral data collected by the hyperspectral camera has the characteristics of high resolution and wide band coverage, the amount of spectral data generated is large. In order to reduce the calculation amount of the classification model, by extracting the important bands in the spectral data, the spectral information corresponding to the important bands is used to reduce the calculation amount of the classification model under the premise of ensuring the prediction accuracy, and improve the efficiency of honeysuckle origin identification. The permutation importance algorithm is a method for evaluating the contribution of features to model performance. First, the spectral data in the hyperspectral data is used to establish a classification model based on the support vector machine. The classification prediction results of the trained classification model are compared with the validation set, and the classification accuracy of the classification model for the original data is recorded as the benchmark performance. Then, the features of the original data are shuffled based on the permutation importance algorithm, and the shuffled features are used to perform classification prediction again through the classification model. According to the classification prediction results compared with the validation set, the performance changes of the shuffled model are recorded, and then the features with high contribution to the prediction accuracy of the classification model are determined, and the spectral band features that have an important impact on the classification prediction accuracy of honeysuckle samples are obtained, that is, the spectral band information is screened.
[0117] S303, performing L1 feature band extraction on the screened spectral band information again by using L1 regularization to obtain optimized spectral band information.
[0118] In order to further reduce the amount of calculation of the classification model, L1 regularization is used to screen the spectral bands again, and the support vector machine classification model is established using the number of characteristic bands between 1 and 50. L1 regularization is introduced into the classification model to optimize the classification model, and the bands that contribute more to the prediction accuracy of the classification model are selected through the weight coefficient of each band. Figure 5As shown, under the premise of giving priority to the bands with greater contribution, when the number of bands reaches 23, the prediction accuracy of the classification model compared with the validation set is stable at 100%, and the logarithmic loss function shows a downward trend; when the number of bands reaches 31, the prediction accuracy of the classification model compared with the validation set is still 100%, and the logarithmic loss function tends to be stable, so 31 optimal bands can be selected as the optimized spectral band information. Under the premise of ensuring the prediction accuracy of the classification model, the minimum number of bands is selected, which reduces the data volume of the spectral information, reduces the calculation amount of the model, and improves the efficiency of identifying the origin of honeysuckle samples. In this embodiment, the distribution of the 31 selected characteristic bands on the original spectral data is shown in Figure 2. Figure 6 As shown, including 959-963nm, 1081-1112nm, 1203-1213nm, 1241-1298nm, 1347-1389nm, 1421-1431nm, 1485-1509nm, 1591-1605nm and 1645nm, these bands can be regarded as important bands that contribute more to the prediction accuracy of the classification model.
[0119] S304, determining optimized spectral information according to the optimized spectral band information and the spectral information, and classifying the optimized spectral information using a support vector machine to obtain a first bottom-level classification feature vector.
[0120] S305, determining the optimized hyperspectral texture information according to the optimized spectral band information and the hyperspectral texture information, screening the optimized hyperspectral texture information using a permutation importance algorithm to obtain screened hyperspectral texture information, and classifying the hyperspectral texture information using a random forest to obtain a second bottom-level classification feature vector.
[0121] S306, determining the optimized hyperspectral color information according to the optimized spectral band information and the hyperspectral color information, screening the optimized hyperspectral color information using a permutation importance algorithm to obtain screened hyperspectral color information, and classifying the hyperspectral color information using a random forest to obtain a third bottom-level classification feature vector.
[0122] S307: Classify the visible light texture information using a support vector machine to obtain a fourth bottom-level classification feature vector.
[0123] S308, classifying the visible light color information using random forest to obtain a fifth bottom-level classification feature vector.
[0124] S309, using the first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector and the fifth bottom-level classification feature vector to perform feature integration to obtain a top-level input feature vector, using a support vector machine to classify the top-level input feature vector, and using the classification result to identify the origin of honeysuckle.
[0125] This embodiment optimizes and screens the spectral data in the hyperspectral data through the permutation importance algorithm and L1 regularization, extracts the bands that contribute more to the prediction accuracy of the classification model, and uses the optimized spectral band information after screening to select the spectral information with the minimum number of bands while ensuring the accuracy of the classification model, thereby reducing the calculation amount of the classification model and improving the efficiency of the classification model in identifying the origin of honeysuckle.
[0126] Embodiment 4
[0127] Figure 7 This is a flow chart of a method for identifying the origin of honeysuckle according to the fourth embodiment of the present invention. This embodiment is optimized based on the above embodiment. In this embodiment, the hyperspectral data and visible light data are preprocessed, and the specific optimization is as follows:
[0128] Extracting regions of interest from the hyperspectral data and the visible light data, respectively, to obtain hyperspectral image regions of interest and visible light image regions of interest;
[0129] Extracting texture information and color information from the hyperspectral image region of interest to obtain hyperspectral texture information and hyperspectral color information;
[0130] Texture information extraction and color information extraction are performed on the visible light image region of interest to obtain visible light texture information and visible light color information.
[0131] Accordingly, the method for identifying the origin of honeysuckle provided in this embodiment specifically includes:
[0132] S401, collecting hyperspectral data and visible light data of a honeysuckle sample, extracting regions of interest from the hyperspectral data and the visible light data, respectively, to obtain hyperspectral image regions of interest and visible light image regions of interest, respectively.
[0133] After the hyperspectral image and visible light image of the honeysuckle sample are collected by the hyperspectral camera and the visible light camera, since the honeysuckle sample is concentrated in a certain area of the image, in order to reduce the influence of irrelevant areas on the classification model, it is necessary to extract the region of interest from the collected image and distinguish the area where the honeysuckle sample is distributed from the background in the image. By selecting the grayscale image with a large color difference between the honeysuckle sample and the background in the images of multiple bands of the hyperspectral data, the grayscale image in the hyperspectral data is used for threshold segmentation to obtain a binary image, and then the rectangular area where the sample contour is distributed is found through morphological transformation, and then the part outside the contour is shielded through mask processing. The masked image can be used to perform morphological transformation on the binary image obtained by threshold segmentation to remove the noise in the background, so as to improve the feature expression ability of the image, reduce the influence of irrelevant factors on the classification model, and finally obtain the hyperspectral image region of interest. The visible light image can use the COLOR_BGR2GRAY function in the OpenCV library to convert the RGB image into a grayscale image, and then perform the region of interest extraction process in the above process to obtain a visible light grayscale image of interest.
[0134] S402, extracting texture information and color information from the hyperspectral image region of interest to obtain hyperspectral texture information and hyperspectral color information.
[0135] In addition to spectral information, hyperspectral images also contain texture information and color information. The grayscale gradient co-occurrence matrix is used to extract the texture information in the hyperspectral image. In the grayscale gradient co-occurrence matrix, the hyperspectral texture information of the hyperspectral image is extracted through the small gradient advantage, large gradient advantage, grayscale distribution unevenness, gradient distribution unevenness, energy, grayscale average, gradient average, grayscale variance, gradient variance, correlation, grayscale entropy, gradient entropy, mixed entropy, inertia and inverse moment features. The color moment is used to extract color information from the hyperspectral image. The color distribution in the image is extracted through the first-order moment (mean), second-order moment (variance), and third-order moment (skewness) to obtain the hyperspectral color information.
[0136] S403, extracting texture information and color information from the visible light image region of interest to obtain visible light texture information and visible light color information.
[0137] Converting the collected original visible light image RGB image into HSV image is more conducive to the classification model to identify the color information in the image, respectively extract the images corresponding to the three channels H, S, V (hue, saturation, brightness) in the HSV image, and use the grayscale gradient co-occurrence matrix to extract the texture information of each channel to obtain the visible light texture information. Use the color moment to extract the color information of the three channels H, S, V to obtain the visible light color information.
[0138] S404, using a permutation importance algorithm and L1 regularization to perform feature screening on the spectral information to obtain optimized spectral band information.
[0139] S405, determining optimized spectral information according to the optimized spectral band information and the spectral information, and classifying the optimized spectral information using a support vector machine to obtain a first bottom-level classification feature vector.
[0140] S406, determining the optimized hyperspectral texture information according to the optimized spectral band information and the hyperspectral texture information, screening the optimized hyperspectral texture information using a permutation importance algorithm to obtain screened hyperspectral texture information, and classifying the hyperspectral texture information using a random forest to obtain a second bottom-level classification feature vector.
[0141] S407, determining the optimized hyperspectral color information according to the optimized spectral band information and the hyperspectral color information, screening the optimized hyperspectral color information using a permutation importance algorithm to obtain screened hyperspectral color information, and classifying the hyperspectral color information using a random forest to obtain a third bottom-level classification feature vector.
[0142] S408: Classify the visible light texture information using a support vector machine to obtain a fourth bottom-level classification feature vector.
[0143] S409, classifying the visible light color information using random forest to obtain a fifth bottom-level classification feature vector.
[0144] S410, using the first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector and the fifth bottom-level classification feature vector to perform feature integration to obtain a top-level input feature vector, using a support vector machine to classify the top-level input feature vector, and using the classification result to identify the origin of honeysuckle.
[0145] This embodiment extracts the region of interest from the hyperspectral data and visible light data, shields the irrelevant regions outside the honeysuckle samples, reduces the influence of irrelevant factors on the classification model, makes the model more focused on the area where the honeysuckle samples are distributed, improves the prediction performance of the classification model, and reduces the amount of calculation of the classification model. By extracting spectral information, hyperspectral texture information, and hyperspectral visible light information from the hyperspectral data, and extracting visible light texture information and visible light color information from the visible light data, data sources for the classification of honeysuckle origin are provided from five different angles, so that when the classification model is finally used to classify and identify the origin of honeysuckle, multiple data sources can be used to consider more comprehensive information of honeysuckle samples, and the classification results obtained from a combination of multiple angles can improve the accuracy of the classification model.
[0146] An optional implementation of this embodiment is that when using hyperspectral data to extract spectral information, the hyperspectral grayscale image of interest is corrected with a black and white plate to extract reflectance spectral data. In the extracted hyperspectral grayscale image of interest, the black and white plate correction is used to identify the coordinates of black pixels, traverse the hyperspectral image of each band and the hyperspectral image corrected with the black and white plate according to the coordinates, and correct the light intensity curve to obtain the reflectance spectral data of the honeysuckle sample. The formula is as follows:
[0147]
[0148] Among them, R cal is the reflectivity of the honeysuckle sample, R raw is the value in the hyperspectral image of the honeysuckle sample, R dark is the system error of the hyperspectral camera, R white is the value in the hyperspectral image of the whiteboard.
[0149] The Monte Carlo method is used to remove outlier samples from the reflectance spectral data. The Monte Carlo cross-validation method can monitor outlier samples in the reflectance spectral data. 70% of the samples in the reflectance spectral data are randomly extracted and repeated multiple times. The samples randomly extracted each time are used to establish a classification model based on a support vector machine. The performance of the classification model is evaluated by comparing it with the remaining 30% of the samples. The obtained cumulative probability is used to identify samples with obvious offsets, and the accuracy of the support vector machine classification model is verified again after the offset samples are removed to verify the impact of the offset samples on the accuracy of the classification model. The offset samples with less impact on the accuracy can be retained, and the offset samples with greater impact are considered as outlier samples and removed. For example, Figure 8 As shown in the figure, three outlier samples, 53, 68 and 107, were detected in a batch of honeysuckle samples. These three samples were deleted respectively. According to the logarithmic loss function of the support vector machine classification results after deletion, it can be seen that after removing the three samples 53, 68 and 107, the prediction performance of the model is equivalent to that of the model after only removing the two samples 53 and 107, as shown in the figure. Fig. 9 As shown, this indicates that sample No. 68 has a relatively small impact on the model performance and there is no need to remove sample No. 68.
[0150] After the outliers are removed, the hyperspectral data is enhanced to obtain spectral information. The hyperspectral data after the outliers are removed are enhanced through first-order derivative (1st Der), second-order derivative (2nd Der), multivariate scatter correction (MSC), standard normal variable transformation (SNV), SG smoothing, mean centering (MC), maximum-minimum normalization (MMS) and standardization (AS) to improve the quality of the spectral data and thus improve the performance of the classification model. The first-order derivative and the second-order derivative are used to enhance the subtle changes in the spectral data, the multi-source scatter correction and the standard normal variable transformation are used to correct the scattering effect in the spectral data, the SG smoothing and mean centering are used to smooth the data, and the maximum-minimum normalization and standardization are used to scale the data. Through the above-mentioned data enhancement processing methods and the combination of the methods, the classification model is established based on the support vector machine using the processed spectral data, and the performance of different processing methods and combinations is evaluated. Fig.10 As shown, the data processed by the three combinations of first-order derivative and standardization (1st Der-AS), standard normal variable transformation and standardization (SNV-AS), and SG smoothing and first-order derivative and standardization (SG-1st Der-AS) have lower logarithmic loss function values compared with the validation set after being predicted by the classification model. Therefore, one of these three combinations can be selected to perform data enhancement processing on the hyperspectral data after removing outlier samples to obtain spectral information. In this embodiment, the spectral bands can be subsequently feature screened by the permutation importance algorithm. After the permutation importance algorithm is used to screen the classification model established based on the support vector machine, the logarithmic loss function value of the combination of SG smoothing and first-order derivative and standardization combined with the permutation importance algorithm (SG-1st Der-AS-PI) is the smallest, as shown in FIG. Fig.11 As shown, it is 0.029, so this combination is used as a preprocessing and screening method for spectral information in hyperspectral data.
[0151] This embodiment removes outlier samples from the extracted reflectance spectral data to reduce interference to the classification model, improves the prediction performance of the classification model, and then performs data enhancement processing. By selecting a better combination of data processing methods compared with the validation set, the spectral data is processed using the combined data processing method, which can improve the quality of the spectral data and help the classification model to obtain more accurate classification prediction results.
[0152] Optionally, extracting texture information and color information from the visible light image region of interest to obtain visible light texture information and visible light color information includes:
[0153] The grayscale image of the visible light data is used to determine the region of interest of the visible light data. The COLOR_BGR2GRAY function in the OpenCV library is used to convert the RGB image into a grayscale image, and the obtained grayscale image of the visible light data is subjected to threshold segmentation to obtain a binary image. After that, the rectangular area where the sample contour is distributed is found through morphological transformation, and then the part outside the contour is shielded through mask processing to determine the region of interest of the visible light data.
[0154] Convert visible light data into HSV images. Extract the images of the three color channels H, S, and V in the HSV image respectively. Use the COLOR_BGR2HSV function in the OpenCV library to convert the RGB image of the original visible light image into an HSV image. The HSV image describes the characteristics of the visible light image from different angles through the three channels of hue, saturation, and brightness, which is more conducive to the classification model to recognize the color information in the image, instead of using the different color brightness of the three RGB channels for expression.
[0155] The regions of interest of each color channel are extracted using the regions of interest of the visible light data. The regions of interest of each color channel are extracted using the regions of interest of the visible light data, and other regions are shielded by a mask, so as to finally form the regions of interest of each color channel.
[0156] The grayscale gradient co-occurrence matrix is used to extract texture information from the image region of interest in each color channel to obtain visible light texture information. The grayscale gradient co-occurrence matrix combines the grayscale co-occurrence matrix (GLCM) and gradient information, and can describe the texture information of the image more comprehensively from more angles. It uses multiple eigenvalues to describe the texture features of the image from different angles to form visible light texture information.
[0157] The color information is extracted from the image region of interest of each color channel, and the color moment is used to calculate the visible light color information. The color information of the three channels is extracted from the images corresponding to the hue, saturation and brightness channels using the color moment. The color moment reflects the average intensity, color distribution dispersion and color distribution asymmetry of each color channel through the first-order moment (mean), second-order moment (variance) and third-order moment (skewness). The color moment features represent the honeysuckle samples identified in the image, forming the visible light color information of the honeysuckle samples.
[0158] Embodiment 5
[0159] Fig.12 The figure is a schematic diagram of the structure of a device for identifying the origin of honeysuckle according to the fifth embodiment of the present invention. In this embodiment, the device for identifying the origin of honeysuckle comprises:
[0160] The data acquisition module 810 is used to collect the hyperspectral data and visible light data of the honeysuckle sample;
[0161] A data preprocessing module 820 is used to preprocess the collected hyperspectral data and visible light data;
[0162] The feature screening module 830 is used to screen the spectral information to obtain optimized spectral band information;
[0163] A first bottom-level classification module 840, configured to generate a first bottom-level classification feature vector according to the optimized spectral band information and the spectral information;
[0164] A second bottom-level classification module 850 is used to generate a second bottom-level classification feature vector according to the optimized spectral band information and the hyperspectral texture information;
[0165] A third bottom-level classification module 860, for generating a third bottom-level classification feature vector according to the optimized spectral band information and the hyperspectral color information;
[0166] A fourth bottom-level classification module 870, configured to generate a fourth bottom-level classification feature vector according to visible light texture information;
[0167] A fifth bottom-level classification module 880, configured to generate a fifth bottom-level classification feature vector according to visible light color information;
[0168] The top-level classification module 890 is used to generate a top-level input feature vector based on the first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector and the fifth bottom-level classification feature vector, and identify the origin of honeysuckle according to the classification result of the top-level input feature vector.
[0169] This embodiment collects hyperspectral data and visible light data of honeysuckle samples, preprocesses the collected data, selects spectral information of important bands of spectral data in hyperspectral data, uses spectral information to establish a classification model based on support vector machine, and uses the result of the model as the first bottom-level classification feature vector, then uses texture information and color information in hyperspectral data to establish classification models based on random forest, and uses the results of the model as the second bottom-level classification feature vector and the third bottom-level classification feature vector, respectively, then uses texture information and color information in visible light data to establish classification models based on support vector machine and random forest, respectively, and uses the results of the model as the fourth bottom-level classification feature vector and the fifth bottom-level classification feature vector, uses five different data sources based on different classification models as bottom-level classifiers, uses the classification prediction results of the bottom-level classifiers as the input of the top-level classifier, and uses the top-level classifier to perform final classification prediction, so that the final classification result obtained can consider the characteristics of the honeysuckle samples more comprehensively from multiple angles, and can accurately identify the origin of the honeysuckle according to the classification results.
[0170] Embodiment 6
[0171] Fig.13 This is a structural diagram of a server according to Embodiment 6 of the present invention. Fig.13 A block diagram of an exemplary server 12 suitable for use in implementing embodiments of the present invention is shown. Fig.13 The server 12 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0172] like Fig.13 As shown, the server 12 is in the form of a general purpose computing device. The components of the server 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that connects various system components (including the system memory 28 and the processing unit 16).
[0173] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor or a local bus using any of a variety of bus architectures. By way of example, these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0174] The server 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the server 12, including volatile and non-volatile media, removable and non-removable media.
[0175] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The server 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Fig.13 not shown, usually called a "hard drive"). Although Fig.13 Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present invention.
[0176] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in the memory 28, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. The program modules 42 generally perform the functions and / or methods of the embodiments described herein.
[0177] The server 12 may also communicate with one or more external devices 14 (e.g., keyboards, pointing devices, displays 24, etc.), may communicate with one or more devices that enable a user to interact with the device / server / server 12, and / or may communicate with any device that enables the server 12 to communicate with one or more other computing devices (e.g., network cards, modems, etc.). Such communication may be performed via an input / output (I / O) interface 22. In addition, the server 12 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other modules of the server 12 via a bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the server 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0178] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, such as implementing the method for identifying the origin of honeysuckle provided in the embodiment of the present invention.
[0179] Embodiment 7
[0180] Embodiment 7 of the present invention further provides a storage medium comprising computer executable instructions, which, when executed by a computer processor, are used to execute the method for identifying the origin of honeysuckle as provided in the above embodiments.
[0181] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.
[0182] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0183] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0184] Computer program code for performing the operations of the present invention may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0185] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A method for identifying the origin of honeysuckle, characterized in that: include: Collecting hyperspectral data and visible light data of the honeysuckle sample, and preprocessing the collected hyperspectral data and visible light data, wherein the preprocessed hyperspectral data includes spectral information, hyperspectral texture information and hyperspectral color information, and the preprocessed visible light data includes visible light texture information and visible light color information; Using the permutation importance algorithm and L1 regularization to perform feature screening on the spectral information to obtain optimized spectral band information; Determine the optimized spectral information according to the optimized spectral band information and the spectral information, classify the optimized spectral information using a support vector machine model, and obtain a first bottom-level classification feature vector; According to the optimized spectral band information and the hyperspectral texture information, the optimized hyperspectral texture information is determined, the optimized hyperspectral texture information is screened by using the permutation importance algorithm to obtain the screened hyperspectral texture information, and the hyperspectral texture information is classified by using the random forest model to obtain the second bottom-level classification feature vector; According to the optimized spectral band information and the hyperspectral color information, the optimized hyperspectral color information is determined, the optimized hyperspectral color information is screened by using the permutation importance algorithm to obtain the screened hyperspectral color information, and the hyperspectral color information is classified by using the random forest model to obtain the third bottom-level classification feature vector; Using a support vector machine to classify the visible light texture information, to obtain a fourth bottom-level classification feature vector; Using random forest to classify the visible light color information, to obtain a fifth bottom-level classification feature vector; The first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector and the fifth bottom-level classification feature vector are used to perform feature integration to form a top-level input feature vector, the top-level input feature vector is classified using a support vector machine, and the classification result is used to identify the origin of honeysuckle.
2. The method according to claim 1, characterized in that The step of integrating the first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector, and the fifth bottom-level classification feature vector to obtain a top-level input feature vector includes: Based on stacking generalization, the first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector, and the fifth bottom-level classification feature vector are respectively concatenated to form a plurality of top-level input feature vector combinations; The support vector machine is trained using each top-level feature vector combination, and the optimal top-level input feature vector combination is selected according to the logarithmic loss function; A top-level input feature vector is formed according to the optimal top-level input feature vector combination and the first, second, third, fourth and fifth bottom-level classification feature vectors.
3. The method according to claim 1, characterized in that The method of using the permutation importance algorithm and L1 regularization to perform feature screening on the spectral information to obtain optimized spectral band information includes: The permutation importance algorithm is used to extract PI characteristic bands from the spectral information to obtain the filtered spectral band information; L1 regularization is used to perform L1 feature band extraction on the screened spectral band information again to obtain the optimized spectral band information.
4. The method according to claim 1, characterized in that: The preprocessing of the hyperspectral data and the visible light data includes: Extracting regions of interest from the hyperspectral data and the visible light data, respectively, to obtain hyperspectral image regions of interest and visible light image regions of interest; Extracting texture information and color information from the hyperspectral image region of interest to obtain hyperspectral texture information and hyperspectral color information; Texture information extraction and color information extraction are performed on the visible light image region of interest to obtain visible light texture information and visible light color information.
5. The method according to claim 4, characterized in that The method further comprises: Performing black and white plate correction on the hyperspectral image region of interest to extract reflectance spectrum data; Monte Carlo cross validation method is used to remove outlier samples from reflectance spectrum data; After removing outlier samples, data enhancement processing is performed on the hyperspectral data to obtain spectral information.
6. The method according to claim 4, characterized in that The extracting of texture information and color information from the visible light image region of interest to obtain visible light texture information and visible light color information includes: Determine a region of interest of the visible light data using the grayscale image of the visible light data; Convert the visible light data into HSV images, and extract the images of the three color channels H, S, and V in the HSV images respectively; Extracting the image region of interest of each color channel respectively using the region of interest of the visible light data; The grayscale gradient co-occurrence matrix is used to extract texture information from the image region of interest of each color channel to obtain visible light texture information; The color information is extracted from the image region of interest of each color channel respectively, and the visible light color information is obtained by color moment calculation.
7. The method according to claim 1, characterized in that The method of collecting the hyperspectral data and visible light data of the honeysuckle sample includes: Spread the honeysuckle sample flat on a black tray, and place the tray on a conveyor belt moving at a constant speed; Use a hyperspectral camera to collect hyperspectral data of honeysuckle samples; Use industrial cameras to collect visible light data of honeysuckle samples.
8. A device for identifying the origin of honeysuckle, characterized in that: include: Data acquisition module, used to collect hyperspectral data and visible light data of honeysuckle samples; A data preprocessing module is used to preprocess the collected hyperspectral data and visible light data; A feature screening module is used to screen the spectral information to obtain optimized spectral band information; A first bottom-level classification module, used for generating a first bottom-level classification feature vector according to the optimized spectral band information and the spectral information; A second bottom-level classification module, used for generating a second bottom-level classification feature vector according to the optimized spectral band information and the hyperspectral texture information; A third bottom-level classification module, used for generating a third bottom-level classification feature vector according to the optimized spectral band information and the hyperspectral color information; A fourth bottom-level classification module, used to generate a fourth bottom-level classification feature vector according to visible light texture information; A fifth bottom-level classification module, used to generate a fifth bottom-level classification feature vector according to visible light color information; The top-level classification module is used to generate a top-level input feature vector based on the first bottom-level classification feature vector, the second bottom-level classification feature vector, the third bottom-level classification feature vector, the fourth bottom-level classification feature vector and the fifth bottom-level classification feature vector, and identify the origin of honeysuckle according to the classification result of the top-level input feature vector.
9. A server, characterized in that: The server comprises: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method for identifying the origin of honeysuckle as described in any one of claims 1-7.
10. A storage medium comprising computer executable instructions, wherein the computer executable instructions, when executed by a computer processor, are used to execute the method for identifying the origin of honeysuckle as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Hyperspectral image classification method based on texture features and affine propagation cluster algorithm
CN108446582A
Chrysanthemum production place identification method, model construction method and electronic equipment
CN116933043A
Lily producing area identification method and device based on near-infrared hyperspectral image, electronic equipment and readable storage medium
CN117333712A
Terahertz time-domain spectroscopy detection method for identifying millets from multiple producing areas
CN118837324A
Method for identifying frostbite condition of grain seeds using spectral feature wavebands of seed embryo hyperspectral images
US20200150051A1