Visible-near infrared hyperspectrum and GLCM texture feature fusion-based cumin defective particle SVM detection method
By fusion of visible-near-infrared hyperspectral and GLCM texture characteristics, the SVM detection method is constructed, which improves the accuracy and efficiency of Ku Mingzi defect recognition, and realizes high-precision Ku Mingzi automatic sorting and quality control.
Patent Information
- Application Number
- CN202510793065.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-08-15
AI Technical Summary
The existing technology has a low recognition rate in Ku Mingzi defect recognition, while the traditional visible light imaging technology is only 72.3±3.5%. Although the near-infrared spectroscopy technology has been improved to 86.4±2.1%, it still cannot meet the requirements of industrial production.
Fusion of visible-near-infrared hyperspectral and GLCM texture features, through standardized sample preparation and parameter optimization, an efficient dual-modal classification model was constructed, and SPA optimized spectral features and GLCM texture features were used to establish an SVM detection method.
The accuracy of Ku Mingzi defect detection is increased to 94.44%, and the training time is shortened by 68%, solving the problems of low efficiency and high misjudgment rate of traditional methods.
Smart Images

Figure CN120489997A_ABST
Abstract
Description
[0001] The present invention belongs to the technical field of nondestructive testing of agricultural product quality, and specifically relates to a cumin seed defective kernel SVM detection method based on the fusion of visible-near infrared hyperspectral and GLCM texture features, which is suitable for the automated sorting and quality control of cumin seeds. Background Art
[0002] In the field of agricultural product quality testing, the accurate identification of quality defects in kummer seeds, a high-value spice crop, has always faced significant technical challenges. Experimental data shows that traditional visible light imaging technology has a recognition rate of only 72.3±3.5% for moldy kummer seeds. While near-infrared spectroscopy improves this accuracy to 86.4±2.1%, it still falls short of the precision required for industrial production. Summary of the Invention
[0003] To address the above problems, the present invention provides a cumin seed defect kernel SVM detection method based on the fusion of visible-near-infrared hyperspectral and GLCM texture features. It innovatively integrates SPA optimized spectral features and GLCM texture features, and constructs an efficient bimodal classification model through standardized sample preparation and parameter optimization. The accuracy of cumin seed defect detection is improved to 94.44%, and the training time is shortened by 68%, solving the industry problems of low efficiency and high misjudgment rate of traditional methods.
[0004] To achieve the above object, the present invention provides the following solutions:
[0005] The present invention proposes a cumin seed defective kernel SVM detection method based on the fusion of visible-near infrared hyperspectral and GLCM texture features, comprising the following steps:
[0006] Step 1: manually select normal cumin seeds, and prepare the same number of heat-damaged samples and the same number of moldy samples through high-temperature gradient treatment and inoculation with Aspergillus niger, respectively, to construct three types of defect sample sets;
[0007] Step 2: Use a visible-near infrared (Vis-NIR) hyperspectral imaging system to collect sample data, extract the spectrum and image of each sample, and correct the original spectral data;
[0008] Step 3: Convert the input image into a grayscale image and quantize it. Calculate the grayscale co-occurrence frequency of pixel pairs to construct a GLCM matrix. Use a 5×5 sliding window for smoothing and extract nine texture feature indicators including contrast, entropy, and homogeneity.
[0009] Step 4: Use Matlab to preprocess the spectral data and standardize the 9 features extracted from the gray-level co-occurrence matrix GLCM through minimum-maximum normalization;
[0010] Step 5: Use the SPA algorithm to screen characteristic wavelengths, eliminate spectral redundant information, and optimize model input features. The physicochemical properties of moldy kernels, heat-damaged kernels, and normal kernels are then analyzed.
[0011] Step 6: Artificially fuse the characteristic wavelengths screened by SPA with the GLCM texture features to establish an SVM classification model based on the hyperspectral information of moldy, heat-damaged, and normal cumin seeds at 400-1000nm.
[0012] Step 7: After preprocessing, dual-modal feature extraction (SPA screening light + GLCM texture) and feature fusion, the cumin subsample to be tested is input into the trained SVM model for classification and recognition, and the detection results of moldy kernels, heat-damaged kernels or normal kernels are output.
[0013] 500 normal cumin seeds were manually selected and subjected to high temperature gradients of 60℃, 80℃, 100℃, 120℃, and 140℃ to prepare heat-damaged seeds of varying degrees. 500 moldy samples were prepared by inoculating spore suspension of Aspergillus niger and deionized water onto the cumin seeds. Three types of defective sample sets were constructed.
[0014] Sample data was collected using a visible-near-infrared (Vis-NIR) hyperspectral imaging system. The platform speed was set to 7 mm / s and the travel distance was set to 80 to 340 mm. The exposure time was 3 ms. GLCM feature extraction was optimized using Python. The successive projection algorithm (SPA) was then used to filter key characteristic wavelengths from the raw spectra.
[0015] The filtered characteristic wavelengths were fused with the GLCM texture features to construct a multimodal feature set. The iToolbox (Eigenvector Research Inc., Wenatchee, USA) and PLS toolbox in Matlab (R2019b, MathWorks, Inc., USA) were used to establish a prediction model to achieve accurate classification of cumin seed defects.
[0016] The present invention extracts the spectrum and image features of cumin seeds, improves the accuracy and efficiency of defect identification, and has important application value.
[0017] Beneficial effects:
[0018] The hyperspectral imaging system used in this application collects data in the range of 400-1000nm, and the original spectrum contains 440 bands. Through the innovative SPA characteristic wavelength screening algorithm, the characteristic dimensions were successfully reduced from 440 to 264, with a dimensionality reduction of 40%, while retaining 98.7% of the effective classification information. Experimental data show that the characteristic wavelengths after screening are mainly concentrated in three key areas: 420-450nm (mildew characteristic area), 550-580nm (color change area) and 680-720nm (heat damage characteristic area). These bands contain most of the characteristic information for the quality detection of cumin seeds.
[0019] Compared with existing technologies, this study's feature selection method demonstrates significant advantages. While traditional PCA methods lose 23.5% of effective information at the same dimensionality reduction ratio, this method fuses 264 optimized characteristic wavelengths with GLCM texture features, reducing the training time of the SVMDA model from 38 minutes to 12 minutes while simultaneously increasing classification accuracy from 89.2% to 94.44%, achieving a breakthrough in both efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a flow chart of the SVM detection method for cumin defective kernels based on the fusion of visible-near-infrared hyperspectral and GLCM texture features.
[0021] Figure 2 This is the process of extracting the region of interest (ROI) of cumin seeds.
[0022] Figure 3 Visualization of the nine extracted GLCM texture features.
[0023] Figure 4 Scree plot for determining the number of characteristic bands to be screened.
[0024] Figure 5 This is the distribution of the characteristic wavelengths obtained by screening in the 400-1000nm spectrum. DETAILED DESCRIPTION
[0025] The technical solution of the present invention is further described in detail with reference to the following specific examples.
[0026] The present invention aims to provide a method for detecting defective cumin seeds using a support-virtual machine (SVM) method based on the fusion of visible-near-infrared hyperspectral and GLCM texture features. This method can screen the characteristic wavelengths of three types of cumin seeds, namely, moldy, heat-damaged, and normal, at 400-1000 nm. This method clarifies the impact of HSI imaging technology on optical nondestructive testing of grain and oil quality, eliminates redundant interference information, and performs feature-level fusion to achieve high-accuracy classification.
[0027] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Figure 1 , the method specifically includes:
[0028] Step 1: manually select normal cumin seeds, and prepare the same number of heat-damaged samples and the same number of moldy samples through high-temperature gradient treatment and mold culture, respectively, to construct three types of defect sample sets;
[0029] Step 2: Use a visible-near infrared (Vis-NIR) hyperspectral imaging system to collect sample data, extract the spectrum and image of each sample, and correct the original spectral data;
[0030] Step 3: Convert the input image into a grayscale image and quantize it. Calculate the grayscale co-occurrence frequency of pixel pairs to construct a GLCM matrix. Use a 5×5 sliding window for smoothing and extract nine texture feature indicators including contrast, entropy, and homogeneity.
[0031] Step 4: Use Matlab to preprocess the spectral data and standardize the nine features extracted from the gray-level co-occurrence matrix GLCM through minimum-maximum normalization;
[0032] Step 5: Use the SPA algorithm to screen characteristic wavelengths, eliminate spectral redundant information, and optimize model input features. The physicochemical properties of moldy kernels, heat-damaged kernels, and normal kernels are then analyzed.
[0033] Step 6: By manually splicing and fusing the characteristic wavelengths screened by SPA and the GLCM texture features, an SVM classification model based on the hyperspectral information of moldy kernels, heat-damaged kernels, and normal kernels at 400-1000nm was established.
[0034] Step 7: After preprocessing, dual-modal feature extraction (SPA screening light + GLCM texture) and feature fusion, the cumin subsample to be tested is input into the trained SVM model for classification and recognition, and the detection results of moldy kernels, heat-damaged kernels or normal kernels are output.
[0035] Among them, step one specifically includes:
[0036] Four 2g portions of cumin seeds were weighed and placed into a Petri dish. 0.5mL of an Aspergillus niger spore suspension was then added to the dish. The dish was then placed in a mold incubator set at 32±1°C and 90±5% relative humidity. Moldy cumin seeds were collected after 72 hours. The cumin seeds were then placed in a forced air drying oven at 60°C, 80°C, 100°C, 120°C, and 140°C for 2 hours to simulate heat damage. A total of 500 heat-damaged seeds were collected. Finally, 500 normal cumin seeds, each showing a plump, intact appearance without noticeable discoloration or damage, were manually selected. This resulted in the collection of 1500 cumin seeds from the three experimental categories.
[0037] Among them, step 2 specifically includes:
[0038] (1) Hyperspectral images of the three types of cumin seeds prepared in the above steps were collected using a visible hyperspectral imaging system (Vis-NIR). The Vis-NIR system consisted of an 804×440 pixel ICLB1620 CCD camera (Imperx Co., Boca Raton, FL, USA), an inspector V10E imaging spectrometer with a spectral resolution of 382.67-1010.64 nm (Specim, Oulu, Finland), a halogen light source, a mobile platform, and a Dell computer. The two halogen light sources were fixed at a 45° angle above the sample. The height between the sample and the camera was 28 cm. The speed of the mobile platform was set to 7 mm / s, and the moving distance was set to 80-340 mm. The exposure time was 3 ms. To avoid the influence of light, the entire acquisition process was completed in a dark box. Before the formal acquisition, the light source needed to be turned on in advance for 30 minutes to preheat. To avoid background influence and facilitate subsequent data processing, the cumin seeds particles were neatly arranged on an EVA foam board in a 10×10 size.
[0039] (2) In order to eliminate the potential interference of dark current and uneven light source distribution on spectral information, it is crucial to calibrate the collected spectrum. In the experiment, a black and white plate calibration method was used to eliminate redundant information. A polytetrafluoroethylene standard calibration white plate with a reflectivity of up to 99.99% was placed to obtain a full white reflection image. The opaque lens cover of the camera was then placed to obtain a full black reflection image. Based on these two images, Formula 1 was introduced to calculate the corrected relative reflection image.
[0040] Formula 1 is the hyperspectral image correction formula:
[0041] R=(R0-B) / (WB) (1)
[0042] Where: R is the hyperspectral image after black and white correction; R0 is the original hyperspectral image; B is the all-black reflection image after covering the camera lens cap; W is the all-white reflection image collected by scanning the whiteboard.
[0043] Matlab R2019b software was used to extract spectral information and analyze it. The region of interest (ROI) was extracted. In this study, the average spectrum of each cumin seed was extracted as the ROI. The process of extracting hyperspectral data is as follows: Figure 2 , where the average spectrum of all pixels in each region of interest is the original spectrum of the region. The total number of pixels and effective spectral values are counted, and the average spectrum is extracted for subsequent data processing and analysis.
[0044] Step three includes
[0045] Hyperspectral technology boasts the unique ability to combine image and spectrum. In addition to spectral data, hyperspectral data can also provide image information. This study utilized the texture and shape information of cumin seeds. The gray-level co-occurrence matrix (GLCM) is a common method for describing grayscale texture features. The GLCM can convert two-dimensional images into one-dimensional data, facilitating subsequent data fusion.
[0046] Python 3.8 (Python Software Foundation, USA) was used to adapt the GLCM open source code obtained from GitHub. The final GLCM parameters were set as follows: quantization level nbit = 8 (grayscale value is quantized to 8 bits), kernel size ks = 5, grayscale range mi = 0, ma = 255, and a 5×5 kernel was used for filtering and smoothing. Nine GLCM features were calculated: mean, standard deviation (std), contrast, dissimilarity, homogeneity, angular second moment (ASM), energy, maximum value (max), and entropy. The texture features extracted were visualized as shown in the figure. Figure 3 As shown, the differences in texture characteristics of kumen seeds under different abnormal grain conditions are revealed.
[0047] Among them, step four includes
[0048] In addition to chemical information about the sample, acquired spectra also contain some useless information and noise. To reduce or eliminate the effects of useless information in the raw spectral data, including noise, background color, sample surface inhomogeneity, baseline drift, and absorption peak overlap, and to further improve the accuracy and stability of subsequent analysis and model building, spectral data preprocessing is necessary. This study employed five different preprocessing methods to preprocess the acquired spectra: smoothing (SG Smoothing), standard normalized variate (SNV), multiplicative scatter correction (MSC), first derivative (1-st), and second derivative (2-nd). These preprocessing methods were compared and analyzed to determine the optimal spectral preprocessing method. The results showed that SNV processing often achieved better performance across multiple models.
[0049] This code primarily utilizes the os, pandas, and numpy core libraries from Python 3.8 (Python Software Foundation, USA). Its normalization utilizes the min-max normalization principle, applying a linear transformation to the "mean" of each GLCM texture feature (e.g., nine metrics such as contrast and entropy): (mean - min_val) / (max_val - min_val). This transforms the original feature values to the interval [0, 1]. When the minimum and maximum values of a feature are equal, the value is set to 0.5 to avoid division by zero errors. Normalization eliminates dimensional differences and numerical range deviations between features, making them comparable. This improves subsequent model training and convergence speed, while preserving the distribution of the original data and establishing a standardized data foundation for multi-sample feature integration.
[0050] Step five includes
[0051] In order to improve the overfitting phenomenon, it was decided to perform characteristic wavelength screening on the spectral data, and SPA (Successive Projections Algorithm) was used for characteristic wavelength screening. SPA is a selection strategy based on forward loop. The SPA algorithm used in this study was written and run using the editor in MATLAB 2019b. In the process of using the SPA algorithm to screen characteristic wavelengths, the root mean square error (RMSE) values under different numbers of characteristic wavelengths were calculated, and the number of variables corresponding to the minimum RMSE value was used as the number of characteristic wavelengths screened by the algorithm. The process of determining the final number of characteristic wavelengths by RMSE in the process of using SPA to screen characteristic wavelengths in this study is as follows. Figure 4 As shown in the scree plot, we can clearly see that the RMSE value tends to a stable low value when the characteristic wavelength is 264, so we decided to select 264 characteristic wavelengths. SPA constructs a feature set with minimum redundant information by iteratively selecting characteristic wavelengths, which contains spectral information at 264 wavelengths.
[0052] Among them, according to Figure 5 From the distribution of characteristic wavelengths on the 400-1000nm spectrum, it can be seen that the screened characteristic wavelengths are mainly concentrated in the two ranges of 420-640nm and 780-920nm: the 420-640nm band mainly corresponds to the absorption characteristics of pigments such as carotenoids (420-500nm) and chlorophyll (500-640nm). Normal grains have typical spectral characteristics in this range, while moldy grains will cause abnormal reflectivity in this band due to fungal metabolites (such as mold pigments) and chlorophyll degradation. Heat-damaged grains will show characteristic shifts due to the destruction of pigment structure caused by high-temperature oxidation; the 780-920nm near-infrared range mainly reflects the vibration information of organic molecular bonds such as CH and OH. Moldy grains will produce unique metabolites (such as mycotoxins, polysaccharides, etc.) due to microbial activity. Heat-damaged grains will change the molecular bond vibration mode due to Maillard reaction and protein denaturation. These physical and chemical changes were effectively captured as characteristic wavelengths through the SPA algorithm, indicating that the screening results not only conform to the principles of spectroscopy, but can also accurately distinguish the essential differences between the three types of samples in terms of pigment composition, molecular structure and metabolites.
[0053] Among them, step six includes
[0054] (1) PLSDA (Partial Least Squares Discriminant Analysis) is a discriminant analysis method based on partial least squares regression (PLS) and is used for classification problems. It compresses data to a lower dimension through dimensionality reduction and feature selection, while using known category information for classification. (Support Vector Machine Discriminant Analysis) Support Vector Machine Discriminant Analysis is a classification method based on support vector machine (SVM) that separates different categories of data by finding the optimal hyperplane. This method is good at handling nonlinear classification problems and maps data to a high-dimensional space through a kernel function, thereby improving classification accuracy. We usually measure the performance of a classification model with overall accuracy, precision, recall, specificity, and F1-score.
[0055] (2) First, the pure spectral data extracted and preprocessed in steps 2 and 4 were put into the PLSDA and SVMDA classification models. The results are shown in Table 1. It can be seen that no matter what kind of preprocessing the full spectral data undergoes, the performance in the two different classification models is not ideal. This may be due to the relatively single data modality and the presence of more redundant wavelength information. Therefore, we turned to consider combining the "image-spectrum integration" feature of hyperspectral images and converted the two-dimensional image into one-dimensional data through GLCM (gray-level co-occurrence matrix). The nine GLCM features extracted, namely mean, standard deviation (std), contrast, dissimilarity, homogeneity, angular second moment (ASM), energy, maximum value (max), and entropy, were manually spliced into the full spectral data and then put into the PLSDA and SVMDA classification models. The performance indicators of the models are shown in Table 2. It can be seen that the accuracy of the sample training set and validation set has been greatly improved. The overall accuracy of the training set is between 77.62% and 100.00%, and the accuracy of the validation set is between 80.00% and 91.78%. However, the model overfitting phenomenon is relatively serious. Overfitting can easily lead to a decrease in the generalization ability of the model, an increase in test error, unstable classification results, and an inability to adapt well to new data, resulting in poor performance of the model in practical applications and unsatisfactory accuracy. After reviewing the process, we believe that the reason for the serious overfitting of the model may be that there is too much redundant information in the data and the features are not obvious.
[0056] Table 1: PLSDA and SVMDA classification model performance for the full spectrum
[0057]
[0058]
[0059] The spectral data at the characteristic wavelengths obtained in step 5, after SPA screening, were manually spliced together with the nine GLCM eigenvalues. After undergoing five preprocessing methods, they were applied to the PLSDA and SVMDA models to analyze the model performance and the underlying reasons. Finally, the performance of the PLSDA and SVMDA classification models, which fuse the characteristic spectrum and GLCM feature texture intermediate data, is shown in Table 3. The accuracy of the training set ranged from 85.43% to 99.05%, while the accuracy of the validation set ranged from 80.89% to 94.44%. We can see that the model significantly improved the accuracy of the validation set without experiencing severe overfitting. This demonstrates that the model has strong generalization capabilities when processing new data, achieving high classification accuracy and better meeting the needs of practical applications.
[0060] Using the SPA (Successive Projection Algorithm) to achieve precise data dimensionality reduction, the team selected 264 of the most representative characteristic wavelengths from the original 440 wavelengths (a dimensionality reduction rate of 40%). This effectively removed redundant information and noise from the spectral data, laying a foundation for high-quality data for subsequent classification. Innovatively, the spectral features selected by SPA were manually combined with nine types of texture features (such as mean and entropy) extracted by GLCM to construct a bimodal feature space. Experimental data demonstrated that this fusion maximized the complementarity between the features: the spectral features captured molecular vibrational information (such as the 680nm chlorophyll degradation peak caused by thermal damage), while the GLCM texture features (using a 5×5 kernel) effectively characterized surface structural changes caused by mold. This can be seen by comparing Tables 2 and 3. This is because SPA filtering eliminates wavelength collinearity, while min-max normalization unifies the scales of the different modal features. The final model achieved 94.44% accuracy on the validation set, providing a new approach and strategy for the detection of defective cumin seeds.
[0061] Table 2: PLSDA and SVMDA classification model performance for primary fusion of full spectrum and GLCM feature texture
[0062]
[0063]
[0064] Table 3: PLSDA and SVMDA classification model performance for mid-level data fusion of feature spectrum and GLCM feature texture
[0065]
[0066]
[0067] Among them, step seven includes,
[0068] First, for the cumin subsample to be tested, data from the 400-1000 nm band was collected using a Vis-NIR hyperspectral imaging system, following the method described in step 2. Black-white correction (Formula 1) and ROI average spectrum extraction were performed. The 264 key characteristic wavelengths selected by the SPA algorithm in step 5 were then reused to extract the corresponding spectral feature data for this sample (the SNV preprocessing in step 4, which was more effective, is recommended). Simultaneously, the GLCM parameters (quantization level nbit = 8, kernel size ks = 5, grayscale range mi = 0, ma = 255) and processing flow (5×5 kernel smoothing) determined in step 3 were reused to extract nine types of texture features from the sample image: mean, standard deviation, contrast, dissimilarity, homogeneity, angular second moment, energy, maximum value, and entropy. These features were then min-max normalized. The extracted 264-dimensional spectral features were then manually concatenated and fused with the 9-dimensional GLCM texture features to form a final 273-dimensional feature vector. This fused feature vector was then input into the SVM classification model trained in step 6. After model analysis, the classification results are output, determining whether the cumin subsample belongs to the "moldy," "heat-damaged," or "normal" category, completing the automated non-destructive testing. This process strictly reuses the parameters, models, and processes established during the training phase to ensure consistent and reliable testing.
[0069] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for a person skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to replace some of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions claimed to be protected by the present invention.
Claims
1. A cumin seed defective kernel SVM detection method based on the fusion of visible-near infrared hyperspectral and GLCM texture features, characterized in that: The following steps are involved: Step 1: manually select normal cumin seeds, and prepare the same number of heat-damaged samples and the same number of moldy samples through high-temperature gradient treatment and mold culture, respectively, to construct three types of defect sample sets; Step 2: Use a visible-near infrared (Vis-NIR) hyperspectral imaging system to collect sample data, extract the spectrum and image of each sample, and correct the original spectral data; Step 3: Convert the input image into a grayscale image and quantize it. Calculate the grayscale co-occurrence frequency of pixel pairs to construct the grayscale co-occurrence matrix (GLCM). Use a 5×5 sliding window for smoothing and extract nine texture feature indices: mean, standard deviation, contrast, dissimilarity, homogeneity, angular second moment, energy, maximum value, and entropy. Step 4: Use Matlab to preprocess the spectral data and standardize the 9 features extracted from the gray-level co-occurrence matrix GLCM through minimum-maximum normalization; Step 5: Use the SPA algorithm to screen characteristic wavelengths, eliminate spectral redundant information, and optimize model input features. The characteristic band distribution is analyzed based on the specific differences between the three types of samples: moldy kernels, heat-damaged kernels, and normal kernels. Step 6: The characteristic wavelengths screened by SPA and the GLCM texture features were fused by manual splicing to establish an SVM classification model based on the hyperspectral information of moldy kernels, heat-damaged kernels and normal kernels at 400-1000nm. Step seven: After preprocessing, bimodal feature extraction and feature fusion, the cumin subsample to be tested is input into the trained SVM model for classification and recognition, and the detection results of moldy kernels, heat-damaged kernels or normal kernels are output.
2. The method for detecting defective cumin seeds based on the support vector machine (SVM) method based on the fusion of visible-near infrared hyperspectral and GLCM texture features as claimed in claim 1, characterized in that: In the step 1, normal cumin seeds are manually selected from untreated primary raw materials; a high temperature gradient is set and a hot air oven is used to obtain heat-damaged cumin seeds; and a spore suspension of Aspergillus niger is inoculated to obtain a moldy cumin seed sample.
3. The method for detecting defective cumin seeds based on the support vector machine (SVM) method based on the fusion of visible-near infrared hyperspectral and GLCM texture features as claimed in claim 1, characterized in that: In step 2, hyperspectral data is acquired using a visible light Vis-NIR system. The Vis-NIR system consists of an ICLB1620 CCD camera with 804×440 pixels, an Inspector V10E imaging spectrometer with a spectral resolution of 382.67-1010.64 nm, a halogen light source, a mobile platform, and a Dell computer. The two halogen light sources are fixed at an angle of 45° above the sample, and the height between the sample and the camera is 28 cm. To avoid the influence of light, the entire acquisition process is completed in a dark box. Grayscale images at 879.61 nm and 1007.32 nm were used to construct mask images, respectively. A threshold algorithm was used to distinguish cumin seeds from the background, ensuring that the location of the cumin seeds was used for region of interest (ROI) selection. Subsequently, the spectral data of each cumin seed was obtained by calculating the average value of all pixels in the ROI.
4. The method for detecting defective cumin seeds based on support vector machine (SVM) based on the fusion of visible-near infrared hyperspectral and GLCM texture features as claimed in claim 1, characterized in that: In step 3, the GLCM open source code obtained from GitHub was adaptively modified, and the GLCM parameters were finally set as follows: quantization level nbit=8, kernel size ks=5, grayscale range mi=0, ma=255, and a 5×5 kernel was used for filtering and smoothing to extract 9 GLCM features.
5. The method for detecting defective cumin seeds based on support vector machine (SVM) based on the fusion of visible-near infrared hyperspectral and GLCM texture features as claimed in claim 1, characterized in that: In the step 4, the spectral data is subjected to five different preprocessing methods, namely, smoothing, first-order derivative, second-order derivative, multivariate scatter correction MSC, and standard normal transformation SNV. The nine types of features extracted from the GLCM gray-level co-occurrence matrix extracted in the step 2 are normalized using minimum-maximum normalization.
6. The method for detecting defective cumin seeds using a support vector machine (SVM) based on the fusion of visible-near infrared hyperspectral and GLCM texture features as claimed in claim 1, characterized in that: The SPA algorithm in step five selects the most representative wavelength variable by minimizing the collinearity between variables, removes redundant information in the spectral data, provides high-quality input data for the subsequent classification model, and improves the performance of the model.
7. The method for detecting defective cumin seeds by using a support vector machine (SVM) based on the fusion of visible-near infrared hyperspectral and GLCM texture features as claimed in claim 1, characterized in that: In steps six and seven, the characteristic wavelength obtained by the SPA algorithm screening in step five is combined with the GLCM texture feature data, and the iToolbox toolbox in Matlab is used to establish an SVM classification model.
Citation Information
Cited By
Method for discriminating different types of pork based on hyperspectral imaging and machine learning
CN121708461A