Peanut kernel producing area identification method based on near infrared spectrum
By using a portable near-infrared spectrometer and fuzzy discrimination analysis method, combined with a K-nearest neighbor classifier, the problem of low accuracy in identifying the origin of peanut kernels was solved, achieving efficient and accurate identification of the origin of peanut kernels.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHUZHOU VOCATIONAL & TECHN COLLEGE
- Filing Date
- 2024-02-04
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies are insufficient to effectively distinguish peanut kernels from different origins, resulting in low accuracy in identifying the origin of peanut kernels.
A portable near-infrared spectrometer was used to collect spectral data of peanut kernels. The data was preprocessed using multivariate scattering correction and Savitzky-Golay smoothing filter. Fuzzy discrimination analysis and K-nearest neighbor classifier were combined to construct a fuzzy scattering matrix for feature decomposition and transformation, thereby achieving accurate identification of the peanut kernel origin.
It improves the accuracy of peanut kernel origin identification, and achieves rapid, non-destructive, green and environmentally friendly origin identification.
Smart Images

Figure CN122045797A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of pattern recognition and artificial intelligence, specifically to a near-infrared spectroscopy method for identifying the origin of peanut kernels, which is applied to the accurate identification of the origin of peanut kernels. Background Technology
[0002] Peanuts are a widely cultivated oilseed crop worldwide, rich in nutrients, containing abundant oils, protein, carbohydrates, vitamin E, and various trace elements. Furthermore, peanuts possess certain health benefits, such as combating diabetic complications, cancer, cognitive impairment, and cardiovascular disease. Besides being eaten directly, peanuts are commonly used in cooking, such as stir-frying, making peanut butter, or as an ingredient in desserts. The texture, taste, and oil content of peanuts can vary depending on their origin. Peanuts from different regions exhibit differences in quality and yield due to factors such as soil, climate, and sunlight exposure. Therefore, it is necessary to provide a simple, easy-to-use, and efficient method for identifying the origin of peanuts.
[0003] Near-infrared spectroscopy is a non-destructive testing technique used to assess and analyze the molecular composition of materials through the interaction of samples with near-infrared light. It is characterized by its simplicity, non-destructive nature, speed, and low cost, and is therefore widely used in origin tracing research. The wavelength of near-infrared spectra is approximately between 800 nm and 2500 nm. The number of amino and hydroxyl groups in peanut kernels from different origins varies, resulting in differences in spectral characteristics in the near-infrared band. By scanning and processing the near-infrared spectra of peanut kernels, their origin can be effectively identified.
[0004] Fuzzy identification is a data analysis technique based on fuzzy mathematics theory, used to process and analyze data containing uncertainty and fuzziness. The core of fuzzy identification is applying fuzzy set theory to the classification decision-making process. In practice, many large datasets contain a certain amount of noise; fuzzy identification can effectively reduce the impact of noise on classification results. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention proposes a near-infrared spectroscopy method for identifying the origin of peanut kernels. A fuzzy scattering matrix is constructed, and a new fuzzy intra-class scattering matrix is calculated to obtain a set of identification vectors in the value domain and null space. The original data is projected onto this set of identification vectors to obtain the transformed data. Then, the K-nearest neighbor method is used to classify the data and derive the identification accuracy. This near-infrared spectroscopy method for identifying the origin of peanut kernels is used to identify the origin of peanut kernels, enabling the method to extract identification information from the near-infrared spectrum of peanut kernels and perform origin identification, thereby improving the identification accuracy.
[0006] A near-infrared spectroscopy method for identifying the origin of peanut kernels includes the following steps:
[0007] S1. A portable near-infrared spectrometer was used to collect spectral data of peanut kernel samples to obtain near-infrared diffuse reflectance spectral data of the samples.
[0008] S2 employs multivariate scattering correction and Savitzky-Golay (SG) smoothing filters to preprocess the spectral data, reducing noise and scattering effects in the spectral data.
[0009] S3 employs a fuzzy discrimination analysis method to compress and extract discrimination information from the spectral data of S2. The specific steps are as follows:
[0010] S4.1, Initialization: The peanut kernel spectral data after preprocessing in S2 is divided into training samples and test samples. The number of training samples is n. tr Number of test samples n te The weight index m and the number of categories c, where m > 1;
[0011] S4.2, Calculate the training sample x k Fuzzy membership degree u belonging to class i (1≤i≤c) ik :
[0012]
[0013] Where c is the number of categories; m is the weight index; Let be the mean of the i-th class of training samples. Let be the mean of the j-th class of the training samples;
[0014] S4.3, Calculate the fuzzy inter-class discrete matrix S of the training samples. fB ,Fuzzy total scattering matrix S fT and the fuzzy intraclass discreteness matrix S fW :
[0015]
[0016]
[0017]
[0018] in, This indicates that when m is known, the training sample x k Fuzzy membership degree belonging to class i (1≤i≤c); The total mean of the training samples. Let be the mean of the training samples of the i-th class.
[0019] S4.4, regarding the fuzzy global scattering matrix S fT Perform singular value decomposition and calculate s(s=rank(S) fT The eigenvectors form a matrix P. A new fuzzy scattering matrix is constructed: For the new fuzzy intraclass scattering matrix Eigenvalue decomposition yields a matrix V = [v1,...,v2] composed of eigenvectors. q ,v q+1 ,...,v p Here, q = sc, p = s-1. Let V1 = [v1,...,v...] q ], V2=[v q+1 ,...,v p ].
[0020] S4.5, Constructing the matrix λ -1 S fB The eigenvectors obtained from eigenvalue decomposition are: z1,...,z s The discrimination vector g is calculated. k =PV2z k ,k=1,...,t,t≤c-1.
[0021] S4.6, Constructing the matrix Calculate matrix The eigenvectors are: y1,...,y d The discrimination vector is calculated as: u v =PV1y v v = 1, ..., d, d ≤ c - 1.
[0022] S4.7, construct the transformation matrix W = [GU], where G = g1,...,g t U = u1,...,u d Then, the training samples and test samples are multiplied by the transformation matrix W respectively to achieve spatial transformation between the training samples and the test samples.
[0023] S5 uses the K-nearest neighbor classifier to classify the test samples transformed in S4.7 and calculates the classification accuracy.
[0024] The beneficial effects of this invention are:
[0025] The present invention provides a method for identifying the origin of peanut kernels using near-infrared spectroscopy, which can extract the optimal identification vector of the near-infrared spectrum of peanut kernels, improve the identification accuracy, and enable rapid identification of the origin of peanut kernels.
[0026] The present invention provides a method for identifying the origin of peanut kernels using near-infrared spectroscopy. Due to the use of a portable near-infrared spectrometer, the method of the present invention can quickly and non-destructively detect peanut kernels, making it green and environmentally friendly. Attached Figure Description
[0027] Figure 1 This is a flowchart of a near-infrared spectroscopy method for identifying the origin of peanut kernels according to the present invention.
[0028] Figure 2 Near-infrared spectral data for peanut kernels;
[0029] Figure 3 The image shows the near-infrared spectrum of peanut kernels after preprocessing with the multivariate scattering correction method (MSC) and Savitzky-Golay (SG) smoothing filter.
[0030] Figure 4 This is a fuzzy membership graph. Detailed Implementation
[0031] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.
[0032] like Figure 1 As shown, a near-infrared spectroscopy method for identifying the origin of peanut kernels specifically includes the following steps:
[0033] S1. A portable near-infrared spectrometer was used to collect spectral data of peanut kernel samples to obtain near-infrared diffuse reflectance spectral data of the samples.
[0034] Four types of peanut kernels, sourced from Qingdao, Yantai, Heze, and Linyi respectively, were selected for the experiment. Sixty samples were collected from each variety, totaling 240 samples. Under constant ambient temperature and humidity, the near-infrared spectra of the peanut kernels were collected after the portable near-infrared spectrometer NIR-M-R2 was preheated for 30 minutes. The obtained near-infrared spectral data are shown below. Figure 2 As shown.
[0035] S2, using multivariate scattering correction and a Savitzky-Golay (SG) smoothing filter to preprocess the spectral data, reduces noise and scattering effects. The preprocessed spectrum is shown below. Figure 3 As shown.
[0036] S3 employs a fuzzy discrimination analysis method to compress and extract discrimination information from the spectral data of S2. The specific steps are as follows:
[0037] S4.1, Initialization: The peanut kernel spectral data after preprocessing in S2 is divided into training samples and test samples. The number of training samples n for each type of peanut kernel is... tr =50, Number of test samples n te=10; the number of peanut kernel categories c=4, then the total number of training samples is 200, and the total number of test samples is 40; the weight index m=2.
[0038] S4.2, Calculate the training sample x k Fuzzy membership degree u belonging to class i (1≤i≤c) ik :
[0039]
[0040] Where c is the number of categories; m is the weight index; Let be the mean of the i-th class of training samples. The mean of the j-th class of training samples; in this embodiment, m=2, c=4 to obtain the initial fuzzy membership degree u. ik like Figure 4 As shown.
[0041] S4.3, Calculate the fuzzy inter-class discrete matrix S of the training samples. fB ,Fuzzy total scattering matrix S fT and the fuzzy intraclass discreteness matrix S fW :
[0042]
[0043]
[0044]
[0045] in, This indicates that when m=2, the training sample x k Fuzzy membership degree belonging to class i (1≤i≤c); The total mean of the training samples. Let be the mean of the training samples of the i-th class. Write down... and value
[0046]
[0047]
[0048]
[0049]
[0050]
[0051] S4.4, regarding the fuzzy global scattering matrix S fT Perform singular value decomposition and calculate s = rank(S) fTThe matrix P is composed of 194 eigenvectors. A new fuzzy scattering matrix is constructed: For the new fuzzy intraclass scattering matrix Eigenvalue decomposition yields a matrix V = [v1,...,v2] composed of eigenvectors. q ,v q +1,...,v p Here, q = sc = 190, p = s⁻¹ = 193. Let V₁ = [v₁, ..., vₙ]. q ], V2=[v q+1 ,...,v p ].
[0052]
[0053]
[0054] S4.5, Constructing the matrix λ -1 S fB The eigenvectors obtained from eigenvalue decomposition are: z1,...,z t The discrimination vector g is calculated. k =PV2z k ,k=1,...,t,t=3. λ=0.01.
[0055]
[0056] S4.6, Constructing the matrix Calculate matrix The eigenvectors are: y1,…,y d The discrimination vector is calculated as: u v =PV1y v v = 1, ..., d, d = 3.
[0057]
[0058] S4.7 Construct the transformation matrix W = [GU], where G = g1, g2, g3 and U = u1, u2, u3.
[0059]
[0060] Then, the training samples and test samples are multiplied by the transformation matrix W respectively to achieve spatial transformation between the training samples and test samples.
[0061] S5 uses the K-nearest neighbor classifier to classify the test samples transformed in S4.7 and calculates the classification accuracy. In this scheme, the number of "neighbors" in the K-nearest neighbor classifier is K=1, and the final classification accuracy is 92.50%.
[0062] The above embodiments are only used to illustrate the design concept and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made based on the principles and design ideas disclosed in the present invention are within the protection scope of the present invention.
Claims
1. A near-infrared spectroscopy method for identifying the origin of peanut kernels, comprising the following steps: S1. A portable near-infrared spectrometer was used to collect spectral data of peanut kernel samples to obtain near-infrared diffuse reflectance spectral data of the samples; S2 employs multivariate scattering correction and Savitzky-Golay (SG) smoothing filter to preprocess the spectral data, reducing noise and scattering effects in the spectral data; S3 uses a fuzzy discrimination analysis method to compress and extract discrimination information from the spectral data of S2. The specific steps are as follows: S4.1, Initialization: Divide the peanut kernel spectral data after S2 preprocessing into training samples and test samples, with n training samples. tr Number of test samples n te The weight index m and the number of categories c, where m > 1; S4.2, Calculate the training sample x k Fuzzy membership degree u belonging to class i (1≤i≤c) ik : in, c represents the number of categories; m represents the weight index; Let be the mean of the i-th class of training samples. Let be the mean of the j-th class of the training samples; S4.3, Calculate the fuzzy inter-class discrete matrix S of the training samples. fB ,Fuzzy total scattering matrix S fT and the fuzzy intraclass discreteness matrix S fW : in, This indicates that when m is known, the training sample x k Fuzzy membership degree belonging to class i (1≤i≤c); The total mean of the training samples. Let be the mean of the training samples of the i-th class; S4.4, regarding the fuzzy global scattering matrix S fT Perform singular value decomposition and calculate s(s=rank(S) fT The eigenvectors form a matrix P, which is used to construct a new fuzzy scattering matrix: For the new fuzzy intraclass scattering matrix Eigenvalue decomposition yields a matrix V = [v1,...,v2] composed of eigenvectors. q ,v q1 ,...,v p Here, q = sc, p = s-1, let V1 = [v1,...,v] q ], V2=[v q+1 ,...,v p ]; S4.5, Constructing the matrix λ -1 S fB The eigenvectors obtained from eigenvalue decomposition are: z1,...,z s The discrimination vector g is calculated. k =PV2z k k = 1, ..., t, t ≤ c - 1; S4.6, Constructing the matrix Calculate matrix The eigenvectors are: y1,...,y d The discrimination vector is calculated as: u v =PV1y v v = 1, ..., d, d ≤ c - 1; S4.7, construct the transformation matrix W = [GU], where G = g1,...,g t U = u1,...,u d Then, the training samples and test samples are multiplied by the transformation matrix W respectively to achieve the spatial transformation between the training samples and test samples; S5 uses the K-nearest neighbor classifier to classify the test samples transformed in S4.7 and calculates the classification accuracy.