A method for identifying mung bean seed varieties based on Raman spectroscopy
Through the identification method based on Raman spectrum, combined with continuous projection algorithm and particle swarm algorithm to improve the support vector machine model, the existing mung bean seed variety identification method has solved the problem of complex operation and low versatility, and achieved rapid and accurate identification of mung bean varieties.
Patent Information
- Application Number
- CN202211539199.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-01
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-12-01
AI Technical Summary
The existing mung bean seed variety identification methods are complex, difficult and have low versatility, making it difficult to achieve standardized operations under different laboratory conditions.
The identification method based on Raman spectroscopy is adopted, combined with continuous projection algorithm and particle swarm algorithm to improve the support vector machine model to achieve rapid and accurate identification of mung bean seed varieties.
It improves the accuracy and speed of mung bean variety identification, simplifies the operation process, and enhances the adaptability and accuracy of the model.
Smart Images

Figure CN115855913B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of variety identification, and in particular relates to a mung bean seed variety identification method based on Raman spectroscopy. Background Art
[0002] In my country, mung bean has a history of more than 2,000 years of cultivation. Its planting area and annual total output are second only to India, ranking second in the world. As one of the main edible beans grown in my country, mung bean has high nutritional and medicinal value, and plays an important role in adjusting the planting structure, enriching people's diet and processing and utilizing agricultural products in my country. Since the mid-1980s, with the diversification of breeding methods in my country, the existing mung bean varieties have reached hundreds. However, some mung bean varieties in the same region have extremely similar appearances and are difficult to distinguish by human eyes. Counterfeit and inferior mung bean varieties are able to circulate in the market, which has a bad impact on the supervision of the mung bean market order. Therefore, the identification of mung bean seed varieties has become one of the indispensable tasks in maintaining and supervising the mung bean market order.
[0003] As an emerging detection method, spectroscopy technology is applied in the fields of chemical industry, medicine, etc. Among them, Raman spectroscopy is rarely used in the agricultural field, but due to its advantages such as high sensitivity, good reproducibility, convenient spectrum analysis and simple operation, seed identification technology based on Raman spectroscopy can effectively improve the identification accuracy and speed.
[0004] For the variety identification of mung bean seeds, researchers used to use traditional methods such as morphological identification methods to identify varieties by observing the morphological characteristics of plants during the growth and development stages. This is not only time-consuming and laborious, but also the accuracy cannot be guaranteed after leaving the specific identification environment. Since then, researchers have also tried to use biochemical detection methods. For example, the Chinese patent publication number is "CN106244681A", and the name is "A method and application for identifying mung bean varieties using genomic SSR and EST-SSR fingerprint maps". The invention provides a method and application for identifying mung bean varieties using genomic SSR and EST-SSR fingerprint maps, which can complete the identification of mung bean seed varieties and genetic diversity evaluation, and reveal the genetic variation and kinship of varieties from the DNA level. However, the use of this method for mung bean seed variety identification is difficult. Due to the conditions and technical differences of different laboratories, it is difficult to form a standardized operating procedure for a certain crop-specific molecular marker identification method during the identification process of this method. It requires high professionalism from researchers and lacks convenience and versatility in operation. Summary of the invention
[0005] In order to solve the problems of complex operation, high operation difficulty and low versatility in the existing mung bean seed variety identification, the present invention provides a mung bean seed variety identification method based on Raman spectroscopy analysis combined with continuous projection algorithm and particle swarm algorithm to improve the support vector machine model.
[0006] The present invention is achieved through the following technical solutions:
[0007] A method for identifying mung bean seed varieties based on Raman spectroscopy specifically comprises the following steps:
[0008] Step 1, preparing mung bean samples: obtaining mung bean samples, and preparing the mung bean samples into a form capable of collecting Raman spectra;
[0009] Step 2, collecting Raman spectrum: placing the processed sample on a sample placement table of a Raman spectrometer to collect the Raman spectrum and obtain the original spectrum;
[0010] Step 3, preprocessing the spectral data: performing appropriate preprocessing on the original spectrum before qualitative analysis;
[0011] Step 4, selecting characteristic wavelengths: After preprocessing, the characteristic wavelengths are automatically selected using a continuous projection algorithm within the entire band;
[0012] Step 5: Establish a classification model, classify the data, and determine the sample type.
[0013] Furthermore, the preparation of mung bean samples in step 1 specifically involves drying the mung beans in a dust-free, clean, and light-proof drying area to remove the shells, dust and impurities, and then washing and dusting with deionized water to prepare the mung bean samples into three forms: whole grains, slices, and powder.
[0014] Furthermore, the slices are made by using a knife to divide the seeds into two parts along the embryo as the center line; the powder is made by putting the cut seeds into a grinding mortar for grinding and then passing through a mesh sieve.
[0015] Furthermore, in step 2, the Raman spectrum is collected as follows:
[0016] For the whole grain samples obtained, the seeds to be tested are placed on a glass slide under a microscope, and the radicle part and cotyledon part of the seed coat are measured multiple times to obtain the original spectrum;
[0017] For the slice samples obtained, the seeds were cut into two halves with a knife, and the embryo and cotyledon on the cross section were measured at multiple points to obtain the original spectrum;
[0018] For the obtained powder samples, the cut seeds were put into a grinding mortar for grinding, and the mung bean seed powder was obtained after passing through a 100-mesh sieve. The spectrum was measured at multiple points to obtain the original spectrum.
[0019] Furthermore, the preprocessing described in step 3 includes one or more combinations of denoising, data normalization, derivatives and baseline correction; wherein the denoising can adopt SG smoothing, moving average smoothing or wavelet transform; the data normalization can adopt maximum and minimum normalization, standard deviation standardization, mean centering, decimal calibration standardization or robust standardization; the derivative can adopt first-order derivative or second-order derivative, and the baseline correction can adopt multivariate scattering correction, standard normal transformation or trend correction.
[0020] Furthermore, the characteristic wavelength in step 4 is selected by using an algorithm for automatically selecting characteristic wavelengths - a continuous projection algorithm; the principle of this algorithm is that it can not only filter out data redundancy, retain characteristic spectral information reflecting the sample differences, improve the prediction performance of the constructed model, but also reduce the model running time.
[0021] Furthermore, in step 5, a classification model is established, specifically, the spectral data preprocessed in step 3 and the characteristic wavelength selected in step 4 are used as input variables, and an original classification model is established through a support vector machine classification algorithm; then a particle swarm algorithm is used to optimize the penalty parameter C and the kernel function parameter gamma in the original support vector machine model to obtain a support vector machine model improved by combining the particle swarm algorithm.
[0022] Furthermore, the support vector machine model improved by combining the particle swarm algorithm has the following specific process:
[0023] The support vector machine parameters are encoded to form an initial population, which is input into the support vector machine model to return the adaptability value in the particle swarm algorithm to determine whether the maximum number of iterations has been reached; if the maximum number of iterations has been reached, the parameters are decoded to obtain the optimal support vector machine parameters; otherwise, the global optimal individual, speed, and position are updated, and the adaptability value is repeatedly calculated until the maximum number of iterations is reached.
[0024] Compared with the prior art, the advantages of the present invention are as follows:
[0025] 1. The use of Raman spectroscopy to identify the types of mung bean samples can effectively improve the accuracy of mung bean variety identification and shorten the identification time. It has the characteristics and advantages of convenient operation, high sensitivity and good reproducibility;
[0026] 2. The spectral data were collected by making samples in three forms: whole grain, slices, and powder. This can eliminate the influence of different sizes and shapes of mung bean samples and the differences between individuals of non-uniform solid grains, and is beneficial for analyzing the classification effects of three different spectral data and confirming the sample form with the best effect.
[0027] 3. The present invention collects Raman spectra of mung beans, and uses denoising, data normalization, derivative, and baseline correction preprocessing methods combined with feature extraction continuous projection algorithm and particle swarm algorithm to improve support vector machine to establish a mung bean variety classification model, thereby improving the model's ability to cope with changes in conditions such as spectral collection and sample preparation time, location, and environment, and improving the model's adaptability and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the specific embodiments or the description of the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn according to the actual scale.
[0029] Figure 1 This is a flow chart of a method for identifying mung bean seed varieties based on Raman spectroscopy according to the present invention;
[0030] Figure 2 This is a flow chart of the preparation of mung bean powder samples and the collection of Raman spectroscopy information of the present invention;
[0031] Figure 3 A specific combination diagram of the method for preprocessing spectral data of the present invention;
[0032] Figure 4 A combination diagram of a specific method for extracting spectral characteristic information of the present invention;
[0033] Figure 5 The spectrum data of each sample in the powder state of Example 1 of the present invention is worth seven mung bean variety spectral lines;
[0034] Figure 6 These are seven types of spectra of the powdered spectral data of Example 1 of the present invention after being preprocessed by two methods: SG smoothing and standard deviation standardization. DETAILED DESCRIPTION
[0035] In order to clearly and completely describe the technical solution and its specific working process of the present invention, the specific implementation methods of the present invention are as follows in conjunction with the accompanying drawings of the specification:
[0036] Example 1
[0037] like Figure 1 FIG. 1 is a flow chart of a method for identifying mung bean seed varieties based on Raman spectroscopy according to the present embodiment. The identification method specifically comprises the following steps:
[0038] Step 1, prepare mung bean samples: dry the mung beans in a dust-free, clean, light-proof drying area to remove impurities such as husks and dust, rinse with deionized water to remove dust, and then prepare the mung bean samples into three forms: whole grains, slices, and powder. The slice preparation method is: use a knife to divide the embryo into two parts; the powder preparation method is: put the cut seeds into a grinding mortar for grinding, and sieve through a mesh sieve. Figure 2 shown.
[0039] Step 2, collect Raman spectra: for the whole grain samples obtained, place the seeds to be tested on a glass slide under a microscope, and measure the spectra of the radicle part and cotyledon part of the seed coat part of the seeds multiple times to obtain the original spectrum. For the sliced samples obtained, cut the seeds in half with a knife, measure the spectra of the embryo and cotyledon on the cross section multiple times to obtain the original spectrum. For the powder samples obtained, put the cut seeds into a grinding mortar for grinding, pass through a 100-mesh sieve to obtain mung bean seed powder, measure its spectrum multiple times to obtain the original spectrum;
[0040] Step 3, perform spectral preprocessing: The causes of noise include fluorescence background, differences in acquisition instruments, and environmental radiation. These interference factors will greatly affect the accuracy of the spectral classification model. Therefore, before conducting qualitative analysis, appropriate preprocessing operations are required. Preprocessing includes data normalization (maximum and minimum normalization, standard deviation standardization, mean centering, decimal calibration standardization, robust standardization), denoising (SG smoothing, moving average smoothing, wavelet transform), derivatives (first-order derivatives, second-order derivatives), baseline correction (multiple scattering correction, standard normal transformation, trend correction), such as Figure 3 shown.
[0041] Step 4, characteristic wavelength selection: Extracting spectral data can not only filter out data redundancy, retain characteristic spectral information that reflects the sample difference, improve the prediction performance of the established model, but also reduce the model running time. This study uses the continuous projection algorithm to extract characteristic wavelengths in the full band range.
[0042] Step 5, establish a classification model: construct the original support vector machine model and the support vector machine model improved by the particle swarm algorithm for the preprocessed and feature extracted spectra. The algorithm flow chart of the support vector machine improved by the particle swarm algorithm is as follows: Figure 4 As shown in the figure, the discrimination effect and running time of the classification model under different modeling data conditions and different modeling methods are analyzed.
[0043] Example 2
[0044] like Figure 1As shown, a method for identifying mung bean seed varieties based on Raman spectroscopy has the following specific steps: The experimental instrument used in the following embodiments is a TriVista CRS555 micro-area three-stage Raman spectrometer produced by Princeton, with an excitation wavelength of 532nm, a slit of 10μm-12mm, and a resolution of 0.6cm -1 ;
[0045] Step 1, obtaining 7 mung bean samples (including Baolu 201323-3, Zhonglu No. 5, Jilu No. 6, Jilu No. 10, Yilu No. 13, Jilu No. 11 and Jilu No. 9); drying the mung beans in a dust-free, clean and light-proof drying area to remove impurities such as shells and dust, and then washing and dust removal with deionized water to make the mung bean samples into three forms: whole grains, slices and powder; the slice production method is: using a knife to divide the slices into two parts with the embryo as the center line; the powder production method is: putting the cut seeds into a grinding mortar for grinding and passing through a 100-mesh sieve;
[0046] Step 2: For the whole grain sample obtained, the seeds to be tested are placed on a glass slide under a microscope, and the radicle part and the cotyledon part in the seed coat part of the seeds are measured multiple times to obtain the original spectrum; for the sliced sample obtained, the seeds cut into two halves with a knife are placed on a glass slide under a microscope, and the embryo and cotyledon on the cross section are measured multiple times to obtain the original spectrum; for the powder sample ground with a grinding mortar, after passing through a 100-mesh sieve, it is placed on a glass slide under a microscope, and the spectrum is measured multiple times to obtain the original spectrum. Figure 5 To obtain the spectrum data of each sample in powder state, the spectrum lines of each mung bean variety were compared;
[0047] Step 3, determine the best preprocessing combination for the collected original spectra; input multiple combinations of preprocessing methods into the SVM classifier, and select the best preprocessing combination based on the classification effect and running time; the best preprocessing combination selected in this example is SG smoothing + standard deviation normalization SS. After these two preprocessing, the spectra of the seven varieties are as follows Figure 6 .
[0048] The average value of the wavelength point at the i-th position of a sample after SG smoothing is:
[0049]
[0050] Among them, X is the sample, w is the smoothing window coefficient, h k is the smoothing coefficient, H is the normalization factor, i represents the i-th window, and in this embodiment, w is 11, h k The value is 2;
[0051]
[0052] Standard deviation standardization means that the data is centered by the mean and then scaled by the standard deviation. The data will follow a standard normal distribution with a mean of 0 and a variance of 1. The standardization formula is as follows:
[0053]
[0054] Here, μ and σ refer to the mean and standard deviation respectively.
[0055] Step 4, characteristic wavelength selection, will extract characteristic wavelengths from all bands for classification research; the automatic extraction algorithm uses wavelength screening-continuous projection algorithm. The continuous projection algorithm is a forward iterative search method, that is, starting from one wavelength, and then adding a new variable in each iteration until the number of selected variables reaches the set value N; the purpose of the continuous projection algorithm is to select the wavelength with the least redundant spectral information to solve the collinearity problem. For the selection of the number of bands and the starting position, the results of different parameters can be compared for analysis; the number of characteristic wavelengths extracted in this embodiment is 22 (respectively 484.5, 501.7, 532.2, 601.8, 719.1, 773.0, 779.7, 800.1, 1097.2, 1145.2, 1267.2, 1298.7, 1346.6, 1371.3, 1482.2, 1502.8, 1526.1, 1563.1, 1613.8, 1667.2, 1691.9, 1731.6), the unit is cm- 1 .
[0056] Step 5: Build a classification model:
[0057] The spectral data preprocessed in step 3 and the characteristic wavelength selected in step 4 are used as input variables, and the original support vector machine model is established through the support vector machine classification method; the SVM model is established, and the low-dimensional input x output y is converted into the inner product of the high-dimensional space through the kernel function; the parameters of the optimized SVM are usually the penalty parameter C and the kernel function parameter gamma; the selection of the penalty parameter C can achieve a compromise between the model complexity and the training error; the kernel function parameter gamma mainly reflects the range characteristics of the training sample data, which directly affects the learning ability of the support vector machine model;
[0058] The particle swarm optimization algorithm is combined with the particle swarm algorithm to optimize the support vector machine model. By using the particle swarm algorithm to optimize the penalty parameter C and the kernel function parameter gamma in the original support vector machine model, the improved support vector machine model combined with the particle swarm algorithm is obtained. The basic idea of the particle swarm optimization algorithm is to find the optimal solution through collaboration and information sharing between individuals in the group. Its advantage is that it is simple and easy to implement and there is no need to adjust many parameters. The basic particle swarm algorithm steps are relatively simple. The particle swarm optimization algorithm is a group of particles moving in the search space, affected by its own best past position pbest and the best past position gbest of the entire group or neighbors. The d-dimensional velocity update formula of particle i in each iteration is:
[0059] v id k+1 =v id k +c1r1(pbest id k -x id k )+c2r2(gbeSt id k -x id k )x id k+1 =x id k +v id k+1
[0060] The speed and position of each particle are initialized randomly. Then the particles move towards the global optimum and the individual optimum. The parameters are explained as follows
[0061] v id k is the velocity of particle i in the kth iteration;
[0062] x id k is the current position of particle i in the kth iteration;
[0063] pbest id k is the position (coordinate) of the individual pole of particle i;
[0064] gbest id k is the location of the global pole of the entire population;
[0065] r1, r2 are random numbers between [0, 1];
[0066] d represents the dimension;
[0067] c1, c2 are acceleration coefficients (learning factors), which represent the statistical acceleration weights of each particle pushed to the pbest and gbest positions.
[0068] The model flow chart of combining particle swarm algorithm and support vector machine is as follows Figure 4 , encode the support vector machine parameters to form the initial population, input it into the support vector machine model, return the adaptive value in the particle swarm algorithm, and judge whether the maximum number of iterations has been reached. If the maximum number of iterations has been reached, the parameters are decoded to obtain the optimal support vector machine parameters; otherwise, the global optimal individual, speed, and position are updated, and the adaptive value is calculated repeatedly until the maximum number of iterations is reached.
[0069] Specifically, in this embodiment, the SVM kernel function is selected: polynomial kernel, the formula is as follows:
[0070]
[0071] Among them, x i ,x j are two different vectors; the SVM parameters automatically selected after particle swarm algorithm optimization are C: 6.8 and gamma: 0.81; the results show that the accuracy of the support vector machine classification model improved by particle swarm algorithm is 96.37%, which is higher than the accuracy of the original support vector machine classification model of 92.90%.
[0072] The preferred embodiments of the present invention are described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, a variety of simple modifications can be made to the technical solution of the present invention, and these simple modifications all belong to the protection scope of the present invention.
[0073] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the present invention will not further describe various possible combinations.
[0074] In addition, various embodiments of the present invention may be arbitrarily combined, and as long as they do not violate the concept of the present invention, they should also be regarded as the contents disclosed by the present invention.
Claims
1. A method for identifying mung bean seed varieties based on Raman spectroscopy. It is characterized in that The specific steps include: Step 1, preparing mung bean samples: obtaining mung bean samples, and preparing the mung bean samples into a form capable of collecting Raman spectra; Step 2, collecting Raman spectrum: placing the processed sample on a sample placement table of a Raman spectrometer to collect the Raman spectrum and obtain the original spectrum; Step 3, preprocessing the spectral data: performing appropriate preprocessing on the original spectrum before qualitative analysis; Step 4, Select characteristic wavelength: After preprocessing, the characteristic wavelength is automatically selected in the whole band using the continuous projection algorithm; Step 5: Establish a classification model, classify the data, and determine the sample type; The step 1 of preparing mung bean samples is to dry the mung beans in a dust-free, clean, and light-proof drying area, remove the shells, dust and impurities, and then rinse and remove dust with deionized water to prepare the mung bean samples into three forms: whole grains, slices, and powder; The method of making slices is: using a knife to divide the seeds into two parts with the embryo as the center line; the method of making powder is: putting the cut seeds into a grinding mortar for grinding and passing through a mesh sieve; In step 2, the Raman spectrum is collected as follows: For the whole grain samples obtained, the seeds to be tested are placed on a glass slide under a microscope, and the radicle part and cotyledon part of the seed coat are measured multiple times to obtain the original spectrum; For the slice samples obtained, the seeds were cut into two halves with a knife, and the embryo and cotyledon on the cross section were measured at multiple points to obtain the original spectrum; For the obtained powder samples, the cut seeds were put into a grinding mortar for grinding, and the mung bean seed powder was obtained after passing through a 100-mesh sieve. The spectrum was measured at multiple points to obtain the original spectrum.
2. A method for identifying mung bean seed varieties based on Raman spectroscopy as claimed in claim 1, It is characterized in that The preprocessing described in step 3 includes one or more combinations of denoising, data normalization, derivatives and baseline correction; wherein the denoising adopts SG smoothing, moving average smoothing or wavelet transform; the data normalization adopts maximum and minimum normalization, standard deviation standardization, mean centering, decimal calibration standardization or robust standardization; the derivative adopts first-order derivative or second-order derivative, and the baseline correction adopts multivariate scattering correction, standard normal transformation or trend correction.
3. A method for identifying mung bean seed varieties based on Raman spectroscopy as claimed in claim 1, It is characterized in that The characteristic wavelength in step 4 is selected by using an algorithm for automatically selecting characteristic wavelengths - a continuous projection algorithm.
4. A method for identifying mung bean seed varieties based on Raman spectroscopy as claimed in claim 1, It is characterized in that In the step 5, a classification model is established, specifically, the spectral data preprocessed in step 3 and the characteristic wavelength selected in step 4 are used as input variables, and an original support vector machine model is established through a support vector machine classification method; then, a particle swarm algorithm is used to optimize the penalty parameter C and the kernel function parameter gamma in the original support vector machine model to obtain a support vector machine model improved by combining the particle swarm algorithm.
5. A method for identifying mung bean seed varieties based on Raman spectroscopy as claimed in claim 4, It is characterized in that The original support vector machine model is established by the support vector machine classification method, and the specific process is as follows: The support vector machine parameters are encoded to form an initial population, which is input into the support vector machine model to return the adaptability value in the particle swarm algorithm to determine whether the maximum number of iterations has been reached; if the maximum number of iterations has been reached, the parameters are decoded to obtain the optimal support vector machine parameters; otherwise, the global optimal individual, speed, and position are updated, and the adaptability value is repeatedly calculated until the maximum number of iterations is reached.
Citation Information
Patent Citations
Method for identifying mung bean variety by utilizing genome SSR and EST-SSR finger-prints and applications of method
CN106244681A
Castor bean leaf identification method
CN110146641A
Rapid identification method for Arabica coffee beans and Roberska coffee beans
CN112782148A