Dried orange peel year identification method and system, terminal equipment and storage medium
Through near-infrared spectroscopy technology and particle swarm optimization support vector machine model, the accuracy problem of tangerine peel year identification was solved, and efficient and low-cost tangerine peel year identification was achieved.
Patent Information
- Application Number
- CN202510824879.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technology makes it difficult to accurately identify the age of dried tangerine peel. The similarities in color, smell and taste make it difficult to distinguish, and the improvement of counterfeiting technology has increased the difficulty of identification.
Near-infrared spectroscopy technology is combined with a support vector machine model optimized by particle swarm optimization. By acquiring and preprocessing the spectral data of tangerine peel, feature screening and parameter optimization are performed, and the support vector machine model is trained to identify the year of tangerine peel.
The accuracy and efficiency of identifying the age of dried tangerine peel are improved, the identification cost is reduced, and the operation is convenient and non-destructive.
Smart Images

Figure CN120673890A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of tangerine peel identification, and in particular to a tangerine peel year identification method, system, terminal device and storage medium. Background Art
[0002] The difficulties in identifying the age of dried tangerine peel mainly include the following aspects: Irregularity of color change: The color change of dried tangerine peel is affected by many factors, such as the temperature and humidity of the storage environment, which leads to irregular color change and it is difficult to judge the age by color alone. Improvement of counterfeiting technology. With the improvement of counterfeiting technology, some unscrupulous merchants artificially make dried tangerine peel old by dyeing, high-temperature drying and other means, which increases the difficulty of distinguishing the authenticity. For example, the inner capsule of dried tangerine peel of high age falls off naturally, while counterfeit dried tangerine peel may make the inner capsule look like it has fallen off through artificial means, which increases the difficulty of identification. Similarity in smell and taste. Dried tangerine peel of different years may have similarities in smell and taste, especially the difference between dried tangerine peel of low age and high age is not obvious enough, making it difficult to accurately judge. Summary of the Invention
[0003] In view of this, the embodiments of the present application provide a method, system, terminal device and storage medium for identifying the age of tangerine peel, which can effectively solve the problem of difficulty in accurately identifying the age of tangerine peel.
[0004] In a first aspect, the present invention provides a method for identifying the age of dried tangerine peel, comprising:
[0005] Acquire spectral data of the tangerine peel to be tested and raw spectral data with labels, and preprocess the spectral data and the raw spectral data;
[0006] Performing feature screening on the pre-processed spectral data to obtain a feature set to be measured;
[0007] The parameters of the support vector machine model are optimized using the particle swarm optimization algorithm;
[0008] Training the support vector machine model after parameter optimization based on the preprocessed original spectral data to obtain a pre-trained support vector machine model;
[0009] The feature set to be tested is input into the pre-trained support vector machine model to determine the year of the tangerine peel to be tested.
[0010] In a first possible embodiment of the first aspect, obtaining spectral data of the tangerine peel to be tested and raw spectral data with labels, and preprocessing the spectral data and the raw spectral data, includes:
[0011] Collecting spectral data of the tangerine peel to be tested after being crushed and spread into a test container;
[0012] The collected spectral data and the original spectral data are sequentially subjected to normal variable transformation processing, wavelet smoothing processing and normalization processing.
[0013] In a second possible embodiment of the first aspect, performing feature screening on the preprocessed spectral data to obtain a feature set to be measured includes:
[0014] The pre-processed raw spectral data is used as a data set, and the data set is divided into a training set;
[0015] The training set is modeled and analyzed by partial least squares method to determine the weight of each wavelength information in the spectral data for the identification of the year of tangerine peel;
[0016] Determine the number of wavelength information retained in each iterative sampling;
[0017] performing multiple rounds of sampling on the wavelength information in the spectral data based on the determined number of wavelength information and the weight of each wavelength information to obtain a screening variable for each round of sampling;
[0018] Based on the analysis and prediction model of the screening variables in each round of sampling, calculating and evaluating the root mean square error of the analysis and prediction model;
[0019] The screening variable corresponding to the minimum root mean square error is used as the feature set to be tested.
[0020] In a third possible embodiment of the first aspect, determining the number of wavelength information retained in each iterative sampling includes:
[0021] Determining exponential decay function parameters based on preset constraints;
[0022] Establishing an exponential decay function for the number of wavelength information retained in each iterative sampling based on the determined exponential decay function parameters;
[0023] The number of wavelength information retained in each iterative sampling is calculated according to the exponential decay function.
[0024] In a fourth possible embodiment of the first aspect, optimizing the parameters of the support vector machine model by using a particle swarm optimization algorithm includes:
[0025] Randomly generate a group of initialized particle swarms, where the position of each particle represents the value of the parameter of the support vector machine model, and the speed of each particle represents the direction and speed of the search;
[0026] Determining a current evolutionary state of the particle swarm according to a distribution state of the particle swarm;
[0027] Based on the determined evolutionary state and inertia factor, the search space is continuously iterated to find the optimal parameters of the support vector machine model.
[0028] In a fifth possible embodiment of the first aspect, determining a current evolutionary state of the particle swarm according to a distribution state of the particle swarm includes:
[0029] Calculating the current evolution factor of the particle group according to the average distance between each of the particles;
[0030] The evolutionary state corresponding to the current evolutionary factor of the particle swarm is determined based on the membership function and preset rules.
[0031] In a sixth possible embodiment of the first aspect, the feature set to be tested is input into the pre-trained support vector machine model to determine the year of the tangerine peel to be tested;
[0032] Mapping the feature set to be tested to a high-dimensional feature space through a kernel function to obtain a linearly separable feature set to be tested;
[0033] The linearly separable feature set to be tested is classified based on the decision boundary of the tangerine peel year classification to identify the year of the tangerine peel to be tested.
[0034] In a second aspect, an embodiment of the present application provides a tangerine peel year identification system, comprising:
[0035] A data acquisition module is used to obtain the spectral data of the tangerine peel to be tested and the raw spectral data with labels;
[0036] a data processing module, configured to preprocess the spectral data and the raw spectral data, and perform feature screening on the preprocessed spectral data to obtain a feature set to be measured;
[0037] Model optimization module, used to optimize the parameters of the support vector machine model through the particle swarm optimization algorithm;
[0038] A model training module is used to train the support vector machine model after parameter optimization based on the preprocessed raw spectral data to obtain a pretrained support vector machine model;
[0039] The classification and recognition module is used to input the feature set to be tested into the pre-trained support vector machine model to determine the year of the tangerine peel to be tested.
[0040] In a third aspect, an embodiment of the present application provides a terminal device, which includes a processor and a memory, wherein the memory stores a computer program, and the processor is used to execute the computer program to implement the above-mentioned method for identifying the age of tangerine peel.
[0041] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed on a processor, implements the above-mentioned method for identifying the age of tangerine peel.
[0042] The embodiments of the present application have the following beneficial effects:
[0043] A method for identifying the age of dried tangerine peel of the present embodiment includes: obtaining spectral data of the dried tangerine peel to be tested and labeled original spectral data, preprocessing the spectral data and the original spectral data; performing feature screening on the preprocessed spectral data to obtain a feature set to be tested; optimizing the parameters of a support vector machine model using a particle swarm optimization algorithm; training the parameter-optimized support vector machine model based on the preprocessed original spectral data to obtain a pretrained support vector machine model; and inputting the feature set to be tested into the pretrained support vector machine model to determine the age of the dried tangerine peel to be tested. The present application provides a real-time method for distinguishing the age of dried tangerine peel with low computational complexity and ease of real-time calculation. Furthermore, obtaining the spectral data of the dried tangerine peel to be tested is non-destructive and easy to carry, making it easy to operate and reducing the cost of dried tangerine peel identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0045] Figure 1 A schematic diagram of a first process of the method for identifying the age of dried tangerine peel according to an embodiment of the present application is shown;
[0046] Figure 2 A second flow chart of the method for identifying the age of dried tangerine peel according to an embodiment of the present application is shown;
[0047] Figure 3 A third flow chart of the method for identifying the age of dried tangerine peel according to an embodiment of the present application is shown;
[0048] Figure 4 A schematic diagram of a confusion matrix for classification of dried tangerine peel by year using near infrared spectroscopy and support vector machine is shown in an embodiment of the present application;
[0049] Figure 5 A schematic diagram of the confusion matrix for classification of tangerine peel by year using near infrared spectroscopy and particle swarm optimized support vector machine is shown in an embodiment of the present application;
[0050] Figure 6 A schematic diagram of the structure of the tangerine peel year identification system according to an embodiment of the present application is shown.
[0051] Description of main component symbols:
[0052] 200-Chenpi year identification system; 210-Data acquisition module; 220-Data processing module; 230-Model optimization module; 240-Model training module; 250-Classification and recognition module. DETAILED DESCRIPTION
[0053] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments.
[0054] The components of the embodiments of the present application generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but rather merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.
[0055] Hereinafter, the terms "including", "having" and their cognates used in various embodiments of the present application are intended only to indicate specific features, numbers, steps, operations, elements, components or combinations of the aforementioned items, and should not be understood as excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the aforementioned items or adding the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the aforementioned items. In addition, the terms "first", "second", "third" and the like are only used to distinguish descriptions and should not be understood as indicating or implying relative importance.
[0056] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art to which the various embodiments of the present application belong. The terms (such as those defined in generally used dictionaries) will be interpreted as having the same meaning as in the context of the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning unless clearly defined in the various embodiments of the present application.
[0057] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.
[0058] Near-infrared spectroscopy (NIRS) is a spectral analysis technique based on molecular vibrational absorption. Its wavelength range lies between the visible and mid-infrared spectra, typically defined as 780 to 2500 nanometers. NIRS is widely used in various fields, with advantages including: high analysis speed: a sample analysis typically takes approximately one minute, significantly improving analytical efficiency; no sample pretreatment required: NIR spectroscopy requires no chemical treatment of the sample, simplifying the analysis process and saving time and cost; no consumables and environmental friendliness: the analysis process does not consume any other materials or damage the sample, making it a green and environmentally friendly detection method; simultaneous multi-component detection: a single test can simultaneously detect multiple components or indicators, making it suitable for the direct and rapid analysis of complex samples; and a wide range of applications: NIR spectroscopy is suitable for the analysis of organic compounds and biomolecules, playing an important role in industries such as pharmaceuticals, agriculture, food, and environmental protection. Due to its unique advantages, NIR spectrometers play an important role in various fields, and with the continuous advancement of technology, their application scope and accuracy will continue to increase.
[0059] In response to the problem that it is difficult to directly identify the age of tangerine peel by characteristics such as color and smell, this application proposes a tangerine peel age identification method, system, terminal device and storage medium, which uses near-infrared spectroscopy and particle swarm optimized support vector machine model to identify the tangerine peel age, thereby improving the accuracy of tangerine peel age identification.
[0060] The method for identifying the age of dried tangerine peel will be described below with reference to some specific embodiments.
[0061] Figure 1 A flow chart of a method for identifying the age of dried tangerine peel according to an embodiment of the present application is shown. Exemplarily, the method for identifying the age of dried tangerine peel includes the following steps:
[0062] S110, obtaining the spectral data of the tangerine peel to be tested and the original spectral data with labels, and preprocessing the spectral data and the original spectral data.
[0063] In one embodiment, the application gathers and smashes the spectral data of the dried tangerine peel to be measured that is tiled in the test container after processing.In the present embodiment, the dried tangerine peel to be measured of different years is tiled in the test container after smashing processing, regulates the near-infrared probe height so that absorbance value size is suitable, before formal test, gathers the data of near-infrared blackboard and whiteboard, for correcting instrument system error.After the test, use near-infrared spectrometer to gather the spectral data of dried tangerine peel to be measured, record the variation curve of absorbance value with wavelength point, obtain the spectral data of dried tangerine peel to be measured.
[0064] For example, the labeled raw spectral data in this application includes spectral data of dried tangerine peel from different years and their corresponding category labels (such as "1 year," "5 years," "10 years," etc.). In this application, dried tangerine peel from different years has different chemical compositions and structural characteristics, which are manifested in the spectral data as specific absorption peak positions, intensities, and shapes. This application collects labeled raw spectral data for subsequent training of a support vector machine model.
[0065] In another embodiment, the present application sequentially performs normal variable transformation, wavelet smoothing, and normalization on the collected spectral data and the original spectral data. The normal variable transformation converts the spectral data into relative absorption intensities, thereby reducing the influence of sample surface characteristics on the spectral signal. Wavelet smoothing is used to smooth and compress the local band data to achieve the purpose of denoising. Normalization is used to scale the spectral data to a uniform range to eliminate dimensional differences.
[0066] In one embodiment, the expression for the normal variable transformation process is:
[0067]
[0068] Among them, X i,SNV represents the spectral data after normal variable transformation, x i,k represents the absorbance value of the i-th spectral data at the k-th wavelength, x i Represents the average absorbance value of all wavelength points corresponding to the i-th spectral data, k = 1, 2, ..., m, k represents the wavelength point; i = 1, 2, ..., n, i represents the spectral data.
[0069] In another embodiment, the present application uses Daubechies wavelet for wavelet decomposition, with the decomposition times being 5, and then reconstructs the decomposed wavelet coefficients again to restore the spectral data. The expression for spectral data normalization is:
[0070]
[0071] Where Z represents the standard score, that is, the normalized spectral data, y represents the absorbance value corresponding to a wavelength point of the original spectral data, μ represents the average absorbance value corresponding to all wavelength points in the spectral data, and σ represents the standard deviation of the absorbance values corresponding to all wavelength points in the spectrum.
[0072] S120 , performing feature screening on the preprocessed spectral data to obtain a feature set to be measured.
[0073] In one embodiment, if Figure 2As shown, this application uses a competitive adaptive reweighted sampling method with variable selection to select spectral feature wavelengths from pre-processed near-infrared spectral data, and selects the spectral wavelength data that is most helpful for year classification performance. The specific steps include the following:
[0074] S121, using the preprocessed original spectral data as a data set, and dividing the data set into a training set.
[0075] S122, modeling and analyzing the training set by partial least squares method to determine the weight of each wavelength information in the spectral data for the identification of tangerine peel year.
[0076] Exemplarily, the present application randomly divides the preprocessed raw spectral data. The division ratio can be selected according to the actual situation. For example, 70% of the data is used to train the model, and the remaining 30% is used to verify the performance of the model. The present application uses partial least squares (PLS) modeling to analyze the relationship between the spectral data (wavelength information) and the target variable (i.e., the year of tangerine peel). The percentage of the absolute value of the regression coefficient is used as the importance of the wavelength information or the explanatory power of the target wavelength information.
[0077] In one embodiment, the spectral data contains a large amount of wavelength information, i.e., a large number of wavelength points and corresponding absorbance values. PLS is used to extract a small amount of wavelength information from the spectral data, which can reflect the relationship between the wavelength information and the tangerine peel year to the greatest extent. During modeling, the partial least squares method utilizes the potential components between the independent variable (wavelength information) and the dependent variable (tangerine peel year) to maximize the covariance between the independent variable and the dependent variable for model training and verification. Among them, the potential component is a weighted combination of multiple wavelength information. For example, if certain wavelength regions (such as 1400nm and 1900nm) are particularly important for distinguishing tangerine peel years, then the potential component may focus more on the information of these wavelengths, and the weight corresponding to the wavelength information of this wavelength region is larger.
[0078] S123: Determine the number of wavelength information items retained in each iterative sampling.
[0079] In one embodiment, the present application sets an exponential decay function to calculate the number of wavelength information retained in each iterative sampling. The expression of the exponential decay function is:
[0080] r q =ae -bq
[0081] r q It represents the number of wavelength information retained in the qth iterative sampling. a and b are both parameters of the exponential decay function, where a represents the amplitude of the exponential decay function and b represents the decay rate parameter, which controls the speed at which the function decreases. When b>0, the function shows a decreasing trend.
[0082] This application determines the parameters of the exponential decay function based on preset constraint conditions. Among them, the preset constraint condition is: r1 = P, r1 represents the number of wavelength information retained in the first iterative sampling, and r N represents the number of wavelength information retained in the Nth iterative sampling, and P represents the total number of wavelength information.
[0083] The parameters of the exponential decay function can be obtained as:
[0084]
[0085] This application establishes an exponential decay function for the number of wavelength information retained in each iterative sampling based on the determined parameters of the exponential decay function, and calculates the number of wavelength information retained in each iterative sampling according to the exponential decay function. Among them, when q < c, it is the fast selection stage, and a large number of irrelevant wavelength information is eliminated. When c < i < N, it is the refined selection stage, and the number of wavelength information deleted each time is less. c represents the classification point between the fast selection stage and the refined selection stage.
[0086] S124, based on the determined number of wavelength information and the weight of each wavelength information, perform multiple rounds of sampling on the wavelength information in the spectral data to obtain the screening variables for each round of sampling.
[0087] In this embodiment, each time sampling is performed, this application filters out the wavelength information with smaller weights according to the weights of the wavelength information, retains the wavelength information with larger weights, and the finally retained number of wavelength information is the calculated number of retained wavelength information.
[0088] S125, based on the analysis and prediction model of the screening variables for each round of sampling, calculate the root mean square error of the evaluation analysis and prediction model.
[0089] S126, use the screening variable corresponding to the minimum root mean square error as the set of待测特征集合.
[0090] In this embodiment, the analysis and prediction model based on the screening variables refers to, after feature selection, using the screened variables to construct a mathematical or statistical model for predicting or classifying the target variable. The specific form of this model can be selected according to the application scenario and data characteristics. For example, it includes but is not limited to regression models, classification models, etc. This application calculates the root mean square error based on the analysis and prediction model of the screening variables, verifies the effectiveness of the analysis and prediction model, calculates the root mean square error of each iteration by setting the number of loop iterations, and determines the best screening variable based on the minimum root mean square error to obtain the set of待测特征集合.
[0091] S130, optimize the parameters of the support vector machine model through the particle swarm optimization algorithm.
[0092] In this embodiment, the support vector machine model parameters include a penalty parameter and a kernel function parameter. In this application, the penalty parameter is used to control the degree of fit of the support vector machine model to the training data. A larger penalty parameter will make the model more concerned with reducing the training error, which may cause overfitting; a smaller penalty parameter value will make the model more inclined to simplify the decision boundary, which may lead to underfitting. The kernel function parameters define the complexity of the feature space. For example, the parameters of the Gaussian kernel (RBF kernel) determine the similarity range between data points. A larger value will make the model pay more attention to local features, which may cause overfitting; a smaller value will make the model pay more attention to global features, which may lead to underfitting. Reasonable selection of penalty parameters and kernel function parameters is the key to improving the performance of support vector machine model parameters.
[0093] In one embodiment, if Figure 3 As shown, this application uses the particle swarm optimization algorithm to optimize the support vector machine model parameters, which specifically includes the following steps:
[0094] S131 , randomly generating a group of initialized particle swarms, wherein the position of each particle represents the value of a parameter of the support vector machine model, and the speed of each particle represents the direction and speed of the search.
[0095] S132, determining the current evolution state of the particle swarm according to the distribution state of the particle swarm;
[0096] In one embodiment, the present application calculates the current evolution factor of the particle swarm based on the average distance between each particle; and determines the evolution state corresponding to the current evolution factor of the particle swarm based on the membership function and preset rules.
[0097] In one embodiment, the present application calculates the average distance d of each particle l relative to other particles. l , select d l The best value is d g , d g Indicates the average distance from the particle at the global optimal position to other particles, and calculates d l The maximum distance d max , minimum distance d min , calculate the evolution factor f. The calculation formula of average distance and evolution factor f is:
[0098]
[0099] Where M is the number of particles in the particle swarm, l = 1, 2, 3…M, j = 1, 2, 3…M, D is the maximum dimension of the search space, d represents the dimension value of the search space, d = 1, 2, 3,…, D, Indicates the position of the lth particle when the dimension value of the search space is d, Indicates the position of the jth particle when the dimension of the search space is d.
[0100] This application uses the membership function to determine the evolutionary state at the moment of evolution based on the calculated evolution factor f. The four evolutionary states are T1: exploration, T2: discovery, T3: convergence, and T4: exit. The expression of the membership function is:
[0101] T1:
[0102]
[0103] T2:
[0104]
[0105] T3:
[0106]
[0107] T4:
[0108]
[0109] The membership function ut1(f) represents the membership of the evolution factor f to state T1. The membership range is [0,1]. When ut1(f) = 1, it means that f completely belongs to state T1. When ut1(f) = 0, it means that f does not belong to state T1 at all. ut2(f) represents the membership of the evolution factor f to state T2. ut3(f) represents the membership of the evolution factor f to state T3. ut4(f) represents the membership of the evolution factor f to state T4.
[0110] Without considering the preset rules, this application determines the evolutionary state corresponding to the current evolutionary factor of the particle swarm based on the membership of the evolutionary factor f in the four evolutionary states. For example, if the membership of the evolutionary factor f in state T1 is the largest, the evolutionary state corresponding to the current evolutionary factor of the particle swarm is T1.
[0111] Taking into account the preset rules, the preset rules of this application are: in the above-mentioned selection of T1 and T2, if the previous state is T4, the state of the evolution factor f belongs to T1; if the previous state is T1, since the state cannot be switched excessively and the stability of the division is maintained, the evolution factor f still belongs to the T1 state. If the previous state is T2 or T3, the evolution factor f is divided into the T2 state according to the degree of membership.
[0112] S133, based on the determined evolutionary state and inertia factor, continuously iterates in the search space to find the optimal parameters of the support vector machine model.
[0113] In this embodiment, the search space refers to the set of all possible solutions. In this application, for the support vector machine (SVM) model parameter optimization problem, the search space specifically refers to all possible value ranges of the SVM model parameters. The determination of the evolutionary state is the core of the dynamic adjustment optimization strategy. In this application, for the support vector machine parameter optimization, the exploration stage (T1) is used to search the solution space extensively and try to find potential high-quality parameter combinations; the discovery stage (T2) is used to focus on certain possible high-quality solution areas and further refine the search; the convergence stage (T3) is used to refine the search and improve parameter accuracy when approaching the optimal solution; the exit stage (T4) is used to reintroduce randomness or disturbance when falling into a local optimum to avoid premature convergence. The inertia factor is used to balance the global search capability and the local search capability. The change in its value directly affects the speed update of the particle. The inertia factor is very adaptive in the exploration stage and decreases in the development stage. The relationship between the inertia factor and the evolution factor is as follows:
[0114]
[0115] w(f) represents the value of the inertia factor when the evolution factor is f. In this application, w is initialized to 0.7. In the jump state and exploration state, larger f and larger w are more conducive to global search. On the contrary, in the development state and convergence state, f is smaller and w will be reduced, which is more conducive to local search.
[0116] S140 , training the support vector machine model after parameter optimization based on the preprocessed original spectral data to obtain a pre-trained support vector machine model.
[0117] In one embodiment, the present application uses pre-treated raw spectral data as a training data set to train a support vector machine model. The core goal of training the support vector machine model is to find the optimal decision boundary (hyperplane) for the classification of dried tangerine peel by year. The support vector machine model finds the optimal decision boundary for the classification of dried tangerine peel by year by learning the raw spectral data, minimizes the classification error on the training data set, and maximizes the interval so that the spectral data of different years are clearly separated. During training, samples from a certain year are sequentially classified into one category, and the remaining samples are classified into another category. In this way, samples from M years construct M support vector machine models, realizing the identification of dried tangerine peel from multiple years.
[0118] S150: Input the feature set to be tested into a pre-trained support vector machine model to determine the year of the tangerine peel to be tested.
[0119] In one embodiment, the present application maps the feature set to be tested to a high-dimensional feature space through a kernel function to obtain a linearly separable feature set to be tested, and classifies the linearly separable feature set to be tested based on the decision boundary of the tangerine peel year classification to identify the year of the tangerine peel to be tested.
[0120] In the present embodiment, support vector machine model converts the nonlinear separable problem in low-dimensional feature space into the linear separable problem in high-dimensional feature space by kernel function, and the tested feature set linear separability after mapping, in high-dimensional feature space, the dried tangerine peel samples of different years become easier to distinguish.For example, originally the sample that can't be separated by a straight line in two-dimensional space can be clearly separated by a hyperplane in high-dimensional space.When classification, the wavelength information representing different years in the tested feature set is correctly separated by the decision boundary of year classification, and support vector machine model outputs the year of dried tangerine peel to be measured simultaneously.
[0121] In this embodiment, the support vector machine model is used for the classification task of dried tangerine peel years. The penalty parameters and kernel function parameters of the SVM model are optimized by the particle swarm optimization algorithm to ensure that the model can distinguish dried tangerine peels of different years with optimal performance. Figure 4 The confusion matrix of tangerine peel vintage classification using near-infrared spectroscopy and support vector machine was given. The average classification accuracy of the five vintages was 79.3%, and the highest classification accuracy was 94.1% for the sixth vintage. Figure 5 The confusion matrix of the classification of tangerine peel year using near infrared spectroscopy and particle swarm optimized support vector machine is given. The average classification accuracy of the five years is 86.8%, and the highest classification accuracy is 97.8% for the first year. The average classification accuracy of the particle swarm optimized support vector machine is 7.5% higher than that of the unoptimized one. From the data in the figure, it can be seen that the particle swarm optimized support vector machine proposed in this application has a better recognition effect for the classification of tangerine peel year. Among them, Figure 4 、 Figure 5 In the , TPR (True Positive Rate, Recall Rate) can be understood as how many of all positive classes are predicted as positive classes (positive class prediction is correct), and FNR (False Negative Rate) can be understood as how many of all positive classes are predicted as negative classes (negative class prediction is wrong).
[0122] Figure 6 A schematic diagram of a system 200 for identifying the age of dried tangerine peel according to an embodiment of the present application is shown. Exemplarily, the system 200 for identifying the age of dried tangerine peel includes:
[0123] The data acquisition module 210 is used to acquire the spectral data of the tangerine peel to be tested and the raw spectral data with labels.
[0124] The data processing module 220 is used to preprocess the spectral data and the original spectral data, and input the preprocessed spectral data into the feature screening model to perform feature screening to obtain a feature set to be tested.
[0125] The model optimization module 230 is used to optimize the parameters of the support vector machine model by using a particle swarm optimization algorithm.
[0126] The model training module 240 is used to train the support vector machine model after parameter optimization based on the preprocessed original spectral data to obtain a pre-trained support vector machine model.
[0127] The classification and recognition module 250 is used to input the feature set to be tested into the pre-trained support vector machine model to determine the year of the tangerine peel to be tested.
[0128] It can be understood that the system of this embodiment corresponds to the method for identifying the age of tangerine peel in the above embodiment, and the optional items in the above embodiment are also applicable to this embodiment, so they will not be repeated here.
[0129] The present application also provides a terminal device. Exemplarily, the terminal device includes a processor and a memory, wherein the memory stores a computer program, and the processor runs the computer program to enable the terminal device to execute the above-mentioned tangerine peel age identification method or the functions of each module in the above-mentioned tangerine peel age identification system.
[0130] Among them, the processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including a central processing unit (CPU), a graphics processing unit (GPU) and a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or at least one of other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can implement or execute the disclosed methods, steps and logic block diagrams in the embodiments of the present application.
[0131] The memory may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The memory is used to store a computer program, and the processor may execute the computer program accordingly after receiving an execution instruction.
[0132] The present application also provides a computer-readable storage medium for storing the computer program used in the terminal device. For example, the computer-readable storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0133] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in an alternative implementation, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the structure diagram and / or flowchart, and the combination of boxes in the structure diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0134] In addition, the functional modules or units in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0135] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a smart phone, personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.
[0136] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A method for identifying the age of dried tangerine peel, characterized in that: include: Acquire spectral data of the tangerine peel to be tested and raw spectral data with labels, and preprocess the spectral data and the raw spectral data; Performing feature screening on the pre-processed spectral data to obtain a feature set to be measured; The parameters of the support vector machine model are optimized using the particle swarm optimization algorithm; Training the support vector machine model after parameter optimization based on the preprocessed original spectral data to obtain a pre-trained support vector machine model; The feature set to be tested is input into the pre-trained support vector machine model to determine the year of the tangerine peel to be tested.
2. The method for identifying the age of dried tangerine peel according to claim 1, wherein The step of obtaining the spectral data of the tangerine peel to be tested and the original spectral data with labels, and preprocessing the spectral data and the original spectral data, comprises: Collecting spectral data of the tangerine peel to be tested after being crushed and spread into a test container; The collected spectral data and the original spectral data are sequentially subjected to normal variable transformation processing, wavelet smoothing processing and normalization processing.
3. The method for identifying the age of dried tangerine peel according to claim 1, wherein The pre-processed spectral data is subjected to feature screening to obtain a feature set to be measured, including: The pre-processed raw spectral data is used as a data set, and the data set is divided into a training set; The training set is modeled and analyzed by partial least squares method to determine the weight of each wavelength information in the spectral data for the identification of the year of tangerine peel; Determine the number of wavelength information retained in each iterative sampling; performing multiple rounds of sampling on the wavelength information in the spectral data based on the determined number of wavelength information and the weight of each wavelength information to obtain a screening variable for each round of sampling; Based on the analysis and prediction model of the screening variables in each round of sampling, calculating and evaluating the root mean square error of the analysis and prediction model; The screening variable corresponding to the minimum root mean square error is used as the feature set to be tested.
4. The method for identifying the age of dried tangerine peel according to claim 3, wherein: Determining the number of wavelength information retained in each iterative sampling includes: Determining exponential decay function parameters based on preset constraints; Establishing an exponential decay function for the number of wavelength information retained in each iterative sampling based on the determined exponential decay function parameters; The number of wavelength information retained in each iterative sampling is calculated according to the exponential decay function.
5. The method for identifying the age of dried tangerine peel according to claim 1, wherein: Optimizing the parameters of the support vector machine model using the particle swarm optimization algorithm includes: Randomly generate a group of initialized particle swarms, where the position of each particle represents the value of the parameter of the support vector machine model, and the speed of each particle represents the direction and speed of the search; Determining a current evolutionary state of the particle swarm according to a distribution state of the particle swarm; Based on the determined evolutionary state and inertia factor, the search space is continuously iterated to find the optimal parameters of the support vector machine model.
6. The method for identifying the age of dried tangerine peel according to claim 5, wherein: Determining the current evolutionary state of the particle swarm according to the distribution state of the particle swarm includes: Calculating the current evolution factor of the particle group according to the average distance between each of the particles; The evolutionary state corresponding to the current evolutionary factor of the particle swarm is determined based on the membership function and preset rules.
7. The method for identifying the age of dried tangerine peel according to claim 1, wherein: The feature set to be tested is input into the pre-trained support vector machine model to determine the year of the tangerine peel to be tested; Mapping the feature set to be tested to a high-dimensional feature space through a kernel function to obtain a linearly separable feature set to be tested; The linearly separable feature set to be tested is classified based on the decision boundary of the tangerine peel year classification to identify the year of the tangerine peel to be tested.
8. A system for identifying the age of dried tangerine peel, characterized in that: include: A data acquisition module is used to obtain the spectral data of the tangerine peel to be tested and the raw spectral data with labels; a data processing module, configured to preprocess the spectral data and the raw spectral data, and perform feature screening on the preprocessed spectral data to obtain a feature set to be measured; Model optimization module, used to optimize the parameters of the support vector machine model through the particle swarm optimization algorithm; A model training module is used to train the support vector machine model after parameter optimization based on the preprocessed raw spectral data to obtain a pretrained support vector machine model; The classification and recognition module is used to input the feature set to be tested into the pre-trained support vector machine model to determine the year of the tangerine peel to be tested.
9. A terminal device, characterized in that: The terminal device includes a processor and a memory, the memory stores a computer program, and the processor is used to execute the computer program to implement the tangerine peel year identification method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that It stores a computer program, which, when executed on a processor, implements the method for identifying the age of tangerine peel according to any one of claims 1 to 7.