Method for detecting content of ergosterol in lentinus edodes based on near infrared spectrum and application
By employing near-infrared spectroscopy and the CPO-LSSVM model, combined with characteristic wavelength screening and parameter optimization, the complexity and noise issues in ergosterol detection in shiitake mushrooms have been resolved, enabling rapid and accurate detection of ergosterol content and supporting efficient quality control and raw material screening for edible fungi enterprises.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies for detecting ergosterol content in shiitake mushrooms are cumbersome, time-consuming, labor-intensive, and cannot be performed on a large scale. Near-infrared spectral data processing is complex and subject to significant noise. Improper model establishment leads to large quantitative analysis errors.
Near-infrared spectroscopy is employed, combined with Mahalanobis distance to remove outliers, spectral preprocessing, feature wavelength extraction, and the CPO-LSSVM model to establish a fast and accurate detection model. The VCPA-GA algorithm is used to screen feature wavelengths, and the LSSVM model parameters are optimized through CPO to improve the model's adaptability and prediction accuracy.
It enables rapid and accurate detection of ergosterol content in shiitake mushrooms, reduces testing costs, provides efficient quality evaluation and raw material screening support, and is suitable for quality control and product grading in edible fungi processing enterprises.
Smart Images

Figure CN121740773A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of ergosterol content detection technology, and in particular relates to a method and application for detecting ergosterol content in shiitake mushrooms based on near-infrared spectroscopy. Background Technology
[0002] Domestic and international studies have found that edible fungi such as shiitake mushrooms are rich in ergosterol, an important sterol compound with various biological activities. Ergosterol is a precursor to vitamin D2 and can be converted into vitamin D2 under ultraviolet irradiation, playing an important role in regulating calcium and phosphorus metabolism and promoting bone development. At the same time, ergosterol also possesses various pharmacological activities such as antioxidant, antibacterial, anti-inflammatory, and immunomodulatory effects, making it highly valuable for food health and pharmaceutical development. Currently, domestic and international methods for detecting ergosterol in edible fungi and their products include ultraviolet spectrophotometry, thin-layer chromatography, liquid chromatography, ultra-high performance liquid chromatography, and gas chromatography-mass spectrometry. However, these quantitative analysis methods are cumbersome, time-consuming, and labor-intensive, and cannot be performed on a large scale. There is an urgent need to develop a rapid and simple analytical method to determine the ergosterol content in shiitake mushrooms.
[0003] Near-infrared spectroscopy has advantages such as speed, accuracy, and no pollution. Therefore, establishing a rapid and accurate near-infrared prediction model for detecting ergosterol in shiitake mushrooms can shorten the time and reduce costs. This can not only provide a reference for accurately evaluating the quality of shiitake mushrooms, but also provide enterprises with efficient and convenient rapid screening technology support in raw material procurement, quality control, and product grading.
[0004] However, near-infrared spectroscopy also faces several challenges in quantitative analysis, including: 1. Near-infrared spectral data typically contains significant noise, which may originate from sensors, environmental conditions, or other factors during the acquisition process. The presence of noise can challenge the accuracy and repeatability of the data, requiring complex data processing and filtering techniques to reduce its impact; 2. Near-infrared spectral data is usually high-dimensional, containing hundreds or even thousands of spectral wavelengths. Processing and analyzing such large amounts of data requires substantial computational resources and complex algorithms; 3. While near-infrared spectral data provides detailed spectral information, using this information for quantitative analysis requires building complex models. Furthermore, inappropriate model building can lead to significant errors in the quantitative analysis. Summary of the Invention
[0005] The purpose of this invention is to provide a method and application for detecting ergosterol content in shiitake mushrooms based on near-infrared spectroscopy, so as to achieve rapid and efficient detection of ergosterol content in shiitake mushrooms, accurately evaluate the quality of shiitake mushrooms, and provide enterprises with efficient and convenient rapid screening technology support in raw material procurement, quality control and product grading.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: This invention provides a method for detecting ergosterol content in shiitake mushrooms based on near-infrared spectroscopy, the method comprising the following steps: S1. Use a near-infrared spectrometer to scan the absorption spectrum data of the shiitake mushroom sample to obtain full-wavelength spectral data; S2. Use Mahalanobis distance to remove abnormal shiitake mushroom samples, obtain a sample set, and divide the sample set. S3. Preprocess the spectral data after removing abnormal samples in step S2. S4. Extract characteristic wavelengths from the preprocessed spectral data obtained in step S3. S5. The content of ergosterol in shiitake mushroom samples was determined by chemical assay. S6. Establish a detection model using the spectral data values of the full wavelength and characteristic wavelength and the ergosterol content determination value obtained in step S5, and evaluate the detection results of the detection model to determine the performance of the detection model; the optimal detection model is the least squares support vector machine CPO-LSSVM based on the porcupine optimization algorithm. The parameter settings for the CPO-optimized LSSVM model are: the number of particle swarms N=6, the maximum number of iterations Gmax=10, the lower bound of the decision variable is [0.1, 0.1], and the upper bound of the decision variable is [900, 900].
[0007] Furthermore, the KS algorithm is used to divide the sample set into a training set and a prediction set.
[0008] Furthermore, the spectral preprocessing method includes any one of SG, MSC, NORM, SNV, SG-MSC, SG-SNV, and SG-NORM.
[0009] Furthermore, the characteristic wavelength extraction process includes the following steps: S11. Use VCPA to rapidly reduce the variable space for near-infrared spectral wavelengths using EDF looping, and roughly screen out important variables. S12. Use GA to perform a fine search on the coarse screening results of VCPA to find the optimal combination of variables.
[0010] Furthermore, the range of the characteristic wavelengths is 900 nm to 1300 nm, 1450 nm to 1650 nm, 1720 nm to 1740 nm, 1810 nm to 2000 nm, and 2300 nm to 2400 nm.
[0011] Furthermore, the process of determining the ergosterol content in shiitake mushroom samples by chemical determination is as follows: Shiitake mushroom powder is mixed with a low alcohol solution containing an alkaline catalyst and subjected to ultrasonic-assisted extraction to obtain an extract; the extract is subjected to solid-liquid separation, and the supernatant is filtered through a microporous membrane and then diluted to a final volume to obtain a sample solution; the absorbance of the sample solution is measured at the ultraviolet characteristic absorption wavelength of ergosterol to quantify the ergosterol content.
[0012] Further, in step S6, the parameter calculation formula for evaluating the detection model is as follows:
[0013]
[0014]
[0015]
[0016]
[0017] in It is the predicted value of ergosterol content in the a-th shiitake mushroom sample; The true value of ergosterol content in the a-th shiitake mushroom sample. This represents the average ergosterol content of the samples in the training set. Predict the average ergosterol content of the concentrated samples, n c n is the number of samples in the training set. p This represents the number of samples in the prediction set.
[0018] Furthermore, in step S6, to evaluate the sensitivity of the established optimal detection model, its limit of detection and limit of quantitation are calculated respectively, taking into account the error and sensitivity of the calibration model. The specific calculation formula is as follows:
[0019]
[0020] Wherein the limit of detection (LOD), the limit of quantitation (LOQ), σ is the standard deviation of the prediction residuals of the calibration set, S is the slope of the calibration curve, and 3.3 and 10 are statistical confidence factors.
[0021] A second objective of this invention is to provide a system for detecting ergosterol content in shiitake mushrooms, the system being based on the aforementioned method and comprising at least the following components built within the system: The data acquisition module is configured to acquire absorption spectral data of shiitake mushroom samples scanned by a near-infrared spectrometer; the near-infrared spectral wavelength range is 850~2500 nm. The dataset building module is configured to use Mahalanobis distance to remove abnormal shiitake mushroom samples, obtain a sample set, and divide the sample set. The preprocessing module is configured to preprocess the spectral data after removing abnormal samples in step S2; The feature wavelength extraction module is configured to extract feature wavelengths from the obtained preprocessed spectral data; The detection model construction module is configured to establish a detection model by using the spectral data values of the full wavelength and the characteristic wavelength and the ergosterol content determination value obtained by the chemical determination method in step S5, respectively, and to evaluate the detection results of the detection model and determine the performance of the detection model; the optimal detection model is the least squares support vector machine CPO-LSSVM based on the porcupine optimization algorithm.
[0022] A third objective of the present invention is to provide a computer-readable storage medium storing a program that can be executed by one or more processors to implement the above-described method for detecting ergosterol content in shiitake mushrooms.
[0023] The abbreviations used in this invention are explained as follows: KS: Kennard-Stone algorithm; SG: Savitzky-Gore filter smoothing algorithm; SNV: Standard Normal Transform Algorithm; MSC: Multiplicative Scattering Correction Algorithm; NORM: Normalization algorithm; VCPA-GA: Variable Combination Population Analysis - Genetic Algorithm; EDF: Empirical Distribution Function; PLSR: Partial Least Squares Regression Algorithm; SVR: Support Vector Regression algorithm; CPO-LSSVM is a hybrid prediction or classification model that combines the Crowned Porcupine Optimization Algorithm with Least Squares Support Vector Machine. It utilizes the powerful global optimization capability of CPO to automatically select the most critical hyperparameters in LSSVM, thereby improving the model's generalization ability and prediction accuracy. RMSE: Root Mean Square Error; RMSEC: Root Mean Square Error of Training Set; RMSEP: Root Mean Square Error of Prediction Set; RPD: Residual Prediction Bias; LOD: Limit of Detection; LOQ: Limit of Quantification.
[0024] Compared with the prior art, the beneficial effects of the technical solution provided by the present invention are as follows: (1) This invention provides a method for detecting ergosterol content in shiitake mushrooms. Near-infrared spectroscopy is used to rapidly detect ergosterol content in shiitake mushrooms. The spectral data of the sample to be tested is input into the optimal detection model to obtain the predicted value of ergosterol content, thereby realizing rapid detection of ergosterol content. This invention uses a combination of SG smoothing and normalization algorithms to preprocess the spectrum, which can effectively remove noise generated when the sensor acquires spectral data and perform baseline and scattering correction. The VCPA-GA algorithm is used to perform wavelength screening and extraction of the original spectrum, which can effectively extract the required important variables, efficiently remove redundant information, and improve the running speed of the model. The CPO-LSSVM model is used to predict the ergosterol content in shiitake mushrooms. This model has strong adaptability and high prediction accuracy.
[0025] (2) The method provided by the present invention makes up for the shortcomings of the existing detection technology of ergosterol content in shiitake mushrooms, realizes rapid and accurate detection of ergosterol content in shiitake mushrooms, provides a rapid detection technology for edible fungi processing enterprises to detect ergosterol content in shiitake mushroom raw materials, and provides a new way for intelligent monitoring of edible fungi quality and high-value processing. Attached Figure Description
[0026] Figure 1 This is a schematic diagram of the method for detecting ergosterol content in shiitake mushrooms provided by the present invention. Figure 2 This is a diagram illustrating the outlier removal process based on the Mahalanobis distance algorithm in an embodiment of the present invention. Figure 3 This is the near-infrared spectrum of shiitake mushrooms after outlier removal in an embodiment of the present invention; Figure 4 This is a diagram showing the characteristic wavelength results extracted by VCPA-GA based on CPO-LSSVM in an embodiment of the method of the present invention; Figure 5 This is an example of a near-infrared spectral prediction model for ergosterol content based on CPO-LSSVM in the embodiments of the present invention; Figure 6 A schematic diagram of the system structure provided by the present invention is shown. Figure 7 A block diagram of an electronic device suitable for implementing an information acquisition method according to an embodiment of the present invention is shown schematically. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] like Figure 1 The diagram shows a flowchart of the method for rapid detection of ergosterol content in shiitake mushrooms based on near-infrared spectroscopy according to the present invention, which includes the following steps: (1) Near-infrared (850~2500 nm) spectral data of shiitake mushroom samples were collected using a near-infrared spectrometer; (2) Use Mahalanobis distance to remove abnormal shiitake mushroom samples; (3) Use the KS algorithm to partition the sample set; (4) Preprocess the spectral data to select the optimal preprocessing method; (5) Extract characteristic wavelengths from the spectral data processed by the optimal preprocessing method; (6) The content of ergosterol in each shiitake mushroom sample was determined by chemical assay. (7) Establish a detection model for ergosterol content based on full wavelength and characteristic wavelength, evaluate the model detection results, and determine the model performance; Example 1 This invention provides a method for detecting the ergosterol content in shiitake mushrooms, specifically including the following steps: The test materials consisted of 124 shiitake mushroom samples. The shiitake mushroom samples were provided by the Wuhan Academy of Agricultural Sciences and some were commercially available. The Wuhan Academy of Agricultural Sciences provides the following shiitake mushroom varieties: Y11(W1), Y93(W1), Y107(W1), Y1493(W1), Y116(W1), Y49(W1), Y44(W2), Y4(W2), Y47(W2), Y39(W2), Y53(W2), Y48(W2), Y29(W2), Z67(C), Y45(W2), Z82(C), Y65(W2), Y27(W2), XG172(C), Z49(C), Z28C, Z50(C), Z51(C), Y70(C), Y37(C), Z6(C), Y69(C), Y67(C), Z42(C), Z3(C), Z46(C), Z6 The following are listed: 7(C), Y66(C), Z31(C), Z48(C), Y16(C), Y119(W1), Z23(C), Z77(C), Z72(C), Z60(C), Z30(C), Z2(C), Y68(C), Y40(W1), Y82, Y1332(W1), Y98(W1), Y30(W2), Y21(W1), Y84(W2), Y19(W2), Y99(W2), XG1346(C), Y58(C), Y55(C), Y3357, Y79(W2), Y28(W2), Y8(C). Commercially available shiitake mushroom varieties include: Shiitake 0912, Shiitake 9608, Shiitake 939, Shiitake 808, and Shiitake K5.
[0029] I. Acquisition of Near-Infrared Scanning Spectra The shiitake mushroom samples were dried to constant weight, pulverized using a pulverizer, thoroughly mixed, and sieved through a 40-mesh sieve. They were then placed in a glass desiccator and cooled to room temperature before being packed into small round sample cups. The samples were leveled and compacted to ensure uniformity, with a total weight of approximately 3 g. Near-infrared spectra of the shiitake mushroom samples in the 850-2500 nm range were collected using a FOSS DS2500F near-infrared spectrometer. The measurement method was diffuse reflectance, with a resolution of 2 nm and 7 scans. The instrument was set to repeat the sample loading 3 times. The average spectrum was used for modeling and verification.
[0030] II. Using Mahalanobis distance to remove abnormal shiitake mushroom samples The Mahalanobis distance of the empirical threshold method was used to screen shiitake mushroom samples, and two abnormal samples were identified and removed.
[0031] III. Partitioning the Sample Set Using the KS Algorithm To evaluate and validate the performance and generalization ability of the constructed model, and to avoid overfitting, the dataset is typically divided into a training set and a prediction set. The training set is used to train the model, and the prediction set is used to evaluate the performance of the trained model. This invention utilizes the KS algorithm to divide 122 shiitake mushroom spectral datasets into a training set and a prediction set in a 4:1 ratio. The training set contains 98 datasets, and the prediction set contains 24 datasets. The training set is used to build the model and for internal cross-validation, while the prediction set is used for external validation.
[0032] IV. Determination of Ergosterol Content in Shiitake Mushroom Samples Preparation of ergosterol standard solution: Accurately weigh 10 mg of ergosterol standard, dissolve in methanol, sonicate for 5 min, and dilute to 10 mL. Accurately pipette 5.00 mL of the standard stock solution, dilute with methanol, and dilute to 50 mL. Pipettes 0.00 mL, 0.5 mL, 2.50 mL, 5.00 mL, 10.00 mL, and 25.00 mL of the ergosterol standard intermediate solution, dilute with methanol, and dilute to 50 mL. The concentrations of this standard series are 0.00 μg / mL, 1.00 μg / mL, 5.00 μg / mL, 10.0 μg / mL, 20.0 μg / mL, and 50.0 μg / mL, respectively. Prepare immediately before use. Store the above solutions at -20 ℃ protected from light.
[0033] The preparation method of the shiitake mushroom sample solution was based on the literature published by Yang Haoyu et al. (Yang Haoyu, Cheng Yanfen, Wang Shengnan, et al. Extraction, characterization and antioxidant and cholesterol-lowering properties of ergosterol in shiitake mushroom [J]. Journal of Edible Fungi, 2023, 30(02):75-84.) and detection: The shiitake mushroom sample was placed in a 50 ℃ oven until constant weight, pulverized by a pulverizer, passed through a 40 mesh sieve, sealed in a dry container, and stored at room temperature for later use. 0.5 g of shiitake mushroom powder was weighed, and methanol was added at a material-to-liquid ratio of 1:30 (g / mL). 0.03 g / mL KOH was added, and the mixture was extracted by ultrasonication (400 W) at 50 ℃ for 80 min, followed by centrifugation (10000 r·min). -1 After 10 min, the supernatant was filtered through a 0.45 μm microporous organic filter membrane and the volume was adjusted to 100 mL. The sample solution was stored at low temperature (-20 ℃) protected from light. 282 nm was used as the detection wavelength for ergosterol and scanned using a full-wavelength microplate reader.
[0034] The statistical results of ergosterol content data in the training and prediction sets of shiitake mushroom samples are shown in Table 1. The ergosterol content in the training set of shiitake mushroom samples ranged from 9.42 mg / g to 24.85 mg / g, while the ergosterol content in the prediction set ranged from 9.00 mg / g to 24.95 mg / g. The mean and standard deviation of both sets (training set 14.95±3.85 mg / g, prediction set 17.02±4.76 mg / g) showed good consistency. Furthermore, the distribution range of chemical components in the samples differed significantly, indicating that the samples are representative to a certain extent.
[0035] Table 1. Determination of ergosterol content in shiitake mushrooms
[0036] V. Selection of Spectral Preprocessing Methods for Ergosterol Preprocessing can improve the quality and usability of spectral data, reduce noise and interference, and provide a more reliable data foundation for subsequent analysis, modeling, and applications. This invention employs seven methods—SNV, SG, MSC, NORM, SG-SNV, SG-MSC, and SG-NORM—to preprocess spectral data. Then, PLS, SVR, and CPO-LSSVM models are established using the original and preprocessed spectral data, respectively. The smaller the root mean square error (RMSEC) of the training set, the root mean square error (RMSEP) of the prediction set, and the smaller the absolute difference, the higher the coefficient of determination (R²) of the training set. 2 c The coefficient of determination R of the prediction set 2 pThe higher the residual prediction deviation (RPD), the more accurate the prediction model. A comparative analysis of the model results was conducted to determine the optimal preprocessing method for each model. The modeling results are shown in Table 2. For the PLSR model, the optimal preprocessing method is SG-NORM; for the SVR model, the optimal preprocessing method is SG-SNV; and for the CPO-LSSVM model, the optimal preprocessing method is SG-NORM.
[0037] Table 2. Modeling results of different preprocessing methods
[0038] VI. Feature Wavelength Extraction VCPA initially selected 100 characteristic wavelengths, effectively eliminating many irrelevant variables. To screen for characteristic variables highly correlated with the detection indicators, GA was used for secondary screening, ultimately extracting 47 key characteristic wavelengths. The extracted characteristic wavelengths are mainly distributed in the ranges of 900 nm–1300 nm, 1450 nm–1650 nm, 1720 nm–1740 nm, 1810 nm–2000 nm, and 2300 nm–2400 nm, as shown in the diagram. Figure 4 As shown.
[0039] VII. Establishment of Quantitative Analysis Model The CPO-LSSVM algorithm, based on LSSVM, transforms the inequality constraints of Support Vector Machines into equality constraints and employs a least-squares loss function, thus simplifying the problem into solving a system of linear equations. The CPO optimization algorithm globally searches for regularization and kernel function parameters in LSSVM to obtain the optimal hyperparameter combination, effectively improving the model's generalization ability and overall performance. The specific steps are as follows: Step (1): Set the basic parameters of the particle swarm optimization algorithm, including the number of particles. Maximum number of iterations Decision variable dimensions The optimization variable is the regularization parameter of LSSVM. Width of RBF kernel function The search spaces are set as follows: L=[0.1,0.1], U=[900,900]. The initial particle position is generated by the following formula:
[0040] in Let be the initial hyperparameter combination (γ, σ) for the i-th particle. Let it be the lower bound vector. Let it be the upper bound vector. For the number of particles, For decision variables.
[0041] The initial velocity is set to a zero vector. Then, chaotic initial values are generated for each particle. And use Logistic mapping to generate chaotic sequences:
[0042] in Let be the chaotic variable in the current iteration. 4 represents the chaotic variable for the next iteration, and 4 represents the Logistic mapping control parameter. t This represents the number of iterations.
[0043] Step (2) involves processing the parameter vector corresponding to each particle. We construct an LSSVM regression model and calculate its fitness value. Training an LSSVM is equivalent to solving the following system of linear equations:
[0044] in For model bias terms (scalar); Lagrange multiplier vectors ( ). For dimension A column vector of all 1s. Its transpose is used to express the LSSVM bias constraint. The true response value of the sample. for identity matrix These are the regularization parameters for LSSVM. Kernel matrix. Calculated using the RBF kernel:
[0045] in For the first The feature vectors of each training sample For the first The feature vectors of each training sample Let σ be the Euclidean distance between two samples, σ be the width of the RBF kernel function, and the model prediction be expressed as:
[0046] in For the model to input samples The predicted value, n The number of training samples. These are Lagrange multipliers (weighting coefficients), reflecting the first... The contribution of each training sample to the prediction; The similarity calculated for the RBF kernel function measures the similarity of the training samples. With the sample to be predicted Distance and similarity; For the first Feature vectors of training samples; x This is the input feature vector that needs to be predicted. This is the model bias term.
[0047] To accurately evaluate the generalization ability of parameter combinations, RMSECV with 10-fold cross-validation is used as the fitness function:
[0048] in n This represents the number of training samples; For the first The true reference value for each sample; For the first step in cross-validation The predicted value for each sample.
[0049] The objective of the CPO is to minimize RMSECV.
[0050] Step (3): Update the individual and global optimal positions, where the historical optimal position of each particle is denoted as . The global optimal position is denoted as . No. The formulas for updating particle velocity and position are as follows:
[0051] The inertia weight adopts a linear decreasing strategy, and its expression is as follows:
[0052] This allows for a gradual transition from global search to local search. The particle's position in the next generation is then obtained according to the following position update formula:
[0053] in For the first The particle in the first The velocity vector of the generation; For linearly decreasing inertia weights, and As a learning factor, and Derived from chaotic sequences, used to improve the diversity of particle search. For the first The best position in the history of each particle; This is the optimal position among all particles in the current iteration; For the first The particle in the first The updated position vector is used to prevent particles from going out of bounds. The updated position is truncated within the search interval. If no fitness improvement occurs for several consecutive generations, some particles are reinitialized through chaotic perturbation to enhance global search capabilities. This process is repeated until the maximum number of iterations is reached. .
[0054] Step (4): After the iteration is complete, CPO outputs the global optimal particle position. , which serves as the final hyperparameter of LSSVM. The final LSSVM model is built using the complete training set:
[0055] A spectral detection model was established by comparing the characteristic wavelengths and full wavelengths extracted using CPO-LSSVM with the measured ergosterol content of shiitake mushroom samples. Using MATLAB software, the spectral data of the characteristic wavelengths and full wavelengths of the shiitake mushroom samples were correlated with the actual measured ergosterol content and linearly fitted. The spectral data and ergosterol content measurements of 98 samples from the training set and 24 samples from the prediction set were input, and detection models were established using PLSR, SVR, and CPO-LSSVM algorithms, respectively. The training set and prediction set were then input into the model, and the determination coefficient R of the training set was used as the reference value. 2 c, Root Mean Square Error (RMSEC) of the training set, Coefficient of Determination (R) of the prediction set 2 The optimal model is obtained by taking p as the root mean square error (RMSEP) of the prediction set and the residual prediction bias (RPD).
[0056] The results are shown in Table 3. CPO-LSSVM performed poorly before wavelength selection, but achieved the best performance after selection. This is mainly because the model itself lacks a built-in mechanism for handling ultra-high-dimensional noisy data. CPO optimization, which optimizes kernel and regularization parameters during the modeling process, cannot compensate for this fundamental deficiency. CPO-LSSVM requires an external pre-selection wavelength filter to provide high-quality input features, among which the feature wavelength model (R...)... p 2 =0.9420, RMSEP=1.0950 mg / g, RPD=4.1517) compared to the full-wavelength model R 2The irgosterol concentration (IRD) increased by 0.0938, RMSE decreased by 1.6293 mg / g, and RPD increased by 2.2628. A better prediction model was built using fewer wavelengths, indicating that the selected characteristic wavelengths are highly correlated with irgosterol. The CPO-LSSVM model, after wavelength selection, performed best overall, followed by SVR and then PLSR. All three models built based on characteristic wavelengths outperformed the models built using all wavelengths, not only reducing model complexity but also improving model efficiency. The results showed that the CPO-LSSVM model established based on the characteristic wavelengths extracted by VCPA-GA had superior performance across all indicators and was considered the optimal model. Based on the standard deviation of the calibration model, the detection limit and quantitation limit of this model were calculated to be 2.8602 mg / g and 8.6672 mg / g, respectively. The ergosterol content range in the samples of this invention (9.003-24.953 mg / g) is significantly higher than the quantitation limit of this method (8.6672 mg / g). These results indicate that the established CPO-LSSVM model possesses reliable quantitative ability across this wide concentration range and is suitable for accurately predicting the ergosterol content in shiitake mushroom samples. The correlation diagram between the actual and predicted ergosterol content in the training and validation sets is attached. Figure 5 .
[0057] Table 3. Modeling Results
[0058] Example 2 This embodiment, based on the design of Embodiment 1, discloses a system for detecting ergosterol content in shiitake mushrooms, such as... Figure 6 As shown, this system, based on the aforementioned detection method, includes at least the following components built into the system: The data acquisition module is configured to acquire the absorption spectrum data of shiitake mushroom samples scanned with a near-infrared spectrometer to obtain full-wavelength spectral data; The dataset building module is configured to use Mahalanobis distance to remove abnormal shiitake mushroom samples, obtain a sample set, and divide the sample set. The preprocessing module is configured to preprocess the spectral data of the obtained sample set; The feature wavelength extraction module is configured to extract feature wavelengths from the obtained preprocessed spectral data; The detection model construction module is configured to establish a detection model using spectral data values of full wavelength and characteristic wavelength and ergosterol content determination values obtained by chemical determination method, and to evaluate the detection results of the detection model and determine the performance of the detection model; the optimal detection model is the least squares support vector machine CPO-LSSVM based on the crown porcupine optimization algorithm.
[0059] Example 3 Based on the above understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a terminal device (which may be a personal computer, mobile phone, or communication device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0060] Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Figure 7 As shown, the computer device 700 includes a processor 701 and a memory 702.
[0061] Processor 701 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 701 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 701 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 701 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 701 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0062] The memory 702 may include one or more computer-readable storage media, which may be non-transitory. The memory 702 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 702 is used to store at least one instruction, which is executed by the processor 701 to implement the method for detecting ergosterol content in shiitake mushrooms provided in this embodiment of the invention.
[0063] Those skilled in the art will understand that Figure 7 The structure shown does not constitute a limitation on the computer device 700, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0064] This invention also provides a non-transitory computer-readable storage medium, wherein when the instructions in the storage medium are executed by the processor of a computer device, the computer device is able to perform the method for detecting ergosterol content in shiitake mushrooms provided in this disclosure.
[0065] This invention also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the method for detecting ergosterol content in shiitake mushrooms provided in this invention.
[0066] Where there is no conflict, the above embodiments and features described herein can be combined with each other.
[0067] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A method for detecting ergosterol content in shiitake mushrooms based on near-infrared spectroscopy, the method comprising the following steps: S1. Use a near-infrared spectrometer to scan the absorption spectrum data of the shiitake mushroom sample to obtain full-wavelength spectral data; S2. Use Mahalanobis distance to remove abnormal shiitake mushroom samples, obtain a sample set, and divide the sample set. S3. Preprocess the spectral data after removing abnormal samples in step S2. S4. Extract characteristic wavelengths from the preprocessed spectral data obtained in step S3. S5. The content of ergosterol in shiitake mushroom samples was determined by chemical assay. S6. Establish a detection model using the spectral data values of the full wavelength and characteristic wavelength and the ergosterol content determination value obtained in step S5, and evaluate the detection results of the detection model to determine the performance of the detection model; the optimal detection model is the least squares support vector machine CPO-LSSVM based on the porcupine optimization algorithm.
2. The method according to claim 1, characterized in that, The KS algorithm is used to divide the sample set into a training set and a prediction set.
3. The method according to claim 1, characterized in that, The spectral preprocessing method mentioned includes any one of SG, MSC, NORM, SNV, SG-MSC, SG-SNV, and SG-NORM.
4. The method according to claim 1, characterized in that, The characteristic wavelength extraction process includes the following steps: S11. Use VCPA to rapidly reduce the variable space for near-infrared spectral wavelengths using EDF looping, and roughly screen out important variables. S12. Use GA to perform a fine search on the coarse screening results of VCPA to find the optimal combination of variables.
5. The method according to claim 4, characterized in that, The characteristic wavelength ranges are 900 nm to 1300 nm, 1450 nm to 1650 nm, 1720 nm to 1740 nm, 1810 nm to 2000 nm, and 2300 nm to 2400 nm.
6. The method according to claim 1, characterized in that, The process of determining the ergosterol content in shiitake mushroom samples by chemical assay is as follows: Shiitake mushroom powder is mixed with a low alcohol solution containing an alkaline catalyst and subjected to ultrasonic-assisted extraction to obtain an extract; the extract is subjected to solid-liquid separation, and the supernatant is filtered through a microporous membrane and then diluted to a final volume to obtain the sample solution; The absorbance of the sample solution was measured at the ultraviolet characteristic absorption wavelength of ergosterol to quantify the ergosterol content.
7. The method according to claim 1, characterized in that, In step S6, the parameter calculation formula for evaluating the detection model is as follows: in It is the predicted value of ergosterol content in the a-th shiitake mushroom sample; The true value of ergosterol content in the a-th shiitake mushroom sample. This represents the average ergosterol content of the samples in the training set. Predict the average ergosterol content of the concentrated samples, n c n is the number of samples in the training set. p This represents the number of samples in the prediction set.
8. The method according to claim 1, characterized in that, In step S6, to evaluate the sensitivity of the established optimal detection model, its limit of detection and limit of quantitation are calculated, taking into account the error and sensitivity of the calibration model. The specific calculation formula is as follows: Wherein the limit of detection (LOD), the limit of quantitation (LOQ), σ is the standard deviation of the prediction residuals of the calibration set, S is the slope of the calibration curve, and 3.3 and 10 are statistical confidence factors.
9. A system for detecting ergosterol content in shiitake mushrooms, characterized in that, The system, based on the method of any one of claims 1-8, includes at least the following components built within the system: The data acquisition module is configured to acquire the absorption spectrum data of shiitake mushroom samples scanned with a near-infrared spectrometer to obtain full-wavelength spectral data; The dataset building module is configured to use Mahalanobis distance to remove abnormal shiitake mushroom samples, obtain a sample set, and divide the sample set. The preprocessing module is configured to preprocess the spectral data after removing abnormal samples in step S2; The feature wavelength extraction module is configured to extract feature wavelengths from the obtained preprocessed spectral data; The detection model construction module is configured to establish a detection model by using the spectral data values of the full wavelength and the characteristic wavelength and the ergosterol content determination value obtained by the chemical determination method in step S5, respectively, and to evaluate the detection results of the detection model and determine the performance of the detection model; the optimal detection model is the least squares support vector machine CPO-LSSVM based on the porcupine optimization algorithm.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that can be executed by one or more processors to implement the method for detecting ergosterol content in shiitake mushrooms according to any one of claims 1-8.