Establishment method and detection method of tomato fructose content prediction model
By combining near-infrared spectroscopy with chemometrics, a predictive model for tomato fructose content was established, solving the problems of cumbersome and costly traditional detection methods. This enabled rapid and accurate detection of tomato fructose content, simplified the operation process, and reduced costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF VEGETABLES GUANGDONG PROV ACAD OF AGRI SCI
- Filing Date
- 2026-03-04
- Publication Date
- 2026-04-28
AI Technical Summary
Traditional methods for detecting fructose content in tomatoes are cumbersome and costly, making it difficult to meet the needs of rapid on-site analysis of large quantities. Near-infrared spectroscopy has not been used in the detection of fructose content in tomatoes.
A tomato fructose content prediction model was established by combining near-infrared spectroscopy with chemometrics. Spectral information was collected using a Fourier transform near-infrared spectrometer, and measured data were obtained using high-performance liquid chromatography. Dimensionality reduction was performed using backward interval partial least squares and a competitive adaptive reweighting algorithm to establish the fructose content prediction model.
The model enables rapid and accurate detection of fructose content in tomatoes. The cross-validation correlation coefficient and prediction correlation coefficient of the model reached 0.981 and 0.992, respectively, and the mean square error of cross-validation and the mean square error of prediction were 0.416 mg/g and 1.598 mg/g, respectively. The model simplifies the operation process and reduces costs.
Smart Images

Figure CN121933468A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vegetable quality testing technology, specifically relating to a method for establishing a prediction model for fructose content in tomatoes and a detection method. Background Technology
[0002] The flavor of a fruit largely depends on the balance between sweetness and acidity; the quality of the flavor is determined by the degree of harmony between these flavors. Tomato ( Solanum lycopersicum The flavor quality of tomatoes is one of the most important commercial traits, directly affecting their commercial value and economic benefits. The composition and content of sugars and acids in tomatoes, as well as their balance, are closely related to tomato quality, especially soluble sugars. The main soluble sugars in tomatoes are fructose and sucrose, with fructose accounting for up to 50% of the soluble sugars. Appropriate amounts of sugar can enhance and enrich the taste of tomatoes, but excessively high or low sugar content will significantly affect the taste of cooked tomatoes. The types and contents of soluble sugars regulate the flavor of tomatoes. Traditional methods for detecting soluble sugars mainly include gas chromatography-mass spectrometry (GC-MS) and high-performance liquid chromatography (HPLC). However, these methods are often complex in pretreatment, cumbersome in testing processes, costly, and time-consuming, failing to meet the requirements for rapid on-site analysis of large quantities.
[0003] Near-infrared spectroscopy (NIRS) is a rapid analytical method that is simple and fast, enabling both on-site and online analysis. It can determine multiple performance indicators within tens of seconds simply by acquiring and measuring the near-infrared spectrum of the sample. It offers advantages such as stable measurement results and good reproducibility, making it a simple, effective, and environmentally friendly new detection technology. However, there are currently no reports on the quantitative detection of fructose content, a flavor factor in tomatoes, using near-infrared technology. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a method for establishing a predictive model and a detection method for fructose content in tomatoes. This invention is based on near-infrared spectroscopy, and the established model can accurately determine the fructose content in tomatoes. Furthermore, the operation process is simple, rapid, and environmentally friendly.
[0005] This invention provides a method for establishing a prediction model for fructose content in tomatoes, comprising the following steps: Multiple tomato samples were dried and ground separately to obtain powder; The powder was subjected to near-infrared spectroscopy to obtain the original near-infrared spectrum of each tomato sample; the powder was then subjected to liquid chromatography to obtain the measured value of fructose content in each tomato sample. The original near-infrared spectrum was preprocessed, and the dimensionality of the preprocessed spectrum was reduced by backward interval partial least squares method and competitive adaptive reweighting algorithm to extract the near-infrared spectral characteristic wavelengths of fructose index. Using the selected characteristic wavelengths, a fructose content prediction model in tomatoes was established by combining chemometric methods with the measured values of fructose content. The preprocessing methods included multivariate scattering correction, normalization, SG convolution smoothing, or standard normal variable transformation.
[0006] Preferably, the preprocessing method is a standard normal variable transformation.
[0007] Preferably, the near-infrared spectral acquisition is performed using an integrating sphere solid-state sampling method, with conditions including a resolution of 4 cm⁻¹. -1 32 scans were performed, with a scanning range of 4000~12000cm. -1 .
[0008] Preferably, the conditions for the liquid chromatography detection include: a C18 column, a column temperature of 25°C; a mobile phase comprising mobile phase A and mobile phase B, wherein mobile phase A is methanol and mobile phase B is a 0.5 wt% diammonium hydrogen phosphate solution, and the volume ratio of mobile phase A to mobile phase B is 3:97; a flow rate of 0.6 mL / min; an elution time of 20 min; and an injection volume of 10 μL.
[0009] Preferably, before the liquid chromatography detection, the method further includes: mixing the powder and water for extraction to obtain an extract; the ratio of the powder to water is 20 mg: 1.5 mL.
[0010] Preferably, the preprocessing further includes the removal of abnormal samples; the software program used for the preprocessing is the functions prep.msc(), prep.norm(), prep.savgol(), and prep.snv() from the R package mdatools.
[0011] Preferably, the dimensionality reduction process includes smoothing, baseline correction, normalization, or differentiation.
[0012] Preferably, the chemometric method is partial least squares.
[0013] Preferably, the cross-validation correlation coefficient R of the tomato fructose content prediction model is... 2 The cv is 0.981, and the predictive correlation coefficient R is... 2 The p-value was 0.992, the mean squared error of cross-validation was 0.416 mg / g, and the mean squared error of prediction was 1.598 mg / g.
[0014] This invention also provides a method for detecting fructose content in tomatoes, comprising the following steps: Using the method described in the above technical solution, a prediction model for fructose content in tomatoes was obtained; The tomato sample to be tested was dried and ground to obtain the powder to be tested. The powder to be tested is subjected to near-infrared spectroscopy to obtain the original near-infrared spectrum of the sample. The original near-infrared spectrum of the sample to be tested was preprocessed, and the fructose content in the tomato sample was obtained according to the fructose content prediction model in tomatoes.
[0015] Compared with the prior art, the present invention has the following beneficial effects: This invention provides a method for rapidly detecting fructose content, a key determinant of tomato flavor, based on near-infrared spectroscopy. First, near-infrared spectra of tomato samples are collected, and the samples are divided into a calibration set and a validation set, with outliers removed. Second, different preprocessing methods are compared and screened for the near-infrared spectra. Third, biPLS-CARS (backward interval partial least squares-competitive adaptive reweighting algorithm) is used to reduce the dimensionality of the spectra, extracting the characteristic wavelengths of the near-infrared spectrum for fructose content. Finally, using the selected characteristic wavelengths, a predictive model for tomato flavor factor fructose content is established using chemometric methods combined with the absolute fructose content. The predictive model is then validated using a validation set, ultimately yielding the optimal predictive model for tomato flavor factor fructose content. This invention's model effectively compresses useless variables and interference information in near-infrared spectra, accurately detecting the content of tomato flavor factor fructose. Furthermore, the operation is simple, rapid, and environmentally friendly, providing a new technical means for detecting tomato flavor factor fructose content and offering technical support for tomato quality breeding and shortening the breeding process. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 The original near-infrared spectrum of the tomato sample; Figure 2 A diagram showing the impact of different pretreatment methods on RMSECV; Figure 3 A graph showing the effect of the number of iterations on RMSECV; Figure 4 Principal component analysis plots of predicted values and measured values from the biPLS-CARS-PLS model; Figure 5This is a scatter plot of near-infrared predicted values and liquid chromatography measured values of fructose content. Detailed Implementation
[0018] This invention provides a method for establishing a prediction model for fructose content in tomatoes, comprising the following steps: Multiple tomato samples were dried and ground separately to obtain powder; The powder was subjected to near-infrared spectroscopy to obtain the original near-infrared spectrum of each tomato sample; the powder was then subjected to liquid chromatography to obtain the measured value of fructose content in each tomato sample. The original near-infrared spectrum was preprocessed, and the dimensionality of the preprocessed spectrum was reduced by backward interval partial least squares method and competitive adaptive reweighting algorithm to extract the near-infrared spectral characteristic wavelengths of fructose index. Using the selected characteristic wavelengths, a fructose content prediction model in tomatoes was established by combining chemometric methods with the measured values of fructose content. The preprocessing methods included multivariate scattering correction, normalization, SG convolution smoothing, or standard normal variable transformation.
[0019] Unless otherwise specified, all materials and equipment used in this invention are commercially available products in the field.
[0020] This invention provides a method for rapidly detecting the fructose content, a key factor in tomato flavor, based on near-infrared spectroscopy. By combining near-infrared spectral information acquired by a Fourier transform near-infrared spectrometer with measured data of fructose content in tomatoes obtained by high-performance liquid chromatography, a predictive model for fructose content in tomatoes is established, providing a new technical means for the rapid and accurate detection of fructose content in tomatoes.
[0021] This invention involves drying and grinding multiple tomato samples separately to obtain powder.
[0022] In this invention, the tomato (fruit) samples are preferably representative tomato germplasm resources with significant differences in the origin and phenotypic differences; the number of tomato samples can be 200, of which 160 are calibration sets and 40 are validation sets.
[0023] In this invention, the process before drying also includes cutting the tomato sample into pieces; there are no special requirements for the pieces, but the mesocarp from the middle part of the tomato can be used.
[0024] In this invention, the drying is preferably freeze-drying, which specifically includes freezing at -80°C for at least 5 hours, followed by drying at -54°C for 48 hours. The grinding process preferably also includes sieving, preferably using a 50-mesh standard sieve.
[0025] After obtaining the powder, the present invention performs near-infrared spectroscopy on the powder to obtain the original near-infrared spectrum of each tomato sample; and performs liquid chromatography detection on the powder to obtain the measured value of fructose content in each tomato sample.
[0026] In this invention, the near-infrared spectral acquisition preferably uses a Fourier transform near-infrared spectrometer, specifically a PerkinElmer FT-NIR Fourier transform near-infrared spectrometer. The sampling method is preferably integrating sphere solid-state sampling; the preferred acquisition conditions include: a resolution of 4 cm⁻¹. -1 32 scans were performed, with a scanning range of 4000~12000cm. -1 The near-infrared spectrum is preferably a near-infrared diffuse reflectance spectrum.
[0027] In this invention, the conditions for the liquid chromatography detection preferably include: The chromatographic column was a C18 column, and the column temperature was 25℃. The mobile phase consisted of mobile phase A and mobile phase B. Mobile phase A was methanol, and mobile phase B was a 0.5wt% diammonium hydrogen phosphate solution. The volume ratio of mobile phase A to mobile phase B was 3:97. The flow rate was 0.6 mL / min. The elution time was 20 min. The injection volume was 10 μL.
[0028] In this invention, the chromatographic column can specifically be a Waters Atlantis T3 Column with dimensions of 4.6 mm × 250 mm × 5 μm.
[0029] In this invention, the process before liquid chromatography detection further includes: mixing the powder and water for extraction to obtain an extract (which is then injected into the liquid chromatograph for detection). The preferred ratio of powder to water is 20 mg: 1.5 mL; the extraction is preferably performed under vortex conditions, and the vortexing time is preferably 10 min. The process after extraction preferably includes solid-liquid separation and membrane filtration. The solid-liquid separation is preferably centrifugation, with a preferred centrifugation speed of 12000 rpm and a preferred centrifugation time of 5 min; the pore size of the filter membrane used for membrane filtration is preferably 0.45 μm.
[0030] This invention establishes a tomato quality database based on liquid chromatography, using measured fructose content in each tomato sample, with taste as the primary information. In the embodiments of this invention, the fructose content range obtained is 101.35~226.20 mg / g (dry weight).
[0031] After obtaining the original near-infrared spectrum and the measured value of fructose content for each tomato sample, this invention preprocesses the original near-infrared spectrum, using backward interval partial least squares and competitive adaptive reweighting algorithm to reduce the dimensionality of the preprocessed spectrum, extracting the near-infrared spectral characteristic wavelengths of fructose index, and using the selected characteristic wavelengths in combination with the measured value of fructose content through chemometric methods to establish a predictive model for fructose content in tomatoes; the preprocessing methods include multivariate scattering correction, normalization, SG convolution smoothing, or standard normal variable transformation.
[0032] In this invention, the preprocessing process also includes removing abnormal samples because, during data processing, some samples may have large deviations between the predicted and actual values due to measurement errors. The presence of these samples can affect the modeling effect.
[0033] In this invention, the preprocessing methods include multivariate scattering correction (MSC), normalization (Nor), Savitzky-Golay convolution smoothing, or standard normal variable transformation (SNV), preferably standard normal variable transformation. The software programs used for preprocessing are preferably the functions prep.msc(), prep.norm(), prep.savgol(), and prep.snv() from the R package mdatools. The preprocessing methods described in this invention result in a small root mean square error in the cross-validation of the partial least squares regression model, which helps to reduce noise interference and enhances the model's predictive ability and robustness.
[0034] In this invention, the method for extracting (screening) the characteristic wavelengths of the near-infrared spectrum is as follows: The near-infrared spectral model for tomato quality is optimized by combining a region-based wavelength selection method and a single-variable-based wavelength selection method using backward inter-regional partial least squares (biPLS) and a competitive adaptive reweighting algorithm (CARS). The corresponding band range obtained by the preliminary optimization using backward inter-regional partial least squares (biPLS) is 4895.21~5031.27cm. -1 and 5321.21~5500cm -1 .
[0035] In this invention, the method for establishing the prediction model is specifically as follows: A mathematical model for fructose, a taste factor in tomatoes, was established using selected characteristic wavelengths and the partial least squares (PLS) method combined with the absolute content of the index (measured fructose content). The number of factors in the model was determined by cross-validation. The optimal model was selected by comparing the cross-validation correlation coefficient Rcv, prediction correlation coefficient Rp, root mean square error of cross-validation RMSECV, and root mean square error of prediction RMSEP. The closer Rc and Rp are to 1, and the lower the RMSECV and RMSEP are, the better the predictive ability and stability of the model. All the above algorithms, as well as the establishment and evaluation of the final model, were completed using built-in functions of the R package mdatools.
[0036] In this invention, after establishing the prediction model for fructose content in tomatoes, it is preferable to further validate the prediction model. This invention does not have any special requirements for this validation. The optimal tomato taste factor-fructose content prediction model finally established in this embodiment of the invention has been validated: the cross-validation correlation coefficient (R0) is... 2 cv) = 0.981, predictive correlation coefficient (R) 2 p)=0.992, cross-validation mean square error (RMSECV)=0.416mg / g, prediction mean square error (RMSEP)=1.598mg / g.
[0037] This invention also provides a method for detecting fructose content in tomatoes, comprising the following steps: Using the method described in the above technical solution, a prediction model for fructose content in tomatoes was obtained; The tomato sample to be tested was dried and ground to obtain the powder to be tested. The powder to be tested is subjected to near-infrared spectroscopy to obtain the original near-infrared spectrum of the sample. The original near-infrared spectrum of the sample to be tested was preprocessed, and the fructose content in the tomato sample was obtained according to the fructose content prediction model in tomatoes.
[0038] In this invention, the processing method, near-infrared spectral acquisition, and preprocessing method of the tomato sample to be tested are preferably consistent with those used in establishing the prediction model, and will not be repeated here. This invention can quickly obtain the fructose content of the tomato sample to be tested based on the prediction model.
[0039] This invention establishes a predictive model for tomato flavor factor-fructose content by combining near-infrared spectral information acquired by a Fourier transform near-infrared spectrometer with measured data of tomato flavor factor-fructose content obtained by high-performance liquid chromatography. By employing spectral preprocessing, spectral noise, baseline drift, and translation caused by factors such as instrument, sample, and spectral acquisition environment are effectively eliminated. Simultaneously, backward partial least squares (BiPLS) combined with competitive adaptive weighting (CARS) is used to reduce the dimensionality of the spectrum, extract characteristic wavelengths, and remove variables unrelated to tomato acidity and flavor. Furthermore, the optimal biPLS-CARS-PLS model (predictive model for tomato flavor factor-fructose content) is selected by comparing the model's cross-validation correlation coefficient Rc(rcv), prediction correlation coefficient Rp, root mean square error of cross-validation RMSECV, and root mean square error of prediction RMSEP. The model's cross-validation correlation coefficient (Rc(rcv)) is verified to be the best. 2 cv), predictive correlation coefficient (R) 2 The mean squared error of cross-validation (RMSECV) and mean squared error of prediction (RMSEP) were 0.981, 0.992, 0.416 mg / g and 1.598 mg / g, respectively. This model can effectively compress useless variables and interference information in the near-infrared spectrum and accurately detect the content of fructose, a tomato taste factor, providing a new method for rapidly and accurately establishing a near-infrared spectral model of tomato fruit taste.
[0040] To further illustrate the present invention, the method for establishing and detecting the fructose content prediction model in tomatoes provided by the present invention will be described in detail below with reference to the accompanying drawings and embodiments, but these should not be construed as limiting the scope of protection of the present invention.
[0041] In the embodiments of the present invention, the establishment method and the instruments used for quantifying chemical values are shown in Table 1.
[0042] Table 1 Test Instruments
[0043] Example 1: Rapid detection of fructose content, a flavor factor in tomatoes, using near-infrared spectroscopy. 1. Material preparation The experimental materials used to construct the predictive model for fructose content in tomatoes were representative tomato resources with significant differences in germplasm origin (including Peruvian, American, Dutch, and native Chinese resources) and appearance phenotypes (fruit size, skin color), totaling 200 samples. For each tomato fruit, a piece of pulp (from the mesocarp in the middle) was selected, cut into chunks, and frozen at -80℃ for at least 5 hours. Then, it was freeze-dried at -54℃ for 48 hours, and finally ground to prepare standard tomato pulp powder samples with uniform dryness and particle size (passing through a 50-mesh standard sieve). The tomato samples were then randomly divided into a validation set and a calibration set, with 160 samples in the calibration set and 40 in the validation set.
[0044] 2. Spectral Acquisition Near-infrared spectroscopy was performed on representative tomato resources with different characteristics. The raw near-infrared diffuse reflectance spectra of all tomato powder samples were acquired using a PerkinElmer FT-NIR Fourier transform near-infrared spectrometer. Sampling method: integrating sphere solid sampling; acquisition conditions: resolution 4 cm⁻¹. -1 32 scans were performed, with a scanning range of 4000~12000 cm. -1 Before each scan, the sample cup was shaken. The sample thickness, packing density, and particle uniformity were kept consistent throughout the experiment (tomato powder that had passed through a 50-mesh sieve was packed into a cylindrical sample cup with an inner diameter of 10 mm and a depth of 2 mm, filling approximately 4 / 5 of the cup's volume. Then, a flat glass plate was used to gently scrape the rim of the cup, flattening the powder surface until it was flush with the rim and achieving a moderately compacted state). The collected spectra are shown below. Figure 1 (The horizontal axis represents wavenumber, in cm) -1 ).
[0045] 3. Determination of fructose content, a flavor factor in tomato fruit, using liquid chromatography. 1) Preparation of the test solution Take 20 mg of dried tomato pulp powder (all tomato powder samples in step 1), add 1.5 mL of extraction solution (pure water) to each sample, vortex the sample for 10 min after adding the extraction solution, centrifuge the resulting mixture at 12000 rpm for 5 min, collect the supernatant and filter it through a 0.45 μm filter membrane to obtain fructose extract.
[0046] 2) Chromatographic conditions High-performance liquid chromatography (HPLC): Waters Atlantis e2695 quaternary gradient pump system; detector: refractive index detector (RID); Atlantis e2695 column oven, column temperature: 25℃; column: Waters Atlantis T3 column (4.6 mm × 250 mm, 5 μm); mobile phase: A:B = 3:97 (volume ratio, A: methanol, B: 0.5 wt% diammonium hydrogen phosphate solution); flow rate: 0.6 mL / min; elution time: 20 min; injection volume: 10 μL.
[0047] 3) Results Processing The purchased fructose standard (Sigma-Aldrich) was analyzed to obtain a standard curve. The obtained liquid chromatogram was analyzed in combination with the standard curve, peak retention time and peak area to calculate the fructose content of the tomato pulp powder sample, which ranged from 101.35 to 226.20 mg / g (dry weight).
[0048] 4. Remove abnormal samples The principle of outlier removal is as follows: At the spectral information level, the orthogonal and score distances of the sample spectrum under a certain principal component are calculated. At the reference value level, the residuals between the predicted and reference values are calculated to obtain the distance between them and the degree of deviation of the sample.
[0049] Based on the above principle, using the original full spectrum information collected in step 2, the mdatools function categorize() was used to identify 11 extreme or outlier samples from 160 calibration set samples.
[0050] 5. Screening and Determination of Spectral Preprocessing Methods Due to factors such as instrumentation, sample quality, and spectral acquisition environment, near-infrared spectra often exhibit noise, baseline drift, and translation. To eliminate the impact of these adverse factors on the model, the raw spectra (none) acquired in step 2 should be preprocessed. Therefore, four preprocessing methods were employed for data optimization: multivariate scattering correction (MSC), normalization (Nor), Savitzky-Golay convolution smoothing, and standard normal variable transformation (SNV). The software programs used were the R package `mdatools` functions `prep.msc()`, `prep.norm()`, `prep.savgol()`, and `prep.snv()`. The spectral preprocessing method was selected based on minimizing the root mean square error of the partial least squares regression model's cross-validation.
[0051] The optimal PLS model established by standard normal variable transformation (SNV) preprocessing is ( Figure 2 The model's Rcv (cross-validation correlation coefficient) improved from 0.745 without preprocessing to 0.981; the RMSECV (mean squared error of cross-validation) decreased from 0.981 mg / g without preprocessing to 0.572 mg / g. Other preprocessing methods showed little improvement. This indicates that SNV preprocessing is beneficial for reducing noise interference and can enhance the model's predictive ability and robustness. Therefore, the SNV-processed spectra were used for subsequent analyses.
[0052] Backward partial least squares (biPLS) was used to remove variables unrelated to the taste and efficacy of tomatoes. biPLS is a region-based variable selection method built upon interval partial least squares (iPLS). In this embodiment, the entire band range acquired in step 2 was divided into 10 equally wide sub-intervals. One sub-interval was removed. Within the remaining intervals, partial least squares regression models for each combined interval were calculated. The sub-interval removed from the combined model with the smallest mean squared cross-validation error (RMSECV) was selected as the first removed sub-interval. This process was repeated until the program terminated. After multiple iPLS iterations, the final spectral intervals were obtained, and each iteration showed higher predictive performance than the model for the intervals removed in the previous iteration. The obtained intervals are shown below. Figure 1 As shown, the corresponding band range is 4895.21~5031.27 cm. -1 5321.21~5500 cm -1 .
[0053] 6. Quadratic deep optimization based on the CARS model (competitive adaptive reweighting algorithm) Although the biPLS algorithm removes a large amount of information unrelated to tomato fructose and improves model performance, as a spectral variable region selection method, adjacent variables still exhibit high correlation within the selected intervals. A 1ibPLS package built using Matlab was used for CARS, which, after translation, can be applied to the R platform for near-infrared spectroscopy analysis. In the CARS algorithm, each time adaptive reweighted sampling (ARS) is used, points with larger absolute values of regression coefficients in the PLS model are retained as a new subset, while points with smaller weights are removed. Then, a PLS model is built based on the new subset. After multiple calculations, the wavelengths in the subset with the smallest root mean square error (RMSECV) of the PLS model cross-validation are selected as feature wavelengths, resulting in the minimum error in the biPLS-CARS-PLS regression model. The corresponding spectral characteristic wavenumbers are: 4901.48, 4907.99, 4908.65, 4912.25, 4920.17, 4930.83, 4954.26, 4955.63, 4966.02, 4972.66, 4973.36, 4976.17, 4976.88, 5031.27 and 5321.21, 5327.13, 5334.21, 5331.90, 5343.07, 5363.05, 5366.16, 5367.41, 5369.60, 5375.62, 5378.29, 5379.58, respectively. 5380.58, 5388.01, 5389.91, 5396.14, 5406.31, 5404.99, 5409.18, 5415.25, 5425.17, 5435.83, 5455.56, 5458.13, 5467.12, 5473.66, 5474.31, 5477.58, 5478.19, 5500.00 (Unit: cm) -1 In this embodiment, the secondary depth optimization of the CARS-based model is performed using the carpls() function with the following parameters: iteration=50, fold=10, nLV=15.
[0054] After 50 iterations, the lowest RMSECV (root mean square error of cross-validation) at the 25th (optimal) iteration was 1.44, with 15 principal components. (See...) Figure 3 Table 2 shows the impact of different models and variable selection methods on Rc and RMSEC.
[0055] Table 2. The impact of different model and variable selection methods on Rc and RMSEC
[0056] 7. Optimized model building and prediction The built-in function `pls()` in the R package `mdatools` is set with the parameters: `ncomp=20, cv=10, scale=T, method="simpls", center=F`. The fructose content of 20 selected tomato pulp powder samples is predicted, demonstrating that the `predict()` function can perform the prediction.
[0057] Further modeling was performed on the extracted characteristic wavelengths of 200 tomato pulp powder samples, and predictions were made on 10 validation set samples. The prediction results are shown in Table 3. In the sample number, GC represents the planting location, and 297-302 represents which variety in this batch. Taking the sample number in the first row as an example, 1② represents which row, and 1①-2 represents which column and which number.
[0058] Table 3 Comparison of model prediction results and liquid chromatography determination values.
[0059] The results showed that the established biPLS-CARS-PLS model could effectively predict the fructose content of tomato powder samples, with deviations all within ±30%. Finally, a biPLS-CARS-PLS predictive model for tomato taste factors and fructose was established: cross-validation correlation coefficient (R0). 2 cv) = 0.981, predictive correlation coefficient (R) 2 p)=0.992, cross-validation mean square error (RMSECV)=0.416mg / g, prediction mean square error (RMSEP)=1.598 mg / g.
[0060] Figure 4 The plot shows the principal component analysis of the measured values (determined by liquid chromatography) and predicted values (predicted by near-infrared model) of the biPLS-CARS-PLS model, where red dots represent liquid chromatography and blue dots represent near-infrared. It can be seen that the two data are significantly correlated and there is no obvious clustering result.
[0061] Figure 5 The scatter plot shows the near-infrared predicted values and liquid chromatography measured values of fructose content; it can be seen that the predicted values and measured values are correlated, and the data exhibit collinearity.
[0062] 8. Evaluation of the taste and quality of tomatoes For tomatoes whose acidity and taste quality need to be determined, freeze-dry and grind them into powder. Then, collect near-infrared diffuse reflectance spectra according to the method in section 2. After preprocessing the spectral data according to the method in section 5, input the biPLS-CARS-PLS regression model according to the preferred characteristic wavelength to obtain the fructose content of tomatoes and complete the rapid evaluation of tomato taste quality.
[0063] Although the above embodiments have provided a detailed description of the present invention, they are only some embodiments of the present invention, not all embodiments. People can obtain other embodiments based on the present invention without creative effort, and these embodiments all fall within the protection scope of the present invention.
Claims
1. A method for establishing a predictive model for fructose content in tomatoes, characterized in that, Includes the following steps: Multiple tomato samples were dried and ground separately to obtain powder; The powder was subjected to near-infrared spectroscopy to obtain the original near-infrared spectrum of each tomato sample; the powder was then subjected to liquid chromatography to obtain the measured value of fructose content in each tomato sample. The original near-infrared spectrum was preprocessed, and the dimensionality of the preprocessed spectrum was reduced by backward interval partial least squares method and competitive adaptive reweighting algorithm to extract the near-infrared spectral characteristic wavelengths of fructose index. Using the selected characteristic wavelengths, a fructose content prediction model in tomatoes was established by combining chemometric methods with the measured values of fructose content. The preprocessing methods included multivariate scattering correction, normalization, SG convolution smoothing, or standard normal variable transformation.
2. The method for establishing according to claim 1, characterized in that, The preprocessing method is standard normal variable transformation.
3. The method for establishing according to claim 1, characterized in that, The near-infrared spectral acquisition method is integrating sphere solid sampling, with the following conditions: resolution 4 cm⁻¹. -1 32 scans were performed, with a scanning range of 4000~12000cm. -1 .
4. The method for establishing according to claim 1, characterized in that, The conditions for the liquid chromatography detection include: a C18 column, a column temperature of 25℃; mobile phases including mobile phase A and mobile phase B, mobile phase A being methanol and mobile phase B being a 0.5wt% diammonium hydrogen phosphate solution, with a volume ratio of mobile phase A to mobile phase B of 3:97; a flow rate of 0.6 mL / min; an elution time of 20 min; and an injection volume of 10 μL.
5. The method for establishing according to claim 1 or 4, characterized in that, The liquid chromatography detection process further includes: mixing the powder and water for extraction to obtain an extract; the ratio of powder to water is 20 mg: 1.5 mL.
6. The method for establishing according to claim 1, characterized in that, The preprocessing process also includes the removal of abnormal samples; the software program used for the preprocessing is the R package mdatools functions prep.msc(), prep.norm(), prep.savgol(), and prep.snv().
7. The method for establishing according to claim 1, characterized in that, The dimensionality reduction process includes smoothing, baseline correction, normalization, or differentiation.
8. The method for establishing according to claim 1, characterized in that, The chemometric method is partial least squares.
9. The method for establishing according to claim 1, characterized in that, The cross-validation correlation coefficient R of the tomato fructose content prediction model. 2 The cv is 0.981, and the predictive correlation coefficient R is... 2 The p-value was 0.992, the mean squared error of cross-validation was 0.416 mg / g, and the mean squared error of prediction was 1.598 mg / g.
10. A method for detecting fructose content in tomatoes, characterized in that, Includes the following steps: A model for predicting fructose content in tomatoes is obtained by using the method of any one of claims 1 to 9. The tomato sample to be tested was dried and ground to obtain the powder to be tested. The powder to be tested is subjected to near-infrared spectroscopy to obtain the original near-infrared spectrum of the sample. The original near-infrared spectrum of the sample to be tested was preprocessed, and the fructose content in the tomato sample was obtained according to the fructose content prediction model in tomatoes.