Grape SSC detection and origin traceability method based on NIR multi-task ensemble learning

Through the NIR multi-task ensemble learning method, the Boruta algorithm and XGBoost model are used to solve the accuracy and applicability of multi-task analysis in grape SSC detection and origin traceability, and realize high-precision multi-task collaborative optimization and non-destructive rapid detection.

CN120375070APending Publication Date: 2025-07-25AM INC FOR METROLOGY & TESTING TECH SERVICES
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510465757.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing NIR spectroscopy technology is difficult to achieve multi-task analysis in grape SSC detection and origin traceability, with low prediction accuracy and insufficient algorithm applicability.

Method used

Using a method based on NIR multi-task ensemble learning, shared features are identified through Boruta algorithm and combined with XGBoost integrated learning model, grape SSC detection and origin traceability are carried out, and multi-task collaborative optimization is used.

Benefits of technology

It significantly improves the SSC regression prediction accuracy (R2 reaches 0.972) and the accuracy of origin classification (97%), realizes non-destructive rapid detection, is suitable for industrial assembly line applications, and enhances the generalization ability and prediction accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375070A_ABST
    Figure CN120375070A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of agricultural product detection, and discloses a grape SSC detection and origin traceability method based on NIR multi-task ensemble learning. The method comprises the following steps: (1) carrying out near infrared spectrum transmission scanning on a grape sample, and measuring the content of soluble solids; (2) integrating the spectral data, the soluble solid content and different production place categories, and splitting into a training set and a test set; (3) identifying soluble solid content characteristics and production place category characteristics by using a Boruta algorithm, and calculating average importance after the characteristics are combined; (4) setting an importance threshold, screening out shared features, and obtaining a reconstruction training set and a reconstruction test set; (5) constructing an XGBoost integrated learning model by using the reconstructed training set, carrying out global tuning on hyper-parameters, and training to obtain a final model; and (6) performing prediction based on the final model to complete grape SSC detection and origin traceability. The method provided by the invention can overcome the defects of low feature utilization rate and limited precision of a traditional single-task model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of agricultural product detection, and particularly to a method for detecting grape SSC and origin traceability based on NIR multi-task integrated learning. Background Art

[0002] Grapes are typical varieties in Xinjiang and are widely planted in areas such as Turpan, Hotan, and Kashgar. Due to sufficient sunlight in the Xinjiang region, its fruits are plump, rich in nutrients, and sweet and sour, and are deeply favored by consumers. The quality and flavor of grapes determine the commercial decisions of producers and the purchasing trends of consumers, and the most important indicator for measuring the quality and flavor of grapes is the soluble solid content (SSC). The SSC of grapes is the most important indicator of taste sensation, and an appropriate SSC makes the fruits sweet and sour. However, grapes from different regions result in significant differences in SSC, which significantly affects the nutritional quality of 'Munage' grapes, and ultimately leads to false propaganda in some regions that their fruits come from high-quality regions. Therefore, it is necessary to adopt detection technologies to accurately predict the grape SSC and origin.

[0003] Traditional food detection technologies such as high-performance liquid chromatography, stable isotopes, refractometers, acidimeters, etc. require destroying samples and consume a large amount of time and manpower, and cause food waste. In addition, due to the application of equipment, the detection cost is increased and it is difficult to achieve industrial application. Near-infrared (NIR) spectroscopy has received extensive attention in food detection due to its non-destructive, fast, accurate, and low-cost characteristics. The collected spectra cover the NIR band from 700 nm to 2500 nm, and contain the vibration energies of hydrogen-containing molecular groups of many chemical components such as O-H, C-H, and C-O. Currently, NIR spectroscopy has been used for the accurate detection of food characteristics such as component quantification, origin traceability, adulterant and pollutant identification, etc. However, most studies and methods focus on the single-task analysis of NIR spectra, that is, single regression or classification tasks, and rarely perform multi-task analysis. The main difficulties lie in prediction accuracy and algorithm applicability. First, single-task analysis has low requirements for feature selection, while multi-task analysis needs to consider the shared features between tasks, resulting in low prediction accuracy. Second, the multi-task integrated learning algorithm is lacking, and it is difficult to apply to multi-task prediction.

[0004] Therefore, aiming at the problems existing in the actual application of the above NIR spectra, it is of great significance to provide a method for detecting grape SSC and origin traceability with the technical characteristics of multi-algorithm integration, high-precision learning, and multi-task recognition. Summary of the Invention

[0005] The object of the present invention is to overcome the problem that the grape detection method based on NIR spectra in the prior art cannot be effectively applied to multi-task analysis.

[0006] To achieve the above object, the present invention provides a method for detecting grape SSC and tracing the origin based on NIR multi-task integrated learning, the method comprising the following steps:

[0007] (1) Perform near-infrared spectral transmission scanning on grape samples from different origins to obtain spectral data, and measure the soluble solid content of each sample; wherein, the wavelength range of the near-infrared spectral transmission scanning is 700-2500 nm;

[0008] (2) Integrate the spectral data, the soluble solid content and the corresponding origin categories obtained in step (1) to obtain a spectral data set and split it into a training set and a test set;

[0009] (3) Use the Boruta algorithm to identify the soluble solid content features and origin category features of the grape samples, and calculate their average importance after feature combination;

[0010] (4) Set an importance threshold to screen out shared features, and reconstruct the training set and the test set according to the shared features to obtain a reconstructed training set and a reconstructed test set;

[0011] (5) Use the reconstructed training set to construct an XGBoost integrated learning model, use the Bayes algorithm to globally tune hyperparameters, and then train with the optimal hyperparameter combination to obtain a final model;

[0012] (6) Perform prediction based on the final model to complete grape SSC detection and origin tracing, wherein, using the reconstructed test set as input, and using the regression values of the soluble solid content of each grape sample and the classification results of the origin categories as output.

[0013] Compared with the prior art, the method provided by the present invention has at least the following advantages:

[0014] (1) Multi-task collaborative optimization: Synchronously process the quantitative detection of soluble solids (SSC) and the origin classification task through an integrated learning framework, and use the shared feature extraction technology (Boruta algorithm) to mine the key wavelengths related to multi-tasks in the spectral data, overcoming the defects of low feature utilization rate and limited accuracy of traditional single-task models;

[0015] (2) High precision and robustness: Combining the XGBoost integrated model with Bayesian hyperparameter optimization significantly improves the SSC regression prediction (R 2 reaching 0.972) and the origin classification accuracy (97%), and the performance is greatly improved compared with the traditional PLS method (R 2 = 0.821, accuracy 81.5%);

[0016] (3) Non-destructive and efficient detection: Based on near-infrared spectroscopy technology, non-destructive and rapid detection is realized, avoiding the problems of sample destruction, time-consuming and consumable materials of traditional chemical methods (such as high-performance liquid chromatography), and is suitable for industrial production line applications;

[0017] (4) Algorithm innovation: Introduce a multi-task feature importance fusion mechanism, screen shared features through thresholds, effectively reduce data dimensions and suppress noise interference, and enhance the generalization ability of the model;

[0018] (5) Geographical indication protection support: Provide accurate quality evaluation and origin authenticity identification tools for geographical indication products such as grapes, and curb market false labeling behavior; The integrated algorithm to identify shared features can improve the prediction accuracy and enable effective analysis of NIR spectroscopy multi-tasks. Description of the Drawings

[0019] Figure 1 It is the near-infrared spectrogram of grapes involved in the preferred embodiment of the present invention;

[0020] Figure 2 It is the average near-infrared spectrum of grapes involved in the preferred embodiment of the present invention;

[0021] Figure 3 It is the shared features screened out in the preferred embodiment of the present invention;

[0022] Figure 4 It is the scatter plot of the comparison between the SSC of grapes predicted and output by the final model and the actually measured SSC in the preferred embodiment of the present invention;

[0023] Figure 5 It is the confusion matrix predicted and output in the preferred embodiment of the present invention. Detailed Embodiment

[0024] In the ranges disclosed herein, the endpoints and any values are not limited to the exact ranges or values, and these ranges or values should be understood to include values close to these ranges or values. For numerical ranges, between the endpoint values of each range, between the endpoint values of each range and individual point values, and between individual point values can be combined with each other to obtain one or more new numerical ranges, and these numerical ranges should be regarded as specifically disclosed herein.

[0025] As described above, the present invention provides a method for detecting the SSC of grapes and tracing the origin based on NIR multi-task integrated learning, and the method includes the following steps:

[0026] (1) Perform near-infrared spectral transmission scanning on grape samples from different origins to obtain spectral data, and measure the soluble solid content of each sample; wherein, the wavelength range of the near-infrared spectral transmission scanning is 700 - 2500 nm;

[0027] (2) Integrate the spectral data, the soluble solid content, and the corresponding origin categories obtained in step (1) to obtain a spectral dataset, and split it into a training set and a test set;

[0028] (3) Use the Boruta algorithm to identify the soluble solid content features and origin category features of the grape samples, and calculate their average importance after feature combination;

[0029] (4) Set an importance threshold to screen out shared features, and reconstruct the training set and the test set according to the shared features to obtain a reconstructed training set and a reconstructed test set;

[0030] (5) Use the reconstructed training set to construct an XGBoost integrated learning model, use the Bayes algorithm to globally tune the hyperparameters, and then train with the optimal hyperparameter combination to obtain the final model;

[0031] (6) Perform predictions based on the final model to complete grape SSC detection and origin traceability. Among them, use the reconstructed test set as the input, and use the regression values of the soluble solid content of each grape sample and the classification results of the origin categories as the output.

[0032] Preferably, in step (1), the grape samples are selected from at least two of the grapes from three origins: Hotan, Turpan, and Kashgar in Xinjiang.

[0033] Preferably, the grape samples have no rot or bruises, and before performing the near-infrared spectral transmission scan, wash the grape samples first.

[0034] Preferably, in step (1), the number of each grape sample is not less than 100, and origin markings are made.

[0035] Preferably, in step (1), the near-infrared spectral transmission scan is performed using a bench-type or hand-held near-infrared spectrometer, and the scanning directions are evenly distributed around the fruit, and the number of near-infrared spectra collected for each grape fruit is not less than 3.

[0036] Preferably, in step (1), a refractometer is used to measure the soluble solid content of each sample, and the number of measurements is not less than 2 times.

[0037] Preferably, in step (2), the splitting is performed by the method of stratified random sampling, and the sample size of the obtained test set is not less than 20% of the sample size of the spectral dataset. This preferred situation is conducive to ensuring that the distributions of origin categories and soluble solid content in the obtained training set and test set are consistent.

[0038] Preferably, in step (3), the Boruta algorithm is based on random forest and is implemented using R code.

[0039] Preferably, in step (4), set the importance threshold p = 0.6, that is, retain the top 60% of the features.

[0040] Preferably, in step (5), when constructing the XGBoost ensemble learning model using the reconstructed training set, the accuracy of the regression model is evaluated by the coefficient of determination and the root mean square error. Among them, the coefficient of determination is calculated by formula (I), the root mean square error is calculated by formula (II), and the closer the coefficient of determination is to 1 and the smaller the root mean square error, the higher the regression accuracy; and

[0041]

[0042] where y obs is the actually measured soluble solid content of the grape sample; y pre is the predicted soluble solid content of the grape sample; n is the number of grape samples.

[0043] Preferably, in step (5), when constructing the XGBoost ensemble learning model using the reconstructed training set, the accuracy of the analysis model is evaluated by accuracy and recall rate. Among them, based on the confusion matrix, the accuracy is calculated by formula (III), the recall rate is calculated by formula (IV), and the closer the accuracy and the recall rate are to 100%, the higher the classification accuracy; and

[0044] The confusion matrix is:

[0045] where TP is the positive correct sample; TN is the negative correct sample; FP is the positive wrong sample; FN is the negative wrong sample;

[0046] Accuracy (%) = (TP + TN) / (TP + TN + FP + FN) × 100 (III);

[0047] Recall rate (%) = TP / (TP + FN) × 100 (IV).

[0048] The present invention is described in detail below through examples. Unless otherwise specified, the raw materials used are all ordinary commercially available products.

[0049] Example 1

[0050] This embodiment is used to illustrate the grape SSC detection and origin tracing method based on NIR multi-task integrated learning provided by the present invention, which is carried out according to the following steps:

[0051] (1) Respectively collect 108, 110, and 116 grape fruits without rot or bruise from Hotan, Turpan, and Kashgar regions in Xinjiang, wash them, and use them as grape samples, and mark the origin of each grape fruit; use a desktop near-infrared spectrometer to perform near-infrared spectral transmission scanning on each sample to obtain spectral data, and measure the soluble solid content of each sample with a refractometer; among them, the band range of near-infrared spectral transmission scanning is 700 - 2500 nm, and the number of near-infrared spectra collected for each grape fruit is 3; the soluble solid content is measured repeatedly 2 times;

[0052] (2) Integrate the spectral data, the soluble solid content, and the corresponding origin categories obtained in step (1) to obtain a spectral dataset, and use the method of stratified random sampling to split it into a training set (260 samples) and a test set (74 samples);

[0053] (3) Based on random forest, use R code to identify the soluble solid content characteristics and origin category characteristics of the grape samples through the Boruta algorithm, and calculate their average importance after feature combination;

[0054] (4) Set the importance threshold p to 0.6, screen out the shared features, and reconstruct the training set and the test set according to the shared features to obtain a reconstructed training set and a reconstructed test set;

[0055] (5) Use the reconstructed training set to build an XGBoost integrated learning model, use the Bayes algorithm to globally tune the hyperparameters, and then train with the optimal hyperparameter combination to obtain the final model;

[0056] When using the reconstructed training set to build an XGBoost integrated learning model, evaluate the accuracy of the regression model through the coefficient of determination and the root mean square error. Among them, the coefficient of determination is calculated by formula (I), and the root mean square error is calculated by formula (II), and

[0057]

[0058] where y obs is the actually measured soluble solid content of the grape sample; y pre is the predicted soluble solid content of the grape sample; n is the number of samples of the grape sample;

[0059] When constructing the XGBoost integrated learning model using the reconstructed training set, the accuracy of the model is evaluated by accuracy and recall rate. Among them, based on the confusion matrix, the accuracy is calculated by Equation (III), and the recall rate is calculated by Equation (IV). And

[0060] The confusion matrix is as follows:

[0061] where TP is the positive correct sample; TN is the negative correct sample; FP is the positive incorrect sample; FN is the negative incorrect sample;

[0062] Accuracy (%) = (TP + TN) / (TP + TN + FP + FN) × 100 (III);

[0063] Recall rate (%) = TP / (TP + FN) × 100 (IV);

[0064] (6) Perform prediction based on the final model to complete the detection of grape SSC and origin traceability. Among them, using the reconstructed test set as the input and the regression value of the soluble solid content and the classification result of the origin category of each grape sample as the output;

[0065] Among them, the R code involved includes:

[0066]

[0067]

[0068]

[0069]

[0070]

[0071]

[0072]

[0073]

[0074]

[0075]

[0076]

[0077] The spectral data of the grape samples obtained by near-infrared spectral transmission scanning in step (1) are as Figure 1 shown; the average spectra of grape samples from different origins are asFigure 2 As shown in the figure, it can be seen that there are obvious differences in the spectra between different production areas.

[0078] The selected shared features are as Figure 3 shown. The scatter plot of the SSC of grapes predicted by the final XGBoost model obtained by training with this shared feature is as Figure 4 shown. After calculation, the predicted coefficient of determination R 2 = 0.972, the root mean square error RMSE = 0.277%, and its hyperparameters are mtry = 2; min_n = 8; tree_depth = 3; learn_rate = 0.0107; loss_reduction = 0.171; sample_prop = 0.922. The confusion matrix of the predicted output of the grape production area category is as Figure 5 shown. After calculation, its predicted accuracy is 97.0%, and the recall rate is 97.0%. Its best hyperparameters are mtry = 3; min_n = 12; tree_depth = 2; learn_rate = 0.00293; loss_reduction = 0.166; sample_prop = 0.984.

[0079] Comparative Example 1

[0080] In this comparative example, NIR spectra were used to construct traditional PLS models to predict SSC and production areas respectively. The specific method includes:

[0081] (1) 108, 110, and 116 grape fruits without rot or bruise were collected from Hotan, Turpan, and Kashgar regions in Xinjiang respectively, washed and used as grape samples, and the production area of each grape fruit was marked; a desktop near-infrared spectrometer was used to perform near-infrared spectral transmission scanning on each sample to obtain spectral data, and the soluble solid content of each sample was measured by a refractometer; among them, the band range of the near-infrared spectral transmission scanning was 700 - 2500 nm, and the number of near-infrared spectra collected for each grape fruit was 3; the soluble solid content was measured repeatedly 2 times;

[0082] (2) Integrate the spectral data, the soluble solid content, and the corresponding production area categories obtained in step (1) to obtain a spectral data set, and use the method of stratified random sampling to split it into a training set (260 samples) and a test set (74 samples);

[0083] (3) Use the training set to train the traditional PLS model, and use the prediction set to predict the SSC and production area categories of grapes,

[0084] Among them, the R code used by the PLS model includes:

[0085] ##PLS model of SSC

[0086] traindata<-read.csv(file.choose())

[0087] testdata<-read.csv(file.choose())

[0088] # Distribution of the dependent variable after splitting

[0089] hist(traindata$SSC,breaks=10)

[0090] # Formula expansion for constructing the dependent and independent variables. This is very useful for chaining changes in the dependent variable

[0091] form_reg<-as.formula(

[0092] paste0(

[0093] "VC~",

[0094] paste(colnames(traindata)[1:145],collapse="+") ) )

[0097] form_reg

[0098] # Define the range of parameters to be tuned

[0099] param_grid<-expand.grid(ncomp=1:15)

[0100] # Use the train function in the caret package for grid search and five-fold cross-validation

[0101] pls_models<-lapply(param_grid$ncomp,function(ncomp){

[0102] model<-train(

[0103] form_reg,

[0104] data=traindata,

[0105] method="pls",

[0106] trControl=trainControl(method="cv",number=5),

[0107] tuneGrid = data.frame(ncomp = ncomp),

[0108] metric = "RMSE" )

[0110] return(model)

[0111] )

[0112] # View model results

[0113] pls_models

[0114] # Find the index of the best model

[0115] best_model_index <- which.min(sapply(pls_models, function(model) model$results$RMSE))

[0116] # Get the parameters of the best model

[0117] best_model <- pls_models[[best_model_index]]

[0118] # Output the average results of the cross-validation metrics for the optimal hyperparameters

[0119] cat("Cross-validation results for the best parameters:\n")

[0120] print(best_model$results)

[0121] # Output the specific results of the cross-validation metrics for the optimal hyperparameters

[0122] cat("Cross-validation results for the best parameters:\n")

[0123] print(best_model$resample)

[0124] # Prediction

[0125] # Prediction results on the training set

[0126] trainpred <- predict(best_model,

[0127] newdata <- traindata)

[0128] # Training set prediction error metrics

[0129] defaultSummary(data.frame(obs = traindata$SSC,

[0130] pred = trainpred))

[0131] # Create training set data frame

[0132] trainprediction_df <- data.frame(Actual = traindata$SSC, Predicted = trainpred)

[0133] # Test set prediction results

[0134] testpred <- predict(best_model,

[0135] newdata = testdata)

[0136] # Prediction set prediction error metrics

[0137] defaultSummary(data.frame(obs = testdata$SSC,

[0138] pred = testpred))

[0139] # Create test set data frame

[0140] testprediction_df <- data.frame(Actual = testdata$SSC, Predicted = testpred)

[0141] ## PLS model of Origin

[0142] # Modify variable types

[0143] # Convert categorical variables to factor

[0144] for(i in c()){

[0145] traindata[[i]] <- factor(traindata[[i]])

[0146] }

[0147] # Convert categorical variables to factor

[0148] for (i in c()) {

[0149] testdata[[i]] <- factor(testdata[[i]])

[0150] }

[0151] # Ensure the dependent variable is the place of origin

[0152] traindata$Origin <- factor(traindata$Origin, levels = c("A", "B", "C", "D", "E"))

[0153] # Assume the target variable is 'Origin', and all other columns are independent variables

[0154] form_class <- as.formula("Origin~.") # '.' represents all columns except 'Origin'

[0155] # Define the range of parameters to be tuned

[0156] param_grid <- expand.grid(ncomp = 1:15)

[0157] # Use the train function of the caret package for grid search and five - fold cross - validation

[0158] pls_models <- lapply(param_grid$ncomp, function(ncomp) {

[0159] model <- train(

[0160] form_class,

[0161] data = traindata,

[0162] method = "pls",

[0163] trControl = trainControl(method = "cv", number = 5, classProbs =

[0164] TRUE, summaryFunction = multiClassSummary),

[0165] tuneGrid = data.frame(ncomp = ncomp),

[0166] metric = "Accuracy" # Use accuracy as the evaluation metric )

[0168] return(model)

[0169] [[ID=8}]]

[0170] # Find the index of the best model

[0171] best_model_index <- which.max(sapply(pls_models, function(model) model$results$Accuracy))

[0172] best_model <- pls_models[[best_model_index]]

[0173] # Output the average results of the cross-validation metrics for the optimal hyperparameters

[0174] cat("Cross-validation results for the best parameters:\n")

[0175] print(best_model$results)

[0176]

[0177] The prediction accuracy is shown in Table 1.

[0178] Table 1

[0179]

[0180] Comparative Example 2

[0181] In this comparative example, a PLS model using NIR spectroscopy to construct characteristic wavelengths of the Boruta algorithm is used to predict SSC and origin. The specific method includes:

[0182] (1) 108, 110, and 116 grape fruits without rot or bruise were collected from Hotan, Turpan, and Kashgar regions in Xinjiang respectively, washed and used as grape samples, and the origin of each grape fruit was marked; a desktop near-infrared spectrometer was used to perform near-infrared spectral transmission scans on each sample to obtain spectral data, and the soluble solid content of each sample was measured by a refractometer; among them, the wavelength range of the near-infrared spectral transmission scan was 700 - 2500 nm, and the number of near-infrared spectra collected for each grape fruit was 3; the soluble solid content was measured repeatedly 2 times;

[0183] (2) Integrate the spectral data, the soluble solid content, and the corresponding place of origin category obtained in step (1) to obtain a spectral data set, and use the method of stratified random sampling to split it into a training set (260 samples) and a test set (74 samples);

[0184] (3) Directly screen the characteristic wavelengths of the soluble solid content characteristics and the place of origin category characteristics through the Boruta algorithm, reconstruct the training set and the prediction set to obtain a reconstructed training set and a reconstructed test set;

[0185] (4) Use the reconstructed training set and the reconstructed test set to build a traditional PLS model, and then predict the SSC and the place of origin category of grapes;

[0186] The R code used refers to the Boruta algorithm part in Specific Example 1 and the PLS model part in Comparative Example 1;

[0187] The prediction accuracy is shown in Table 2.

[0188] Table 2

[0189]

[0190] From the above results, it can be seen that the method provided by the present invention is applicable to multi-task analysis, and has the characteristics of high precision and strong applicability, and can overcome the defects of low feature utilization rate and limited precision of traditional single-task models.

[0191] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited thereto. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solutions of the present invention, including any other suitable combination of each technical feature. These simple modifications and combinations should also be regarded as the content disclosed by the present invention and fall within the protection scope of the present invention.

Claims

1. A grape SSC detection and origin traceability method based on NIR multi-task integrated learning, characterized in that, The method includes the following steps: (1) Perform near-infrared spectroscopy transmission scanning on grape samples from different origins to obtain spectral data, and measure the soluble solid content of each sample; wherein, the wavelength range of the near-infrared spectroscopy transmission scanning is 700 - 2500 nm; (2) Integrate the spectral data, the soluble solid content and the corresponding origin categories obtained in step (1) to obtain a spectral dataset, and split it into a training set and a test set; (3) Use the Boruta algorithm to identify the soluble solid content characteristics and origin category characteristics of the grape samples, and calculate their average importance after feature combination; (4) Set an importance threshold to screen out shared features, and reconstruct the training set and the test set according to the shared features to obtain a reconstructed training set and a reconstructed test set; (5) Use the reconstructed training set to build an XGBoost integrated learning model, use the Bayes algorithm to globally tune the hyperparameters, and then train with the optimal hyperparameter combination to obtain the final model; (6) Perform prediction based on the final model to complete grape SSC detection and origin traceability, wherein, using the reconstructed test set as the input, and the regression value of the soluble solid content of each grape sample and the classification result of the origin category as the output.

2. The method according to claim 1, wherein In step (1), the grape samples are selected from at least two of the grapes from three origins: Hotan, Turpan and Kashgar in Xinjiang.

3. The method according to claim 1 or 2, characterized in that, In step (1), the number of each grape sample is not less than 100, and origin marks are made.

4. The method according to claim 1 or 2, characterized in that, In step (1), the near-infrared spectroscopy transmission scanning is performed using a bench-type or hand-held near-infrared spectrometer, and the scanning directions are evenly distributed around the fruit, and the number of near-infrared spectra collected for each grape fruit is not less than 3.

5. The method according to claim 1 or 2, characterized in that, In step (2), the splitting is performed by the method of stratified random sampling, and the sample size of the obtained test set is not less than 20% of the sample size of the spectral dataset.

6. The method according to claim 1 or 2, characterized in that, In step (4), set the importance threshold p = 0.

6.

7. The method according to claim 1 or 2, characterized in that, In step (5), when using the reconstructed training set to build an XGBoost integrated learning model, evaluate the regression model accuracy through the coefficient of determination and the root mean square error. Among them, the coefficient of determination is calculated by formula (I), the root mean square error is calculated by formula (II), and the closer the coefficient of determination is to 1, the smaller the root mean square error, the higher the regression accuracy; and where y obs is the soluble solid content of the grape sample actually measured; y pre is the predicted soluble solid content of the grape sample; n is the number of samples of the grape sample.

8. The method according to claim 1 or 2, characterized in that, In step (5), when using the reconstructed training set to build an XGBoost integrated learning model, evaluate the analysis model accuracy through the accuracy rate and the recall rate. Among them, based on the confusion matrix, the accuracy rate is calculated by formula (III), the recall rate is calculated by formula (IV), and the closer the accuracy rate and the recall rate are to 100%, the higher the classification accuracy; and The confusion matrix is: Wherein, TP is the positive correct sample; TN is the negative correct sample; FP is the positive wrong sample; FN is the negative wrong sample; Accuracy rate (%) = (TP + TN) / (TP + TN + FP + FN) × 100 (III); Recall rate (%) = TP / (TP + FN) × 100 (IV).

Citation Information

Cited By

  • Method for improving portable near infrared spectrum data analysis precision

    CN120563944A