A melon and fruit quality prediction method based on spectral technology, conductivity measurement and machine learning
By combining near-infrared spectroscopy and conductivity measurement into a multimodal fusion method, the problems of limited sample size, long time, and high cost in detecting protein and FAA content in bottle gourd have been solved, enabling rapid and accurate quality assessment and expanding the application range of the detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG UNIV OF TECH
- Filing Date
- 2024-11-14
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies for detecting protein and FAA content in bottle gourds suffer from limitations in sample size, long detection time, and high cost, making it difficult to meet the needs of large-scale production.
By combining near-infrared spectroscopy and conductivity measurement, spectral and conductivity data of bottle gourd were obtained through a multimodal fusion method. A predictive model was then established using machine learning algorithms to achieve rapid and accurate assessment of the protein and FAA content of bottle gourd.
It improves the accuracy and efficiency of detection, reduces costs, is suitable for large-scale application, and can simultaneously process the chemical and physical properties of samples, expanding its application scope to the quality detection of other fruits and vegetables.
Smart Images

Figure CN119782696B_ABST
Abstract
Description
Technical Field
[0001] This patent belongs to the field of fruit and vegetable quality testing, specifically involving a method for predicting fruit and vegetable quality based on spectral technology, conductivity measurement, and machine learning. Background Technology
[0002] Bottle gourd, a cultivated variety of the Cucurbitaceae family, is an annual climbing herbaceous plant that typically flowers in the evening. As a vegetable, bottle gourd fruit is rich in nutrients, containing protein, FAA, vitamins, and trace elements. The flesh of the bottle gourd is delicate in texture, white in color, tender and smooth, with a pleasant aroma and delicious taste, making it popular with consumers. Protein and FAA are important quality indicators for bottle gourd. However, traditional methods for detecting protein and FAA in bottle gourd have limitations such as long detection time, low efficiency, and high cost, making them unsuitable for large-scale testing in production practice. Therefore, there is an urgent need to develop a rapid detection technology specifically for protein and FAA in bottle gourd to achieve rapid, efficient detection and large-scale production.
[0003] In recent years, hyperspectral imaging (HSI), near-infrared spectroscopy (NIRS), ultrasonic testing, and electrical property testing have been widely used for rapid quality detection of fruits and vegetables. Due to its convenience, environmental friendliness, and stability, near-infrared spectroscopy has gradually become the mainstream method for rapid detection of agricultural products. The wavelength range of near-infrared spectroscopy (780–2526 nm) lies between the visible and mid-infrared spectra. By recording the harmonic and combined frequency absorption of hydrogen-containing groups (CH, NH, OH), near-infrared spectroscopy can exhibit different absorption peaks at different wavelengths. When using near-infrared spectroscopy for fruit and vegetable quality analysis, a highly stable and accurate mathematical model should be constructed by mining spectral information data and employing chemometric methods.
[0004] For example, the invention patent with patent number CN111948357A proposes a method for quantitatively evaluating the umami quality of bottle gourd. In terms of FAA detection, it uses a Hitachi L-8900 amino acid automatic analyzer to detect the FAA content of bottle gourd samples. While this method can provide accurate FAA data, its shortcomings include complex and time-consuming detection steps, making rapid assessment difficult. It also requires specialized experimental equipment and consumables, resulting in high detection costs, making it unsuitable for the rapid detection needs of large-scale production.
[0005] The invention patent with patent number CN113484278A proposes a non-destructive testing method for the comprehensive quality of tomatoes based on spectral analysis and principal component analysis. This method acquires visible and near-infrared spectral data of tomato samples and performs feature extraction using principal component analysis to calculate the principal component score data for each tomato sample, establishing a predictive model to predict the overall quality of the tomato samples. While this method can quickly assess tomato quality using spectral data, its limitation lies in its reliance on spectral data and inability to incorporate other physical characteristics for a more comprehensive quality assessment.
[0006] For example, invention patent CN118362608A proposes a device and method for measuring the acidity of fruits and vegetables. This method detects the conductivity of fruit and vegetable juices by measuring a conductivity electrode, and calculates the acidity information by combining data from a TDS (Total Dissolved Solids) electrical signal conversion board and a temperature sensor. This device has the advantages of low cost and fast measurement speed, but its drawback is that relying solely on conductivity to detect fruit and vegetable acidity may ignore the influence of other physical or chemical properties, limiting the comprehensiveness of the test results and its application in other quality tests.
[0007] Relying solely on spectral data has certain limitations. To further improve the accuracy and efficiency of detection, this invention proposes to use conductivity measurement as a supplementary data source. Conductivity data reflects the electrical conductivity of the sample, a physical property that is related to the internal structure and composition of the bottle gourd sample. By combining conductivity data with spectral data and applying multimodal fusion technology, the prediction accuracy of protein and FAA content can be significantly improved.
[0008] The method of this invention is mainly applied to the detection of protein and amino acid (FAA) content inside fruits and vegetables. Specifically, it involves a method based on near-infrared spectroscopy combined with conductivity measurement to obtain spectral and conductivity data of fruits and vegetables, and to predict their internal protein and FAA content through multimodal fusion of the two. For experimental convenience, bottle gourd is used as the experimental subject in this patent. Summary of the Invention
[0009] To address the aforementioned technical problems in existing technologies, this application aims to provide a method for predicting the quality of bottle gourds based on spectral technology, conductivity measurement, and machine learning. This method addresses the limitations of existing technologies in detecting bottle gourd quality, such as limited sample size, long detection time, and high cost. The method rapidly acquires spectral data of bottle gourd samples using near-infrared spectroscopy, simultaneously measures their conductivity data, and combines this with machine learning algorithms to establish a multimodal prediction model for the protein and FAA content of bottle gourds. This efficiently predicts the protein and FAA content of bottle gourds, thereby achieving rapid and accurate assessment of bottle gourd quality.
[0010] The technical concept of this patent is as follows: Currently, existing technologies for detecting bottle gourd quality suffer from problems such as limited sample size, long detection time, and high cost. Therefore, this patent proposes to utilize NIRS (Near Infrared Spectroscopy) combined with conductivity measurement and machine learning algorithms to further enhance the accuracy and efficiency of detecting bottle gourd protein and FAA content, thereby achieving a comprehensive assessment of bottle gourd quality.
[0011] The technical solution adopted by this patent to achieve the above-mentioned inventive objective is as follows:
[0012] A method for predicting the quality of melons and fruits based on spectral technology, conductivity measurement, and machine learning includes the following steps:
[0013] S1: Harvest fresh and tender melons and fruits at the ripe stage. After the melons and fruits are freeze-dried, they are ground into powder using a high-speed grinder. The melon and fruit powders are grouped and numbered to establish multiple core germplasm populations of melons and fruits with individual numbers.
[0014] S2: For the fruit samples of each germplasm population in step S1, detect their near-infrared spectral data and electrical conductivity data, as well as nutritional data including protein content and FAA content, to obtain the corresponding dataset.
[0015] S3: Preprocess the near-infrared spectral data and conductivity data in the dataset from step S2 to eliminate noise bias;
[0016] S4: Next, feature wavelength selection is performed on the preprocessed near-infrared spectral data. The absorbance and conductivity data under the selected feature wavelengths are merged into a dataset, and the dataset is divided into a test set and a training set.
[0017] S5: Construct a regression prediction model for fruit protein and FAA using a ridge regression model. Based on the test set and training set divided in step S3, train and test the constructed prediction model. When detecting the content of fruit protein and FAA, first obtain near-infrared spectral data and conductivity data according to the method in steps S2-S3. After preprocessing and feature selection, input them into the trained prediction model to obtain the corresponding fruit protein and FAA content data.
[0018] Further, the fruit mentioned in step S1 is bottle gourd. Uniform, tender fruits are harvested during the commercial maturity period of the bottle gourd. Nine bottle gourd fruits at commercial maturity are collected from each germplasm population, and these nine fruits are divided into three groups, with three fruits in each group. Samples are taken from the middle part of each group of fruits. The samples are freeze-dried in a freeze dryer (LGJ-10) for 3-4 days and ground into powder using a high-speed grinder (JXFSTPRP-24L).
[0019] Furthermore, in step S2, the fruit samples of each germplasm population are divided into three parts for testing. The first part is to detect protein data, the second part is to analyze FAA content using a Hitachi L-8900 amino acid analyzer, and the third part is to detect near-infrared spectral data and conductivity data. The near-infrared spectral data is obtained by scanning with a near-infrared analyzer, and the conductivity data detection process is as follows: the fruit powder is pressed into tablets, and then the conductivity of the tableted sample is measured using a precision conductivity meter.
[0020] Furthermore, the steps for acquiring near-infrared spectral data through scanning with a near-infrared analyzer are as follows:
[0021] 1) To obtain the near-infrared spectrum of bottle gourd, this invention used a Thermo Nicolet ANTARIS II Fourier transform near-infrared (FT-NIR) analyzer to scan the third part of the samples from each numbered core germplasm population in step S2. Scanning was performed under constant temperature (24℃) and constant humidity (60%) conditions. The spectral range was 1000 to 2500 nm, with each sample scanned 64 times at a resolution of 8 cm⁻¹. -1 .
[0022] 2) Place the samples into sample cups with a diameter of 5 cm and a height of approximately 1.5 cm. Maintain consistent sample thickness and compaction to minimize measurement errors caused by uneven sample loading. Scan each sample three times and calculate the average absorption spectrum as the spectral value. Near-infrared spectral information for all samples is collected and stored.
[0023] Further, the steps for measuring the conductivity of the samples are as follows: The conductivity of the compressed samples is measured using a precision conductivity meter. The conductivity meter employs a four-probe method, applying current through external electrodes while measuring voltage changes within the sample through internal electrodes. This reduces the influence of contact resistance on the measurement results and improves measurement accuracy. The conductivity data for each sample is recorded. To ensure the stability of the results, each sample is measured at least three times, and the average value is taken as the final conductivity.
[0024] Furthermore, the near-infrared spectral data mentioned in step S2 is curve data composed of n data points in a rectangular coordinate system, where the horizontal and vertical coordinates of the data points are wavelength and absorbance values, respectively.
[0025] Further, step S3 preprocesses the near-infrared spectral data. To eliminate background noise and interference caused by instrument and position changes during data acquisition, it is necessary to preprocess the raw spectra. Preprocessing the raw spectral data using multiple scattering correction (MSC) and standard normalized variable (SNV) methods can eliminate noise and baseline drift interference. Multiple scattering correction corrects for the effects of multiple scattering by constructing a ratio between the sample spectrum and the reference spectrum. Multiple scattering leads to an increase in optical path length and a decrease in signal intensity. The basic formula for MSC correction is as follows:
[0026] MSC=(R sample -R min ) / (R ref -R min ), (1)
[0027] In the above formula, R sample R is the absorbance measurement value of an original spectral curve to be corrected at a data point at wavelength i. min R is the minimum value of the absorbance values at different wavelengths of an original spectral curve to be corrected. ref The MSC value is the average absorbance of all original spectral curves at the same wavelength i. The MSC value is the absorbance correction value of one original spectral curve to be corrected at wavelength i. i is an integer from 1 to n. The purpose of MSC correction is to remove residual scattering components in the spectrum and improve the accuracy and comparability of sample spectra.
[0028] Variable normalization (SNV) is a normalization technique used to eliminate differences caused by variations in light intensity and baseline drift. The SNV formula is as follows:
[0029] SNV = (X i -μ) / σ, (2)
[0030]
[0031] Among them, X i σ is the absorbance test value of an original spectral curve to be corrected at a data point at wavelength i, μ is the average absorbance value of the original spectral curve to be corrected at different wavelengths, and σ is the standard deviation of the absorbance value of the original spectral curve to be corrected at different wavelengths.
[0032] The SNV correction value is the absorbance correction value of the original spectral curve to be corrected at wavelength i. For each spectrum, the SNV method subtracts the mean from each data point in the spectrum and then divides by the standard deviation. In this way, the overall intensity variation and baseline drift in the spectrum are eliminated, thus highlighting the spectral features.
[0033] Further, in step S3, the conductivity data for each sample is standardized. Since conductivity data and spectral data have different magnitudes, they must be normalized to facilitate model training after being combined with the spectral data. The standardization formula is as follows:
[0034]
[0035] Where Z is the conductivity correction value of the sample to be corrected, Y is the conductivity test value of the sample to be corrected, α is the mean conductivity of all samples, and γ is the standard deviation of the conductivity data of all samples. This method standardizes the conductivity data to a similar order of magnitude as the spectral data, which facilitates subsequent fusion and analysis.
[0036] Furthermore, the specific steps of step S4 are as follows:
[0037] 1) Using the protein content and amino acid FAA content data as evaluation results, respectively, the competitive adaptive reweighted sampling algorithm CARS was used to extract the characteristic wavelengths of the preprocessed near-infrared spectral data, and the characteristic wavelengths of the near-infrared spectral data identified by protein and FAA were obtained respectively.
[0038] Competitive Adaptive Reweighted Sampling (CARS), based on Monte Carlo sampling and partial least squares (PLS) regression coefficients, is widely used for feature wavelength optimization of near-infrared spectral data. CARS is used to extract feature wavelengths from preprocessed near-infrared spectral data. In each iteration (iterations are used to compare different sampling results, continuously filtering and optimizing; through multiple iterations, CARS can find the feature wavelengths that contribute the most to the prediction model and have the least redundancy), the calibration set samples are randomly reselected using the exponential decay function EDF (see the paper https: / / www.sciencedirect.com / science / article / pii / S0003267009008332) to facilitate the selection of feature wavelengths. The data is optimized using an adaptive reweighting method (using the PLS model to calculate RMSECV, optimizing wavelength selection based on regression coefficients). The subset with the smallest root mean square error in cross-validation is selected as the feature wavelengths. This algorithm is implemented in Python using the sklearn library and extracts feature wavelengths from preprocessed near-infrared spectral data.
[0039] Using Competitive Adaptive Reweighted Sampling (CARS) for feature wavelength extraction from near-infrared spectral data is an existing technology. For example, Chinese patent CN116465855A describes in paragraph 0003 of the background section that "Competitive Adaptive Reweighted Sampling (CARS) is an effective variable selection algorithm widely used in feature selection of near-infrared spectral data. CARS proposes to select the most relevant combination of variables (feature wavelengths) in a continuous selection process. Based on the regression coefficients obtained from the Partial Least Square (PLS) model, CARS iteratively selects N subsets of variables from N Monte Carlo sampling processes. In each sampling process, a fixed proportion of sample data is randomly selected to build a calibration model. Next, using the obtained regression coefficients, a two-step variable selection procedure is used to select relevant wavelengths. Finally, cross-validation is used to select the subset (most relevant wavelength combination) with the lowest root mean square error."
[0040] 2) Based on the selected characteristic wavelengths of the spectral data, the conductivity data is added as an additional physical feature and input into the multimodal analysis model along with the selected characteristic wavelengths of the spectral data.
[0041] Assuming the absorbance values of m spectral characteristic wavelengths are extracted, they are expressed as:
[0042] x spectral =[x1,x2,...,x m (4)
[0043] Each sample corresponds to a conductivity value d. The conductivity data is used as an additional feature and concatenated with the spectral feature vector to form a fused feature vector:
[0044] X fuesd =[x1,x2,...,x m ,d] (5)
[0045] Protein recognition fusion feature vector X fused The protein content is passed as input to the regression model, which outputs a predicted value.
[0046] FAA identifies the fused feature vector X fused The data is passed as input to the regression model, which outputs a predicted FAA content value.
[0047] y = f(X) fuesd (6)
[0048] Where f is the regression model, such as ridge regression or support vector machine, and y is the prediction result of the model.
[0049] Introducing conductivity data can help improve the model's understanding of the physical and chemical properties of samples, compensating for the limitations of single spectral information. Near-infrared spectral data after extracting characteristic wavelengths and conductivity data were randomly divided into training and testing sets at an 8:2 ratio. Consistency between the conductivity and spectral data during the partitioning process was maintained to ensure the integrity of modal information during training and testing.
[0050] To predict the protein and FAA content of bottle gourd samples using spectral information and conductivity data, several common regression models were selected for experiments. These models are used to establish the relationship between inputs (spectral information and conductivity data) and outputs (protein and FAA content), thereby enabling prediction of unknown samples. Commonly used methods include ridge regression, random forest regression, support vector machine regression, and fully connected neural networks. Ridge regression is an extension of linear regression, addressing multicollinearity by imposing constraints on the regression coefficients. Ridge regression performs better on relatively small datasets. When the dataset is small, random forests, neural networks, and support vector machines may be too flexible, prone to overfitting, and less responsive to variable correlations. Therefore, considering the high correlation between feature wavelengths and the limited data sample, we adopted a ridge regression-based approach to construct the prediction model. The training dataset and training labels were shaped to match the input requirements of the ridge regression model, and the regularization parameter alpha was set to 0.1.
[0051] Furthermore, the training set in S5 is used to train the prediction model, and the test set is used to test the model's performance; and the determination coefficient R between the predicted and actual values on the test set is used as the basis for the prediction. 2 RMSE is used as an indicator to determine the effectiveness of the model.
[0052]
[0053] In the coefficient of determination R 2 In the formula for root mean square error (RMSE), y i This represents the true value of the i-th sample. This represents the predicted value of the i-th sample. is the mean of the true values of all samples, and m is the number of samples.
[0054] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0055] (1) This patent utilizes near-infrared spectroscopy combined with conductivity measurement to simultaneously acquire spectral and conductivity data of bottle gourd samples through a multimodal fusion method, enabling rapid and accurate prediction of protein and FAA content in bottle gourd samples. Compared to traditional detection methods, combining spectral information and conductivity data further improves the accuracy and stability of prediction. This method requires lower equipment and consumable costs, is easy to operate, and is suitable for large-scale application, significantly reducing detection costs. Simultaneously, automated data processing and model training improve detection efficiency and reduce manual intervention.
[0056] (2) This patent, through multimodal analysis combining spectral and conductivity data, can simultaneously process the chemical and physical properties of samples. Through reasonable experimental design and data processing methods, the system can process multiple samples simultaneously, increasing detection throughput and meeting the needs of large-scale production and research. Conductivity data provides an additional dimension to the physical properties of samples, enhancing the predictive ability for protein and FAA content, and further improving the robustness and predictive performance of the model.
[0057] (3) This method is not only applicable to the quality detection of bottle gourds, but can also be extended to the quality detection of other fruits and vegetables, especially in applications that combine spectral and conductivity data for multimodal analysis. By fusing conductivity data with spectral data, the detection accuracy of different types of fruits and vegetables can be improved, expanding the application scope of this patent in other agricultural products. Attached Figure Description
[0058] Figure 1 This is a flowchart of the method of this patent;
[0059] Figure 2 This is the original near-infrared spectral data image for this patent;
[0060] Figure 3a This is a diagram showing the preprocessing results of the near-infrared spectral data of this patent based on the SNV method;
[0061] Figure 3b This is a diagram showing the preprocessing results of the near-infrared spectral data of this patent based on the MSC method;
[0062] Figure 4a The image shows the results of extracting characteristic wavelengths for protein recognition from the near-infrared spectral data of this patent.
[0063] Figure 4b The image shows the results of extracting characteristic wavelengths from the near-infrared spectral data of this patent for identifying the content of amino acids (FAA). Detailed Implementation
[0064] The specific embodiments of this patent will be described in further detail below with reference to the accompanying drawings.
[0065] Reference Figures 1 to 4b A method for predicting the quality of bottle gourd based on spectral technology, conductivity measurement, and machine learning includes the following steps:
[0066] S1: Preparation of bottle gourd samples, specifically including:
[0067] A total of 206 core germplasm populations of bottle gourd were selected as experimental materials. Uniform and tender fruits were harvested at the commercial maturity stage. Nine fruits from each germplasm population at commercial maturity were collected and divided into three groups of three fruits each. Samples were taken from the middle portion of each group of fruits. The samples were freeze-dried in a freeze dryer (LGJ-10) for 3-4 days and then ground into powder using a high-speed grinder (JXFSTPRP-24L).
[0068] The fruit samples from each germplasm population were divided into three parts for testing. The first part was to detect protein data. The second part was to analyze the FAA content using a Hitachi L-8900 amino acid analyzer. The third part was to detect near-infrared spectral data and conductivity data. The near-infrared spectral data was obtained by scanning with a near-infrared analyzer. The conductivity data detection process was as follows: the fruit powder was pressed into tablets, and then the conductivity of the tableted sample was measured using a precision conductivity meter.
[0069] S2: Determine the protein and FAA content of bottle gourd samples using specific instruments, including:
[0070] Near-infrared spectral data of bottle gourd samples were obtained by scanning them with a near-infrared analyzer, including:
[0071] To obtain the near-infrared spectrum of bottle gourd, this invention used a Thermo Nicolet ANTARIS II Fourier transform near-infrared (FT-NIR) analyzer to scan the samples. Scanning was performed under constant temperature (24°C) and constant humidity (60%) conditions. The spectral range was 1000 to 2500 nm, with each sample scanned 64 times at a resolution of 8 cm⁻¹. -1 The samples were placed into sample cups with a diameter of 5 cm and a height of approximately 1.5 cm. The thickness and compaction of the samples were kept consistent to minimize measurement errors caused by uneven sample loading. Each sample was scanned three times, and the average absorption spectrum was calculated as the spectral value. Near-infrared spectral information for all samples was collected and stored.
[0072] The stored near-infrared spectral data can be visualized, such as... Figure 2 As shown, the near-infrared (NIR) reflectance spectra of the 206 core germplasm populations of *Gnaphalium affine* range from 4000 to 10000 cm⁻¹. -1 Between these values, the spectral trends of all samples were similar. Regarding reflectance, at 4698 cm⁻¹...-1 5102cm -1 and 6697cm -1 There is a significant absorption peak at this point.
[0073] S3: Measurement of electrical conductivity data for bottle gourd samples, specifically including:
[0074] The conductivity of the compressed samples was measured using a precision conductivity meter. The conductivity meter employed a four-probe method, applying current through external electrodes while measuring voltage changes within the sample through internal electrodes. This reduces the influence of contact resistance on the measurement results and improves accuracy. Conductivity data for each sample was recorded. To ensure the stability of the results, each sample was measured at least three times, and the average value was taken as the final conductivity.
[0075] Electrical conductivity data reflects the internal physical properties of a sample, particularly those related to conductivity and ion movement. Combined with spectral data, conductivity can provide additional sample information, helping to improve the accuracy of predicting protein and FAA content.
[0076] S4: Perform preprocessing operations on the near-infrared spectral data and conductivity data obtained by scanning with a near-infrared analyzer, specifically including:
[0077] To eliminate background noise and interference caused by instrument and location changes during data acquisition, it is necessary to preprocess the raw spectra. Preprocessing the raw spectral data using multiple scattering correction (MSC) and standard normalized variable (SNV) methods can eliminate noise and baseline drift interference.
[0078] The multiple scattering correction method corrects the effects of multiple scattering by constructing the ratio of the sample spectrum to the reference spectrum. Multiple scattering leads to an increase in optical path length and a decrease in signal intensity.
[0079] The basic formula for MSC calibration is as follows:
[0080] MSC=(R sample -R min ) / (R ref -R min ), (1)
[0081] In the above formula, R sample R is the absorbance measurement value of an original spectral curve to be corrected at a data point at wavelength i. min R is the minimum value of the absorbance values at different wavelengths of an original spectral curve to be corrected. refThe MSC value is the average absorbance of all original spectral curves at the same wavelength i. The MSC value is the absorbance correction value of one original spectral curve to be corrected at wavelength i. i is an integer from 1 to n. The purpose of MSC correction is to remove residual scattering components in the spectrum and improve the accuracy and comparability of sample spectra.
[0082] Variable normalization (SNV) is a normalization technique used to eliminate differences caused by variations in light intensity and baseline drift. The SNV formula is as follows:
[0083] SNV = (X i -μ) / σ, (2)
[0084]
[0085] Among them, X i σ is the absorbance test value of an original spectral curve to be corrected at a data point at wavelength i, μ is the average absorbance value of the original spectral curve to be corrected at different wavelengths, and σ is the standard deviation of the absorbance value of the original spectral curve to be corrected at different wavelengths.
[0086] The SNV correction value is the absorbance correction value of the original spectral curve to be corrected at wavelength i. For each spectrum, the SNV method subtracts the mean from each data point in the spectrum and then divides by the standard deviation. In this way, the overall intensity variation and baseline drift in the spectrum are eliminated, thus highlighting the spectral features.
[0087] By comparing the raw near-infrared data, both MSC and SNV effectively reduce noise, but their trends are not significantly different. Specific results are as follows: Figures 3a-3b As shown, the subsequent experimental steps of this patent use data processed by MSC as the experimental results.
[0088] The conductivity data for each sample were standardized. Since conductivity data and spectral data have different orders of magnitude, they must be normalized to facilitate model training after being combined with the spectral data. The standardization formula is as follows:
[0089]
[0090] Where Z is the conductivity correction value of the sample to be corrected, Y is the conductivity test value of the sample to be corrected, α is the mean conductivity of all samples, and γ is the standard deviation of the conductivity data of all samples. This method standardizes the conductivity data to a similar order of magnitude as the spectral data, which facilitates subsequent fusion and analysis.
[0091] S5: The preprocessed near-infrared spectral data obtained from S4 is used to select characteristic wavelengths. The absorbance and conductivity data at the selected characteristic wavelengths are then fused into a single dataset, which is subsequently divided into a test set and a training set at an 8:2 ratio. Specifically, this includes:
[0092] Competitive Adaptive Reweighted Sampling (CARS), based on Monte Carlo sampling and partial least squares (PLS) regression coefficients, is widely used for feature wavelength optimization of near-infrared spectral data. CARS is used to sample preprocessed near-infrared spectral data. In each iteration, a calibration set sample is randomly reselected using an exponential decay function to facilitate the selection of feature wavelengths. The data is optimized using an adaptive reweighting method. The subset with the smallest root mean square error in cross-validation is selected as the feature wavelengths. This algorithm is implemented in Python using the sklearn library and extracts feature wavelengths from preprocessed near-infrared spectral data.
[0093] In this study, CARS was used to extract significant characteristic wavelengths from near-infrared spectral data, with the extracted characteristic wavelengths mainly concentrated in the 4000 to 5000 cm⁻¹ range. -1 and 7000 to 10000 cm -1 The number of characteristic wavelengths is significantly reduced compared to the original near-infrared spectral data. By applying the CARS algorithm, we identified 127 highly correlated characteristic wavelengths for the protein, and renumbered these characteristic wavelengths as integers from 1 to 127. The result of extracting characteristic wavelengths from the near-infrared spectral data using protein identification is shown in the figure. Figure 4a In this study, by applying the CARS algorithm, we identified 124 highly correlated feature wavelengths for the FAA (Front-Infrared Spectroscopy). These feature wavelengths were then renumbered as integers from 1 to 124. The results of extracting the feature wavelengths from the near-infrared spectral data using the FAA-identified wavelengths are shown in the figure below. Figure 4a middle.
[0094] Figures 4a-4b The distribution of these characteristic wavelengths is shown, with the horizontal axis representing the wavelength number and the vertical axis representing the spectrum.
[0095] In addition, from Figures 4a-4b As shown by the characteristic wavelength distribution represented by the red line, the characteristic wavelength distribution obtained based on the CARS algorithm is relatively uniform. The CARS algorithm selects and extracts the most representative features through a competitive process, without explicitly favoring any specific wavelength.
[0096] Based on the selected spectral characteristic wavelengths, conductivity data is added as an additional physical feature and input into the multimodal analysis model along with the selected spectral characteristic wavelengths. Assuming that m absorbance values are extracted for each spectral characteristic wavelength, they are expressed as:
[0097] Xspectral =[x1,x2,...,x m (4)
[0098] Each sample corresponds to a conductivity value d.
[0099] The conductivity data is treated as an additional feature and concatenated with the spectral feature vector to form a fused feature vector.
[0100] X fused =[x1,x2,...,x m ,d] (5)
[0101] Protein recognition fusion feature vector X fused The protein content is passed as input to the regression model, which outputs a predicted value.
[0102] FAA identifies the fused feature vector X fused The data is passed as input to the regression model, which outputs a predicted FAA content.
[0103] y = f(X) fuesd (6)
[0104] Where f is the regression model (such as ridge regression, support vector machine, etc.), and y is the prediction result of the model.
[0105] Introducing conductivity data can help improve the model's understanding of the physical and chemical properties of samples, making up for the shortcomings that may arise from single spectral information.
[0106] Near-infrared spectral data after extracting characteristic wavelengths and conductivity data were randomly divided into training and testing sets at an 8:2 ratio. A total of 164 training samples and 42 testing samples were used. Consistency was maintained between the conductivity and spectral data during the partitioning process to ensure the integrity of modal information during training and testing.
[0107] S5: Construct a regression prediction model for bottle gourd protein and FAA using ridge regression. Specifically, this includes:
[0108] To predict the protein and FAA content of bottle gourd samples using spectral information and conductivity data, several common regression models were selected for experiments. These models are used to establish the relationship between inputs (spectral information and conductivity data) and outputs (protein and FAA content), thereby enabling prediction of unknown samples. Commonly used methods include ridge regression, random forest regression, support vector machine regression, ordinary least squares, and fully connected neural networks. Ridge regression is an extension of linear regression, addressing multicollinearity by imposing constraints on the regression coefficients. Ridge regression performs better on relatively small datasets. When the dataset is small, random forests, neural networks, and support vector machines may be too flexible, prone to overfitting, and less responsive to variable correlations.
[0109] Therefore, considering the high correlation between feature wavelengths and the limited data sample, we adopted a ridge regression-based approach to construct the prediction model. The training dataset and training labels were shaped to match the input requirements of the ridge regression model, and the regularization parameter alpha was set to 0.1.
[0110] S6: Based on the prediction model in S5, input the spectral data and conductivity data obtained from near-infrared analyzer scanning to predict the protein and FAA content of bottle gourd. Specifically, this includes:
[0111] The training set from S4 is used to train the prediction model in S5, and the test set from S4 is used to test the performance of the model in S5. The performance is then assessed based on the coefficient of determination R between the predicted and actual values on the test set. 2 RMSE is used as an indicator to determine the effectiveness of the model.
[0112]
[0113] As shown in Tables 1 and 2, the protein and FAA content prediction models based on the ridge regression algorithm have a good fit with the observed data (test set), and the predicted values closely match the true values along the diagonal. Furthermore, compared with random forests, support vector machines, and fully connected neural networks, ridge regression exhibits faster speed and higher R-values. 2 Value. In protein content prediction, the R-value of the ridge regression-based model. 2 The value was 0.957, and the RMSE value was 0.228, which was 15.5% higher than the second place. Furthermore, in the prediction of FAA content, its R... 2 The score was 0.766, and the RMSE score was 0.523, which was 26.4% better than the second place.
[0114] Table 1 Results of the protein prediction model
[0115]
[0116] Table 2 Results of the FAA Prediction Model
[0117]
[0118] The contents described in the embodiments of this specification are merely examples of implementation forms of the patent concept. The scope of protection of this patent should not be regarded as limited to the specific forms stated in the embodiments. The scope of protection of this patent also extends to equivalent technical means that can be conceived by those skilled in the art based on the concept of this invention.
Claims
1. A method for predicting the quality of melons and fruits based on spectral technology, conductivity measurement, and machine learning, characterized in that... Includes the following steps: S1: Harvest fresh and tender melons and fruits at the ripe stage. After the melons and fruits are freeze-dried, they are ground into powder using a high-speed grinder. The melon and fruit powders are grouped and numbered to establish multiple core germplasm populations of melons and fruits with individual numbers. S2: For the fruit samples of each germplasm population in step S1, detect their near-infrared spectral data and electrical conductivity data, as well as nutritional data including protein content and FAA content, to obtain the corresponding dataset. S3: Preprocess the near-infrared spectral data and conductivity data in the dataset from step S2 to eliminate noise bias; S4: Next, feature wavelength selection is performed on the preprocessed near-infrared spectral data. The absorbance and conductivity data under the selected feature wavelengths are merged into a dataset, and the dataset is divided into a test set and a training set. S5: Construct a regression prediction model for fruit protein and FAA using a ridge regression model. Based on the test set and training set divided in step S3, train and test the constructed prediction model. When detecting the content of fruit protein and FAA, first obtain near-infrared spectral data and conductivity data according to the method in steps S2-S3. After preprocessing and feature selection, input them into the trained prediction model to obtain the corresponding fruit protein and FAA content data. The specific steps for step S4 are as follows: 1) Using the protein content and amino acid FAA content data as evaluation results respectively, the competitive adaptive reweighted sampling algorithm CARS was used to extract the characteristic wavelengths of the preprocessed near-infrared spectral data, and the characteristic wavelengths of the near-infrared spectral data identified by protein and FAA were obtained respectively. 2) Based on the selected characteristic wavelengths of the spectral data, the conductivity data is added as an additional physical feature and input into the multimodal analysis model along with the selected characteristic wavelengths of the spectral data. Assuming the absorbance values of m spectral characteristic wavelengths are extracted, they are expressed as: ; Each sample corresponds to a conductivity value d. The conductivity data is used as an additional feature and concatenated with the spectral feature vector to form a fused feature vector: ; Protein recognition fusion feature vector X fused The protein content is passed as input to the regression model, which outputs a predicted value. FAA identifies the fused feature vector X fused The data is passed as input to the regression model, which outputs a predicted FAA content value. The ridge regression model described in step S5 is used to establish the relationship between the input layer data and the output layer data, thereby enabling the prediction of unknown samples. The input layer data is the feature vector X fused from protein recognition or FAA recognition. fused The output layer data is protein content or FAA content.
2. The method for predicting fruit quality based on spectral technology, conductivity measurement, and machine learning as described in claim 1, characterized in that... The melon and fruit mentioned in step S1 are bottle gourds. Nine bottle gourd fruits at the commercial maturity stage are collected from each germplasm population. These nine fruits are divided into three groups, with three fruits in each group. Samples are taken from the middle part of each group of fruits. The samples are freeze-dried in a freeze dryer for 3-4 days and then ground into powder using a high-speed grinder.
3. The method for predicting fruit quality based on spectral technology, conductivity measurement, and machine learning as described in claim 1, characterized in that... In step S2, the fruit samples of each germplasm population are divided into three parts for testing. The first part is to detect protein data. The second part is to analyze the FAA content using a Hitachi L-8900 amino acid analyzer. The third part is to detect near-infrared spectral data and conductivity data. The near-infrared spectral data is obtained by scanning with a near-infrared analyzer. The conductivity data detection process is as follows: the fruit powder is pressed into tablets, and then the conductivity of the tableted sample is measured using a precision conductivity meter.
4. The method for predicting fruit quality based on spectral technology, conductivity measurement, and machine learning as described in claim 1, characterized in that... The near-infrared spectral data mentioned in step S2 is curve data composed of n data points in a rectangular coordinate system, where the horizontal and vertical coordinates of the data points are wavelength and absorbance values, respectively. Step S3 involves preprocessing the near-infrared spectral data using one of the following two methods: 1) By applying the multiple scattering correction (MSC) method to preprocess the raw spectral data, noise and baseline drift interference can be eliminated. The basic formula for MSC correction is as follows: ; In the above formula R sample It is the absorbance measurement value of an original spectral curve to be corrected at a data point at wavelength i. R min It is the minimum value of the absorbance at different wavelengths compared to the original spectral curve to be corrected. R ref It is the average absorbance of all original spectral curves at the same wavelength i, and the MSC value is the absorbance correction value of one original spectral curve to be corrected at wavelength i. i is an integer from 1 to n; 2) The original spectral data is preprocessed using the Standard Normalized Variable (SNV) method to eliminate differences caused by variations in light intensity and baseline drift. The basic formula for SNV correction is as follows: ; Among them, X i is the absorbance measurement value of an original spectral curve to be corrected at a data point at wavelength i, and µ is the average absorbance value of the original spectral curve to be corrected at different wavelengths. σ It is the standard deviation of the absorbance values at different wavelengths of an original spectral curve to be corrected; The SNV correction value is the absorbance correction value of an original spectral curve to be corrected at wavelength i. The standard deviation σ The calculation formula is as follows: 。 5. The method for predicting fruit quality based on spectral technology, conductivity measurement, and machine learning as described in claim 1, characterized in that... Step S3 involves standardizing the conductivity data to facilitate model training after combining it with the spectral data. The standardization formula is as follows: ; Where Z is the conductivity correction value of the sample to be corrected, and Y is the conductivity test value of the sample to be corrected. It is the average conductivity of all samples. It is the standard deviation of the conductivity data of all samples. This method standardizes the conductivity data to a similar order of magnitude as the spectral data, which facilitates subsequent fusion and analysis.
6. The method for predicting fruit quality based on spectral technology, conductivity measurement, and machine learning as described in claim 1, characterized in that... The training set in S5 is used to train the prediction model, and the test set is used to test the model's performance; the performance is then determined based on the coefficient of determination between the predicted and actual values on the test set. And RMSE as an indicator to determine the effectiveness of the model: ; ; In the coefficient of determination R 2 In the formula for root mean square error (RMSE), This represents the true value of the i-th sample. This represents the predicted value of the i-th sample. is the mean of the true values of all samples, and m is the number of samples.