Method for detecting organic acid content of bergamot pears based on near infrared spectrum technology

By employing near-infrared spectroscopy and a multi-step pretreatment method, the problems of narrow applicability and poor stability in the detection of organic acid content in fragrant pears have been solved, achieving efficient and accurate detection results. This method is applicable to fragrant pears of different maturity levels and varieties, and is suitable for orchard harvesting grading and market quality supervision.

CN121830554APending Publication Date: 2026-04-10XINJIANG GUANNONG FRUIT & ANTLER GROUP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing methods for detecting organic acid content in fragrant pears have a narrow scope of application, are prone to deviations in test results, have poor stability, and traditional methods cannot guarantee the reliability of the model, making it difficult to meet the needs of large-scale and diversified testing.

Method used

Near-infrared spectroscopy was employed, and a multi-step preprocessing, model building, and validation method was used, including homogenization, multivariate scattering correction, baseline correction, smoothing, and chemometric model construction. Combined with partial least squares regression and leave-one-out cross-validation, detection parameters and environmental control were optimized to construct a detection model for Korla pears with different maturity levels and varieties.

Benefits of technology

It achieves efficient and accurate detection of organic acid content in fragrant pears, has a wide range of applications, standardized detection procedures, requires no complicated chemical pretreatment, and is suitable for orchard harvesting and grading as well as market quality supervision, thus improving the accuracy and stability of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121830554A_ABST
    Figure CN121830554A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of fruit spectrum detection, and discloses a method for detecting the organic acid content of bergamot pears based on a near infrared spectrum technology, which comprises the following steps: step 1, sample preparation: selecting bergamot pears with different maturity degrees and different varieties as detection samples, manually removing fruit stems and kernels, and detecting the organic acid content of the bergamot pears; putting the pulp into homogenizing equipment for crushing and homogenizing treatment; step 2, spectrum acquisition: optimizing near infrared spectrum detection parameters; step 3, spectrum pretreatment: adopting a mode of combining multiplicative scatter correction, baseline correction and smoothing treatment; step 4, establishing a model: selecting a part of samples to be detected in the step 1, and determining the organic acid content of the samples by adopting a traditional standard detection method; and 5, model verification: collecting different batches of bergamot pear blind samples from different sources. Through multi-sample preparation, spectral parameter optimization, combined pretreatment and closed-loop modeling verification, efficient and accurate detection of organic acid in bergamot pears is realized, the application range is widened, and the process is simplified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fruit spectral detection technology, specifically a method for detecting the organic acid content of fragrant pears based on near-infrared spectroscopy. Background Technology

[0002] Agricultural product quality testing is a crucial area for ensuring the quality and safety of agricultural products and the standardization of market circulation. The organic acid content of fragrant pears is a key indicator for assessing their flavor and ripeness, and its test results directly affect the grading, harvesting, storage, preservation, and market pricing of fragrant pears. Therefore, developing efficient and accurate methods for detecting the organic acid content of fragrant pears has significant practical application value in the field of agricultural product quality evaluation and supervision.

[0003] Current methods for detecting organic acid content in Korla pears largely rely on traditional chemical detection techniques. To improve applicability, a common practice is to develop specific testing procedures for Korla pears at a single maturity level or variety. However, this approach is difficult to adapt to Korla pears at different growth stages, resulting in a narrow scope of application and failing to meet the needs of large-scale, diversified testing. To improve accuracy, existing technologies often employ a single spectral preprocessing method to eliminate detection interference. However, a single preprocessing method cannot comprehensively remove various interfering factors during spectral acquisition, making the test results prone to deviation and exhibiting poor stability. Furthermore, to simplify the operation process, some detection methods omit the validation step of the detection model. While this shortens the testing time, it cannot guarantee the reliability of the model, leading to insufficient accuracy and failing to meet the precision requirements of actual testing. Therefore, this paper proposes a method for detecting organic acid content in Korla pears based on near-infrared spectroscopy. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a method for detecting the organic acid content of pears based on near-infrared spectroscopy, thereby solving the problems mentioned in the background section.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for detecting the organic acid content of pears based on near-infrared spectroscopy, comprising the following steps:

[0006] Step 1, Sample Preparation: Select pears of different maturity and varieties as test samples. After manually removing the fruit stems and cores, place the pulp in a homogenizing device for crushing and homogenization to prepare a uniformly dispersed test sample.

[0007] Step 2, Spectral Acquisition: Optimize near-infrared spectral detection parameters, and use a near-infrared spectrometer to acquire spectral data of the sample prepared in Step 1 to obtain the raw near-infrared spectrum;

[0008] Step 3: Spectral preprocessing: The near-infrared raw spectral data obtained in Step 2 is preprocessed using a combination of multivariate scattering correction, baseline correction, and smoothing to eliminate interference factors.

[0009] Step 4: Model Establishment: Select some of the samples to be tested in Step 1 and determine their organic acid content using traditional standard detection methods. Correlate the determination results with the spectral data after preprocessing in Step 3, and construct a quantitative analysis model by combining chemometric methods.

[0010] Step 5: Model Validation: Collect blind samples of pears from different batches and sources, use the model from Step 4 to predict the organic acid content, compare the prediction results with the standard determination results, and validate the model performance.

[0011] Traditional standard detection methods employ high performance liquid chromatography (HPLC), using a CAPECELL PAK MG S5C18 column, a mobile phase of 0.1% phosphoric acid aqueous solution-methanol (96:4), a flow rate of 1 mL / min, a column temperature of 40℃, and a detection wavelength of 210 nm to accurately determine the content of organic acids.

[0012] A complete near-infrared detection process was constructed to achieve systematic detection of organic acid content in fragrant pears. The accuracy of the detection is ensured through multi-step collaboration, the homogeneity of sample preparation is ensured, the interference is eliminated through spectral preprocessing, and the modeling and verification form a closed loop. Compared with traditional single detection methods, no complex chemical pretreatment is required, the detection efficiency is improved, and it can be adapted to fragrant pears of different maturity and varieties, with a wide range of applications. It provides a standardized solution for rapid detection of fragrant pear quality and helps orchards harvest grading and market quality supervision.

[0013] Preferably, in step one, the homogenizing equipment is a high-speed tissue homogenizer, the homogenization time is controlled at 1-5 minutes, and after homogenization, the sample is passed through an 80-100 mesh sieve to remove unbroken fruit pulp residue and ensure the homogeneity of the sample to be tested.

[0014] The speed of the high-speed tissue homogenizer was set to 8000-12000 r / min. During homogenization, twice the mass of deionized water was added to assist in the crushing process. After sieving, the sample was placed in a sealed centrifuge tube and stored at 4℃ for later use to prevent sample deterioration.

[0015] Clearly define the type and key parameters of the homogenizing equipment to ensure sample homogenization. A high-speed tissue homogenizer, combined with specific time and sieving, can thoroughly break down fruit pulp cells, release organic acid components, and avoid interference from uncrushed residues on spectral detection. The combination of 1-5 minutes of homogenization time and an 80-100 mesh sieve ensures the homogenization effect while avoiding oxidation of components due to over-crushing. After sieving, the sample particle size is uniform, resulting in consistent light scattering during spectral acquisition, improving the stability of spectral data, and providing a high-quality sample basis for subsequent modeling.

[0016] Preferably, in step one, the total number of pear samples tested is no less than 30, covering the three maturity stages of early maturity, mid maturity and late maturity, and including at least 5 mainstream cultivated varieties, covering the main pear planting areas;

[0017] The samples were selected by random sampling, with at least 10 samples at each maturity stage. The five mainstream varieties included Korla fragrant pear and red fragrant crisp pear, and the main planting areas covered core production areas such as Xinjiang, Gansu and Shaanxi to ensure the representativeness of the samples.

[0018] Standardize sample selection criteria to improve model universality; a sample size of no less than 30 samples meets the sample size requirements for chemometric modeling, covering three maturity stages and multiple mainstream varieties, and can comprehensively cover the variation range of organic acid content in fragrant pears; covering major planting areas can eliminate the influence of regional environmental differences, making the constructed model applicable to fragrant pear detection in different production areas; random sampling and stratified distribution ensure sample representativeness, avoid sample bias leading to model failure, and significantly improve the practical application value of the model.

[0019] Preferably, in step two, the optimized near-infrared spectral detection parameters are: a detection wavelength range of 780-2500 nm, 32-128 scans, and a spectral resolution of 4-16 cm⁻¹. -1 Baseline calibration was performed using a standard reference whiteboard before each data collection.

[0020] The near-infrared spectrometer used is a Fourier transform type. During detection, the distance between the probe and the sample is fixed at 2 cm. The diffuse reflectance detection mode is adopted. The spectrum of each sample is collected three times, and the average value is taken as the raw spectral data.

[0021] The optimized spectral detection parameters are highly targeted, ensuring the quality of the original spectrum; the wavelength range of 780-2500nm covers the characteristic absorption peak region of organic acids; the number of scans (32-128) balances detection efficiency and signal intensity; and the wavelength range of 4-16cm... -1 The high resolution clearly distinguishes adjacent characteristic peaks; the standard reference whiteboard baseline calibration eliminates instrument drift interference; the combination of Fourier transform spectrometer and diffuse reflectance mode is suitable for the detection of pear homogenate samples; the obtained raw spectrum has a high signal-to-noise ratio, providing high-quality data support for subsequent preprocessing and modeling, and reducing the difficulty of subsequent calibration.

[0022] Preferably, in step two, the spectral acquisition process is carried out in a constant temperature and humidity environment, with the ambient temperature controlled at 20-25℃ and the relative humidity controlled at 40%-60%, to avoid environmental factors from interfering with the spectral data.

[0023] The constant temperature and humidity environment is achieved through an intelligent constant temperature and humidity chamber. The air velocity inside the chamber is controlled below 0.3 m / s. The sample is placed for 30 minutes before spectral acquisition to ensure that the sample temperature is consistent with the ambient temperature.

[0024] Constant temperature and humidity control eliminates environmental interference and improves detection repeatability; a temperature and humidity range of 20-25℃ and 40%-60% avoids spectral redshift / blueshift caused by temperature changes and reduces the interference of humidity on the -OH absorption peak of organic acids; the combination of intelligent constant temperature and humidity chamber and sample constant temperature pretreatment ensures that the detection environment of different batches of samples is consistent, reducing spectral differences caused by environmental factors; experiments show that this environmental control can reduce the relative standard deviation of spectral data and significantly improve the repeatability and reliability of detection results.

[0025] Preferably, in step three, the spectral preprocessing also includes derivative processing, which uses the first or second derivative and combines it with the Savitzky-Golay filtering algorithm to further eliminate spectral overlap interference and improve the identification of characteristic peaks.

[0026] The Savitzky-Golay filter window width is set to 11-15 points, the first derivative processing interval is 2-4nm, and the second derivative processing interval is 4-6nm. The derivative calculation and filtering are automatically completed by the spectral software.

[0027] The new derivative processing and filtering algorithms enhance spectral feature extraction; Savitzky-Golay filtering effectively smooths noise and preserves characteristic peak information; first / second derivative processing eliminates baseline drift and spectral overlap, improving the identification of organic acid characteristic peaks; compared to single preprocessing methods, this combination can enhance the intensity of characteristic peaks and reduce interference from irrelevant information; automatic processing via spectral software makes it easy to operate and suitable for batch sample analysis, laying the foundation for building high-precision quantitative models, enabling the models to accurately detect even low-content organic acids.

[0028] Preferably, in step three, the execution order of spectral preprocessing is as follows: first, baseline correction is performed, then smoothing and multivariate scattering correction are performed sequentially, and finally derivative processing is performed to ensure synergistic effect of each preprocessing step;

[0029] Each preprocessing step was performed using professional spectral analysis software (such as Unscrambler). After each step, the spectral quality was checked. If the baseline drift was still greater than 0.02, the baseline was corrected again.

[0030] A scientifically ordered preprocessing sequence achieves synergistic effects: baseline correction first eliminates overall offset, followed by smoothing to reduce noise, then multivariate scattering correction to eliminate the influence of particle scattering, and finally derivative processing to improve feature recognition. The steps are progressive and complementary, avoiding information loss or interference residue caused by disordered processing. Experiments have shown that this order can improve the correlation coefficient of the preprocessed spectrum by 0.1-0.2. The standardized processing procedure is easy for operators to master, reduces human error, ensures consistency of spectral data processed by different personnel, and improves the stability of the model.

[0031] Preferably, in step four, the chemometric method specifically adopts partial least squares regression, combined with leave-one-out cross-validation to optimize model parameters, eliminate abnormal sample data, and improve the model's fit.

[0032] Leave-one-out cross-validation involves modeling and validating individual samples one by one by eliminating them, repeating until all samples have been validated. The Mahalanobis distance method is used to identify outliers. When the Mahalanobis distance of a sample is greater than 3, it is determined to be an outlier and is eliminated.

[0033] Partial least squares regression is adapted to the multivariate characteristics of spectral data, and leave-one-out cross-validation optimizes the model. This method can effectively handle the multicollinearity problem of spectral data and build a stable model even with small sample sizes. Leave-one-out cross-validation optimizes model parameters precisely by fully validating the samples, and Mahalanobis distance is used to remove outliers, avoiding interference from outliers on the model. This improves the model fit and reduces the prediction error. Compared with other chemometric methods, this combination is more suitable for the near-infrared detection of organic acids in pear and has stronger model robustness.

[0034] Preferably, in step four, the test samples from step one are divided into a modeling set (70%-80%) and an internal validation set (20%-30%). The modeling set is used to build the model, and the internal validation set is used to initially optimize the model parameters.

[0035] The sample division adopts a stratified random sampling method to ensure that the modeling set and the internal validation set are consistent in terms of variety, maturity and production area distribution. The number of samples in the modeling set is no less than 21 and the number of samples in the internal validation set is no less than 9.

[0036] Reasonable sample division ratios and sampling methods ensure model quality; a 70%-80% modeling set ensures sufficient samples for model construction, while a 20%-30% internal validation set effectively optimizes parameters; stratified random sampling ensures consistent distribution between the two sets of samples, avoiding systematic bias between the modeling and validation sets and ensuring the effectiveness of model optimization; clear division between the modeling and validation sets forms an initial quality control step, enabling timely detection of overfitting or underfitting issues, facilitating early parameter adjustments, and improving the model's generalization ability.

[0037] Preferably, in step five, the evaluation metrics for model validation include the correlation coefficient (R²). 2 ), root mean square error and relative analysis error, requiring R 2 ≥0.95, root mean square error ≤0.05%, relative analysis error ≥3.0, ensuring the model meets actual detection requirements;

[0038] During model validation, the number of blind samples should be no less than 15, covering different varieties and production areas. Evaluation indicators should be calculated using Excel or Origin software. If the requirements are not met, return to step four to re-optimize the model parameters.

[0039] Clear validation indicators and thresholds ensure the practical value of the model; the correlation coefficient, root mean square error, and relative analysis error comprehensively evaluate the model's fit, accuracy, and predictive ability, and thresholds such as R2≥0.95 meet the industry standards for near-infrared detection of agricultural products; blind sample validation covers different batches and production areas, which can truly reflect the actual application effect of the model; through quantitative indicator validation, unqualified models are avoided from being put into use, ensuring that the detection results meet the actual needs of orchard grading, market supervision, etc., and improving the credibility and promotion value of the detection method.

[0040] Compared with existing technologies, this invention provides a method for detecting the organic acid content of pears based on near-infrared spectroscopy, which has the following beneficial effects:

[0041] This invention achieves efficient and accurate detection of organic acid content in fragrant pears by preparing uniform samples from pears of different maturity and varieties, optimizing spectral parameters for data acquisition, preprocessing spectra using a combination of multiple methods, constructing a model by correlating with standard test results, and validating the model with blind samples. It has the advantages of wide applicability, standardized detection process, and no need for complex chemical pretreatment.

[0042] Multiple sample preparation ensures the universality of detection, optimized spectral acquisition and preprocessing eliminate interference and improve data reliability, and model construction and validation form a closed loop to ensure detection accuracy. This solves the problems of traditional detection methods, such as narrow applicability, cumbersome operation, and low detection efficiency and unstable results due to many interfering factors. Attached Figure Description

[0043] Figure 1 This is a flowchart illustrating the steps of the method for detecting organic acid content in fragrant pears based on near-infrared spectroscopy of the present invention.

[0044] Figure 2 This is a schematic diagram of the near-infrared detection model construction and verification system for organic acids in fragrant pear according to the present invention;

[0045] Figure 3 The pear sample processing procedure of the present invention. Detailed Implementation

[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] This invention provides a technical solution: a method for detecting the organic acid content of pears based on near-infrared spectroscopy. Please refer to [link to relevant documentation]. Figure 1 , Figure 2 and Figure 3 It includes the following steps:

[0048] Step 1, Sample Preparation: Select pears of different maturity and varieties as test samples. After manually removing the fruit stems and cores, place the pulp in a homogenizing device for crushing and homogenization to prepare a uniformly dispersed test sample.

[0049] Step 2, Spectral Acquisition: Optimize near-infrared spectral detection parameters, and use a near-infrared spectrometer to acquire spectral data of the sample prepared in Step 1 to obtain the raw near-infrared spectrum;

[0050] Step 3: Spectral preprocessing: The near-infrared raw spectral data obtained in Step 2 is preprocessed using a combination of multivariate scattering correction, baseline correction, and smoothing to eliminate interference factors.

[0051] Step 4: Model Establishment: Select some of the samples to be tested in Step 1 and determine their organic acid content using traditional standard detection methods. Correlate the determination results with the spectral data after preprocessing in Step 3, and construct a quantitative analysis model by combining chemometric methods.

[0052] Step 5: Model Validation: Collect blind samples of pears from different batches and sources, use the model from Step 4 to predict the organic acid content, compare the prediction results with the standard determination results, and validate the model performance.

[0053] Traditional standard detection methods employ high performance liquid chromatography (HPLC), using a CAPECELL PAK MG S5C18 column, a mobile phase of 0.1% phosphoric acid aqueous solution-methanol (96:4), a flow rate of 1 mL / min, a column temperature of 40℃, and a detection wavelength of 210 nm to accurately determine the content of organic acids.

[0054] A complete near-infrared detection process was constructed to achieve systematic detection of organic acid content in fragrant pears. The accuracy of the detection is ensured through multi-step collaboration, the homogeneity of sample preparation is ensured, the interference is eliminated through spectral preprocessing, and the modeling and verification form a closed loop. Compared with traditional single detection methods, no complex chemical pretreatment is required, the detection efficiency is improved, and it can be adapted to fragrant pears of different maturity and varieties, with a wide range of applications. It provides a standardized solution for rapid detection of fragrant pear quality and helps orchards harvest grading and market quality supervision.

[0055] In step one, a high-speed tissue homogenizer is used for homogenization. The homogenization time is controlled at 1-5 minutes. After homogenization, the sample is passed through an 80-100 mesh sieve to remove unbroken fruit pulp residue and ensure the homogeneity of the sample to be tested.

[0056] The speed of the high-speed tissue homogenizer was set to 8000-12000 r / min. During homogenization, twice the mass of deionized water was added to assist in the crushing process. After sieving, the sample was placed in a sealed centrifuge tube and stored at 4℃ for later use to prevent sample deterioration.

[0057] Clearly define the type and key parameters of the homogenizing equipment to ensure sample homogenization. A high-speed tissue homogenizer, combined with specific time and sieving, can thoroughly break down fruit pulp cells, release organic acid components, and avoid interference from uncrushed residues on spectral detection. The combination of 1-5 minutes of homogenization time and an 80-100 mesh sieve ensures the homogenization effect while avoiding oxidation of components due to over-crushing. After sieving, the sample particle size is uniform, resulting in consistent light scattering during spectral acquisition, improving the stability of spectral data, and providing a high-quality sample basis for subsequent modeling.

[0058] In step one, the total number of pear samples tested shall not be less than 30, covering the three maturity stages of early maturity, mid maturity and late maturity, and including at least 5 mainstream cultivated varieties, covering the main pear growing areas;

[0059] The samples were selected by random sampling, with at least 10 samples at each maturity stage. The five mainstream varieties included Korla fragrant pear and red fragrant crisp pear, and the main planting areas covered core production areas such as Xinjiang, Gansu and Shaanxi to ensure the representativeness of the samples.

[0060] Standardize sample selection criteria to improve model universality; a sample size of no less than 30 samples meets the sample size requirements for chemometric modeling, covering three maturity stages and multiple mainstream varieties, and can comprehensively cover the variation range of organic acid content in fragrant pears; covering major planting areas can eliminate the influence of regional environmental differences, making the constructed model applicable to fragrant pear detection in different production areas; random sampling and stratified distribution ensure sample representativeness, avoid sample bias leading to model failure, and significantly improve the practical application value of the model.

[0061] In step two, the optimized near-infrared spectral detection parameters are as follows: the detection wavelength range is set to 780-2500nm, the number of scans is 32-128, the spectral resolution is 4-16cm-1, and baseline calibration is performed using a standard reference white board before each acquisition.

[0062] The near-infrared spectrometer used is a Fourier transform type. During detection, the distance between the probe and the sample is fixed at 2 cm. The diffuse reflectance detection mode is adopted. The spectrum of each sample is collected three times, and the average value is taken as the raw spectral data.

[0063] The optimized spectral detection parameters are highly targeted, ensuring the quality of the original spectrum; the wavelength range of 780-2500nm covers the characteristic absorption peak region of organic acids; the number of scans from 32 to 128 balances detection efficiency and signal intensity; the resolution of 4-16cm-1 can clearly distinguish adjacent characteristic peaks; the standard reference white board baseline calibration can eliminate instrument drift interference; the combination of Fourier transform spectrometer and diffuse reflectance mode is suitable for the detection of pear homogenate samples, and the obtained original spectrum has a high signal-to-noise ratio, providing high-quality data support for subsequent preprocessing and modeling, and reducing the difficulty of subsequent calibration.

[0064] In step two, the spectral acquisition process is carried out in a constant temperature and humidity environment, with the ambient temperature controlled at 20-25℃ and the relative humidity controlled at 40%-60%, to avoid environmental factors from interfering with the spectral data.

[0065] The constant temperature and humidity environment is achieved through an intelligent constant temperature and humidity chamber. The air velocity inside the chamber is controlled below 0.3 m / s. The sample is placed for 30 minutes before spectral acquisition to ensure that the sample temperature is consistent with the ambient temperature.

[0066] Constant temperature and humidity control eliminates environmental interference and improves detection repeatability; a temperature and humidity range of 20-25℃ and 40%-60% avoids spectral redshift / blueshift caused by temperature changes and reduces the interference of humidity on the -OH absorption peak of organic acids; the combination of intelligent constant temperature and humidity chamber and sample constant temperature pretreatment ensures that the detection environment of different batches of samples is consistent, reducing spectral differences caused by environmental factors; experiments show that this environmental control can reduce the relative standard deviation of spectral data and significantly improve the repeatability and reliability of detection results.

[0067] In step three, spectral preprocessing also includes derivative processing, which uses the first or second derivative and combines it with the Savitzky-Golay filtering algorithm to further eliminate spectral overlap interference and improve the identification of characteristic peaks.

[0068] The Savitzky-Golay filter window width is set to 11-15 points, the first derivative processing interval is 2-4nm, and the second derivative processing interval is 4-6nm. The derivative calculation and filtering are automatically completed by the spectral software.

[0069] The new derivative processing and filtering algorithms enhance spectral feature extraction; Savitzky-Golay filtering effectively smooths noise and preserves characteristic peak information; first / second derivative processing eliminates baseline drift and spectral overlap, improving the identification of organic acid characteristic peaks; compared to single preprocessing methods, this combination can enhance the intensity of characteristic peaks and reduce interference from irrelevant information; automatic processing via spectral software makes it easy to operate and suitable for batch sample analysis, laying the foundation for building high-precision quantitative models, enabling the models to accurately detect even low-content organic acids.

[0070] In step three, the execution order of spectral preprocessing is as follows: first, baseline correction is performed, then smoothing and multivariate scattering correction are performed in sequence, and finally derivative processing is performed to ensure synergistic effect of each preprocessing step;

[0071] Each preprocessing step was performed using professional spectral analysis software (such as Unscrambler). After each step, the spectral quality was checked. If the baseline drift was still greater than 0.02, the baseline was corrected again.

[0072] A scientifically ordered preprocessing sequence achieves synergistic effects: baseline correction first eliminates overall offset, followed by smoothing to reduce noise, then multivariate scattering correction to eliminate the influence of particle scattering, and finally derivative processing to improve feature recognition. The steps are progressive and complementary, avoiding information loss or interference residue caused by disordered processing. Experiments have shown that this order can improve the correlation coefficient of the preprocessed spectrum by 0.1-0.2. The standardized processing procedure is easy for operators to master, reduces human error, ensures consistency of spectral data processed by different personnel, and improves the stability of the model.

[0073] In step four, the chemometric method specifically adopts partial least squares regression, combined with leave-one-out cross-validation to optimize model parameters, eliminate outlier data, and improve the model's fit.

[0074] Leave-one-out cross-validation involves modeling and validating individual samples one by one by eliminating them, repeating until all samples have been validated. The Mahalanobis distance method is used to identify outliers. When the Mahalanobis distance of a sample is greater than 3, it is determined to be an outlier and is eliminated.

[0075] Partial least squares regression is adapted to the multivariate characteristics of spectral data, and leave-one-out cross-validation optimizes the model. This method can effectively handle the multicollinearity problem of spectral data and build a stable model even with small sample sizes. Leave-one-out cross-validation optimizes model parameters precisely by fully validating the samples, and Mahalanobis distance is used to remove outliers, avoiding interference from outliers on the model. This improves the model fit and reduces the prediction error. Compared with other chemometric methods, this combination is more suitable for the near-infrared detection of organic acids in pear and has stronger model robustness.

[0076] In step four, the samples to be tested in step one are divided into a modeling set of 70%-80% and an internal validation set of 20%-30%. The modeling set is used to build the model, and the internal validation set is used to initially optimize the model parameters.

[0077] The sample division adopts a stratified random sampling method to ensure that the modeling set and the internal validation set are consistent in terms of variety, maturity and production area distribution. The number of samples in the modeling set is no less than 21 and the number of samples in the internal validation set is no less than 9.

[0078] Reasonable sample division ratios and sampling methods ensure model quality; a 70%-80% modeling set ensures sufficient samples for model construction, while a 20%-30% internal validation set effectively optimizes parameters; stratified random sampling ensures consistent distribution between the two sets of samples, avoiding systematic bias between the modeling and validation sets and ensuring the effectiveness of model optimization; clear division between the modeling and validation sets forms an initial quality control step, enabling timely detection of overfitting or underfitting issues, facilitating early parameter adjustments, and improving the model's generalization ability.

[0079] In step five, the evaluation indicators for model validation include correlation coefficient (R2), root mean square error, and relative analysis error. The requirements are R2 ≥ 0.95, root mean square error ≤ 0.05%, and relative analysis error ≥ 3.0, to ensure that the model meets the actual detection requirements.

[0080] During model validation, the number of blind samples should be no less than 15, covering different varieties and production areas. Evaluation indicators should be calculated using Excel or Origin software. If the requirements are not met, return to step four to re-optimize the model parameters.

[0081] Clear validation indicators and thresholds ensure the practical value of the model; the correlation coefficient, root mean square error, and relative analysis error comprehensively evaluate the model's fit, accuracy, and predictive ability, and thresholds such as R2≥0.95 meet the industry standards for near-infrared detection of agricultural products; blind sample validation covers different batches and production areas, which can truly reflect the actual application effect of the model; through quantitative indicator validation, unqualified models are avoided from being put into use, ensuring that the detection results meet the actual needs of orchard grading, market supervision, etc., and improving the credibility and promotion value of the detection method.

[0082] This method utilizes the characteristic interaction between near-infrared light and organic acid molecules in pear pulp. A spectrometer captures the absorption signal of organic acids at specific wavelengths of near-infrared light. Preprocessing eliminates interference from physical scattering and baseline drift, extracting effective spectral features. A quantitative correlation model between spectral features and organic acid content is constructed using chemometric methods. Blind sample validation ensures the model's reliability, enabling rapid quantitative detection of organic acid content.

[0083] First, prepare homogenized samples of fragrant pears covering multiple varieties and maturity levels according to requirements and sieve them for later use; adjust the spectrometer parameters to the optimized value and collect the original spectra of the samples in a constant temperature and humidity environment; complete the spectral preprocessing in the specified order using software; divide the modeling set and the validation set, and construct and optimize the model using partial least squares regression; validate the model with blind samples, and once it meets the standards, it can be used to detect the organic acid content of unknown fragrant pear samples. Baseline calibration and environmental control need to be performed simultaneously during the detection.

[0084] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0085] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for detecting the organic acid content of fragrant pears based on near-infrared spectroscopy, characterized in that, Includes the following steps: Step 1, Sample Preparation: Select fragrant pears of different maturity and varieties as test samples. After manually removing the fruit stems and cores, place the pulp in a homogenizing device for crushing and homogenization to prepare a uniformly dispersed test sample. Step 2, Spectral Acquisition: Optimize near-infrared spectral detection parameters, and use a near-infrared spectrometer to acquire spectral data of the sample prepared in Step 1 to obtain the raw near-infrared spectrum; Step 3: Spectral preprocessing: The near-infrared raw spectral data obtained in Step 2 is preprocessed using a combination of multivariate scattering correction, baseline correction, and smoothing to eliminate interference factors. Step 4: Model Establishment: Select some of the samples to be tested in Step 1 and determine their organic acid content using traditional standard detection methods. Correlate the determination results with the spectral data after preprocessing in Step 3, and construct a quantitative analysis model by combining chemometric methods. Step 5: Model Validation: Collect blind samples of pears from different batches and sources, use the model from Step 4 to predict the organic acid content, compare the prediction results with the standard determination results, and validate the model performance.

2. The method for detecting the organic acid content of pear based on near-infrared spectroscopy according to claim 1, characterized in that: In step one, the homogenizing equipment is a tissue homogenizer, and the homogenization time is controlled within 1-5 minutes. After homogenization, the sample is passed through an 80-100 mesh sieve to remove unbroken fruit pulp residue and ensure the homogeneity of the sample to be tested.

3. The method for detecting the organic acid content of pear based on near-infrared spectroscopy according to claim 2, characterized in that: In step one, the total number of fragrant pear samples tested shall not be less than 30, covering the three maturity stages of early maturity, mid maturity and late maturity, and including at least 5 mainstream cultivated varieties, covering the main fragrant pear planting areas.

4. The method for detecting the organic acid content of pear based on near-infrared spectroscopy according to claim 1, characterized in that: In step two, the optimized near-infrared spectral detection parameters are specifically: the detection wavelength range is set to 780-2500 nm, the number of scans is 32-128, and the spectral resolution is 4-16 cm⁻¹. -1 Baseline calibration was performed using a standard reference whiteboard before each data collection.

5. The method for detecting the organic acid content of pear based on near-infrared spectroscopy according to claim 4, characterized in that: In step two, the spectral acquisition process is carried out in a constant temperature and humidity environment, with the ambient temperature controlled at 20-25℃ and the relative humidity controlled at 40%-60%.

6. The method for detecting the organic acid content of fragrant pears based on near-infrared spectroscopy according to claim 1, characterized in that: In step three, the spectral preprocessing also includes derivative processing, which uses the first or second derivative and combines it with the Savitzky-Golay filtering algorithm to further eliminate spectral overlap interference.

7. The method for detecting the organic acid content of pear based on near-infrared spectroscopy according to claim 6, characterized in that: In step three, the execution order of spectral preprocessing is as follows: first, baseline correction is performed, then smoothing and multivariate scattering correction are performed in sequence, and finally, derivative processing is performed.

8. The method for detecting the organic acid content of pear based on near-infrared spectroscopy according to claim 1, characterized in that: In step four, the chemometric method specifically employs partial least squares regression, combined with leave-one-out cross-validation to optimize model parameters and eliminate outlier sample data.

9. The method for detecting the organic acid content of pear based on near-infrared spectroscopy according to claim 8, characterized in that: In step four, the samples to be tested in step one are divided into a modeling set (70%-80%) and an internal validation set (20%-30%). The modeling set is used to build the model, and the internal validation set is used to initially optimize the model parameters.

10. The method for detecting the organic acid content of pear based on near-infrared spectroscopy according to claim 1, characterized in that: In step five, the evaluation metrics for model validation include correlation coefficient, root mean square error, and relative analysis error.