Tobacco leaf quality multi-modal evaluation method and system and storage medium
By constructing a fusion evaluation method for near-infrared spectroscopy and pyrolysis data of tobacco leaf quality, the problem of multimodal data fusion was solved, achieving low-cost detection and high-precision prediction, supporting stable application across equipment and operating conditions, and improving the comprehensive evaluation accuracy of tobacco leaf quality analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-07
AI Technical Summary
Existing methods for analyzing tobacco leaf quality lack a unified framework for near-infrared spectroscopy and pyrolysis analysis data, making it difficult to achieve multimodal data fusion and limiting the improvement of the accuracy of comprehensive evaluation of tobacco leaf quality.
A fusion evaluation method for near-infrared spectroscopy and pyrolysis data was constructed, including acquiring multimodal data, reconstructing spectral data, constructing a chemical composition prediction model, establishing a back regression model, and calculating evaluation indicators to quantify the complementarity and consistency of the data, and conducting statistical tests and confidence assessments.
It achieves reduced detection costs while ensuring prediction accuracy, improves detection efficiency and prediction performance by identifying data complementarity to guide data fusion strategies, and ensures the stability and usability of the model in different environments.
Smart Images

Figure CN121808736A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital detection and multimodal data fusion analysis of tobacco raw material quality, specifically to a method, system, and storage medium for multimodal evaluation of tobacco leaf quality. Background Technology
[0002] In the wave of digital and intelligent R&D for cigarette products, building a precise digital characterization system for tobacco raw material quality is the cornerstone for achieving scientific formula design and process optimization. This system is like the digital gene of cigarette products, and its completeness and accuracy directly determine the depth and effectiveness of digital R&D. Currently, the digital characterization of tobacco quality relies heavily on near-infrared spectroscopy (NIR) technology. With its advantages of speed and non-destructive testing, NIR technology can effectively measure the static chemical components of tobacco leaves (such as nicotine, sugar, and moisture), and has become the mainstream method for quality control and formula maintenance in the industry. However, the sensory quality of cigarettes is ultimately determined by the complex chemical reactions during combustion. As a static characterization method, NIR cannot reflect the dynamic process of pyrolysis of chemical components at high temperatures, the release patterns of products, and combustion performance. This leads to a characterization gap between NIR-based predictive models and the final sensory experience.
[0003] To compensate for the limitations of static chemical information, pyrolysis analysis techniques (such as thermogravimetric-infrared spectroscopy) have attracted attention in recent years. This technique can simulate the combustion process of tobacco leaves, revealing their thermal stability, weight loss characteristics, and the release kinetics of gaseous products, providing crucial information for predicting combustion flavor and sensory characteristics from a dynamic perspective. Thus, NIR and pyrolysis techniques characterize tobacco leaf quality from two orthogonal and complementary dimensions: static composition and dynamic behavior. However, the fundamental differences between the two in terms of physical quantity domain, time dimension, and information focus constitute a technical bottleneck for multimodal data fusion. Existing research methods typically model and analyze the two types of data in isolation, lacking a unified framework to quantify their inherent correlation and complementarity. This data silo phenomenon makes it difficult to achieve synergistic effects, limiting further improvements in the accuracy of comprehensive tobacco leaf quality evaluation.
[0004] Therefore, the key to overcoming the existing technological bottlenecks lies in how to establish a new method that can effectively measure the complementary relationship between NIR spectra and pyrolysis data, and on this basis, construct a unified analysis and prediction model across modes. Summary of the Invention
[0005] The purpose of this invention is to provide a NIR-pyrolysis fusion evaluation method, system, and storage medium for tobacco leaf quality, in order to solve the problems of substitution quantification and complementary measurement of near-infrared spectroscopy and pyrolysis analysis data in tobacco leaf quality analysis, so as to achieve optimal decision-making between detection cost and prediction accuracy.
[0006] To achieve the above objectives, embodiments of the present invention provide a multimodal evaluation method for tobacco leaf quality, comprising: Acquire near-infrared spectral data and pyrolysis data of tobacco leaves; Reconstruct near-infrared spectral data based on the pyrolysis data to obtain reconstructed near-infrared spectral data; A chemical composition prediction model was constructed based on the near-infrared spectral data. Based on the reconstructed near-infrared spectral data, predict the chemical composition to obtain the predicted values of the chemical composition of the reconstructed data. A near-infrared data-pyrolysis data inverse regression model was constructed, and the predicted chemical composition values were validated. Calculate evaluation indicators, including the consistency correlation index and the complementarity correlation index of the data system; Based on the evaluation indicators, statistical tests and confidence assessments are performed.
[0007] Optionally, reconstructing near-infrared spectral data based on the pyrolysis data to obtain reconstructed near-infrared spectral data includes: Construct a regression model based on formula (1). (1) in, To reconstruct near-infrared spectral data, For pyrolysis data, For regression coefficients, For the intercept term; The regression model was trained using the pyrolysis data and near-infrared spectral data. Based on the trained regression model, the pyrolysis data is converted into reconstructed near-infrared spectral data.
[0008] Optionally, constructing a chemical composition prediction model based on the near-infrared spectral data includes: A chemical composition prediction model is constructed based on formula (2). (2) in, These are predicted values for chemical composition. Near-infrared spectral data, For regression coefficients, For the intercept term; Cross-validation was performed using independent datasets to obtain the optimal model parameters.
[0009] Optionally, predicting chemical composition based on the reconstructed near-infrared spectral data to obtain predicted values of chemical composition from the reconstructed data includes: Based on formula (3), a regression model for predicting reconstructed data components is established. (3) in, To reconstruct the predicted values of chemical composition from the data, To reconstruct near-infrared spectral data, For regression coefficients, For the intercept term; Cross-validation was performed using independent datasets to obtain the optimal model parameters; The reconstructed near-infrared spectral data is fed into the reconstructed data component prediction regression model to obtain the corresponding predicted values of the reconstructed data chemical composition.
[0010] Optionally, constructing a near-infrared data-pyrolysis data inverse regression model and validating the predicted chemical composition values includes: Construct a back regression model based on formula (4). (4) in, To reverse the pyrolysis data, Near-infrared spectral data, For regression coefficients, For the intercept term; Chemical composition is predicted based on the reverse reconstructed pyrolysis data and the original pyrolysis data, respectively, and the accuracy of the prediction is evaluated by calculating the root mean square error and the coefficient of determination.
[0011] Optionally, the evaluation indicators, including the data system consistency correlation index and the complementarity correlation index, are calculated as follows: The consistency correlation degree of the data system is calculated according to formulas (5) to (7). (5) (6) (7) in, For the consistency and correlation of the data system, To reconstruct the consistency correlation of near-infrared data using pyrolysis data, To reconstruct the consistency correlation of pyrolysis data using near-infrared data, Using reconstructed near-infrared spectral data Predicted components The root mean square error, Using raw near-infrared data Predicted components The root mean square error, Using reconstructed pyrolysis data Predicted components The root mean square error, Predicting composition using raw pyrolysis data The root mean square error.
[0012] Optionally, the evaluation indicators, including the data system consistency correlation index and the complementarity correlation index, are calculated as follows: The complementarity correlation index is calculated according to formulas (8) to (10). (8) (9) (10) in, It is a complementarity correlation index. In the case of reconstructing near-infrared spectral data from pyrolysis data, the original near-infrared spectral data provides information about... Additional information, In the case of reconstructing pyrolysis data using near-infrared spectroscopy data, the original pyrolysis data provides information about... Additional information, joint variables With target chemical components Mutual information between them To reconstruct near-infrared spectral data With target chemical components Mutual information between them joint variables With target chemical components Mutual information between them To reconstruct pyrolysis data With target chemical components Mutual information between them.
[0013] Optionally, based on the evaluation indicators, statistical tests and confidence assessments include: The regression model was evaluated at both the reconstruction layer and the component layer. The confidence intervals for DSA and CCI were estimated using the bootstrap method; Stratified reports are generated for different batches and production areas to verify the applicability of the model and its consistency across batches; Perform segmented direct correction operations and subspace alignment operations.
[0014] On the other hand, the present invention also provides a multimodal evaluation system for tobacco leaf quality, the system including a processor configured to perform any of the methods described above.
[0015] In another aspect, the present invention also provides a computer-readable storage medium storing instructions that, when executed by a processor, implement any of the methods described above.
[0016] The beneficial effects of this invention are: This invention provides a scientific basis for technology selection and process control by quantitatively evaluating the substitutability (DSA) and complementarity (CCI) of multi-source data. When the DSA is close to 1, it supports replacing high-cost methods (such as near-infrared) with low-cost detection methods (such as pyrolysis), significantly reducing calibration and equipment usage costs while ensuring prediction accuracy. At the same time, through complementarity identification, it guides data fusion strategies to achieve the best balance between detection efficiency and prediction performance.
[0017] The embodiments of this invention employ a smoothing constraint, leave-one-batch validation, and bootstrap confidence interval evaluation mechanism to effectively suppress batch effects and data bias, ensuring the statistical reliability of model evaluation conclusions. Combined with piecewise direct correction (PDS) and subspace alignment strategies, calibration migration across devices and operating conditions is achieved, guaranteeing stable model performance and long-term availability under different production lines and instrument environments.
[0018] The embodiments of this invention support flexible replacement and combination of data preprocessing, regression models, and regularization methods, possessing "plug-and-play" engineering adaptability. It can simultaneously perform multi-task joint modeling for various chemical components and integrate prior knowledge such as pyrolysis kinetics to enhance the interpretability and generalization ability of the model, providing a systematic solution for multimodal quality monitoring and intelligent formulation in industries such as tobacco.
[0019] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description
[0020] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings: Figure 1 A flowchart of a multimodal evaluation method for tobacco leaf quality according to an embodiment of the present invention; Figure 2 A flowchart of a method for reconstructing near-infrared spectral data from pyrolysis data to obtain reconstructed near-infrared spectral data according to an embodiment of the present invention; Figure 3 This is a flowchart of a method for predicting chemical composition based on reconstructed near-infrared spectral data to obtain predicted values of chemical composition from reconstructed data, according to one embodiment of the present invention. Detailed Implementation
[0021] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.
[0022] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application all comply with relevant laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.
[0023] like Figure 1 The diagram shows a flowchart of a multimodal evaluation method for tobacco leaf quality according to an embodiment of the present invention. Figure 1 In this evaluation method, the steps may include: In step S10, near-infrared spectral data and pyrolysis data of tobacco leaves are acquired; In step S11, near-infrared spectral data are reconstructed based on pyrolysis data to obtain reconstructed near-infrared spectral data; In step S12, a chemical composition prediction model is constructed based on near-infrared spectral data; In step S13, the chemical composition is predicted based on the reconstructed near-infrared spectral data to obtain the predicted values of the chemical composition from the reconstructed data. In step S14, a near-infrared data-pyrolysis data inverse regression model is constructed, and the predicted chemical composition values are verified. In step S15, evaluation indicators are calculated, including the consistency correlation degree and complementarity correlation degree index of the data system; In step S16, statistical tests and confidence assessments are performed based on the evaluation indicators.
[0024] In such Figure 1 In the multimodal evaluation method for tobacco leaf quality shown, step S10 is used to acquire near-infrared spectral data and pyrolysis data of tobacco leaves. The sample data mainly comes from multiple domestic tobacco-producing areas (including Yunnan, Guizhou, Sichuan, etc.). In this embodiment, 399 samples are randomly selected as the training set and 100 samples are selected as the test set. The division of the training set and the test set uses leave-one-out cross-validation to ensure the stability of the model under different data distributions. The near-infrared spectral data contains 1609 dimensions, representing the near-infrared spectral data of the tobacco leaf samples, and the wavelength range of acquisition is 10000 to 3800 cm⁻¹. -¹. The pyrolysis data contained 8001 dimensions and were analyzed using a Discovery thermogravimetric analyzer (TA Instruments, USA). The pyrolysis atmosphere was nitrogen, and the flow rate was 60.0 ml / min. The pyrolysis temperature program was as follows: increasing the temperature from room temperature to 100°C at a rate of 30°C / min and holding for 5 min to completely remove free water, then increasing the temperature from 100°C to 800°C at a rate of 10°C / min, and finally cooling to 60°C. Further analysis of the chemical composition of the samples was required, including total sugar, nicotine, reducing sugar, chlorine, potassium, total nitrogen, pH, chlorogenic acid, and starch.
[0025] After acquiring near-infrared spectral and pyrolysis data, preprocessing and registration are required. In this example, near-infrared spectral data preprocessing methods may include baseline calibration, filtering, and normalization. Baseline calibration can be performed using a polynomial fitting method. Filtering and smoothing can be achieved by using Savitzky-Golay filtering to smooth the near-infrared spectral data using the first derivative, setting the window size to 15, the order to 3, and the derivative order to 1. Normalization can be performed by standard normalization correction (SNV) or multivariate scattering correction (MSC), selecting a range of 4500 to 9000 cm⁻¹. - ¹ Band range. In this example, the preprocessing method for pyrolysis data can include baseline removal, time axis alignment, and normalization. Baseline removal can involve baseline correction of the pyrolysis data to ensure accuracy. Time axis alignment can be achieved by aligning the time and temperature axes of the pyrolysis data using dynamic time warping (DTW) or fixed sampling methods. Normalization can involve logarithmic or area normalization to ensure comparability between different samples.
[0026] Step S11 is used to reconstruct near-infrared spectral data based on pyrolysis data to obtain reconstructed near-infrared spectral data. In this embodiment, the specific method for obtaining the reconstructed near-infrared spectral data in step S11 can be of various forms known to those skilled in the art. In one example of the present invention, step S11 may include, for example... Figure 2 The steps shown are described. Figure 2 In this context, step S11 may include: In step S20, a regression model is constructed; In step S21, a regression model is trained using pyrolysis data and near-infrared spectroscopy data; In step S22, the pyrolysis data is converted into reconstructed near-infrared spectral data based on the trained regression model.
[0027] In such Figure 2 In the method shown, step S20 is used to construct a regression model. Specifically, in this example, partial least squares regression (PLSR) can be used, and the regression model is in the form of: (1) in, To reconstruct near-infrared spectral data, For pyrolysis data, For regression coefficients, This is the intercept term. The number of latent variables (LVs) is determined through cross-validation, and a smoothing regularization term for the band or temperature range can be introduced. , where D is the difference matrix and F represents the Frobenius norm. This regularization term helps suppress overfitting and enhances the stability of the model across different temperature or wavelength ranges, thereby improving the smoothness of the reconstruction results.
[0028] Steps S21 and S22 are used to establish a mapping relationship between pyrolysis data X and near-infrared spectral data Y by training a regression model, thereby converting pyrolysis data X into reconstructed near-infrared spectral data. In this example, the reconstructed spectrum can also be evaluated using mean square error, spectral angle, or Pearson correlation coefficient. The similarity to the original spectrum Y is used to determine the reconstruction effect.
[0029] Step S12 is used to construct a chemical composition prediction model based on near-infrared spectral data. Specifically, in this example, the raw near-infrared spectral data Y can be used as input features, and a regression model can be trained to predict target chemical components (such as total sugar, nicotine, reducing sugar, etc.). This regression model uses the near-infrared spectral data Y as the independent variable, and the optimal model parameters are determined through cross-validation to ensure prediction accuracy. The model takes the following form:
[0030] in, These are predicted values for chemical composition. For regression coefficients, The intercept term is used. In this example, the regression model preferably uses PLSR, but it can also be replaced by kernel PLSR, SVR, ridge regression, random forest, or neural networks. RMSEP and [other methods] are calculated on independent datasets. As a metric for model evaluation.
[0031] Step S13 is used to predict chemical composition based on reconstructed near-infrared spectral data to obtain predicted chemical composition values from the reconstructed data. In this embodiment, the specific method for obtaining the predicted chemical composition values from the reconstructed data in step S13 can be of various forms known to those skilled in the art. In one example of the present invention, step S13 may include, for example... Figure 3 The steps shown are described. Figure 3 In this context, step S13 may include: In step S30, a regression model for predicting reconstructed data components is established; In step S31, cross-validation is performed using independent datasets to obtain the optimal model parameters; In step S32, the reconstructed near-infrared spectral data is fed into the reconstructed data component prediction regression model to obtain the corresponding predicted values of the reconstructed data chemical components.
[0032] In such Figure 3 In the method shown, step S30 is used to establish a regression model for predicting reconstructed data components. Specifically, in this example, the regression model for predicting reconstructed data components can be established according to formula (3): (3) in, To reconstruct the predicted values of chemical composition from the data, To reconstruct near-infrared spectral data, For regression coefficients, This is the intercept term.
[0033] Step S31 is used to perform validation using the same independent dataset as in step S12, calculating RMSEP and To evaluate the performance of the reconstructed data in component prediction and ensure its feasibility and stability in practical applications, step S32 involves inputting the reconstructed near-infrared spectral data into the reconstructed data component prediction regression model to obtain the corresponding predicted values of the reconstructed data chemical composition.
[0034] Step S14 is used to construct a near-infrared data-pyrolysis data inverse regression model and to verify the predicted chemical composition values. Specifically, in this example, an inverse regression model can be constructed to map near-infrared spectral data Y back to pyrolysis data X, and the inverse model can be used to verify the substitutability of the two types of data in composition prediction. The inverse regression model can be constructed according to formula (4): (4) in, To reverse the pyrolysis data, Near-infrared spectral data, For regression coefficients, This is the intercept term. Further, the data is reconstructed using reverse methods. And the original pyrolysis data X predicts the chemical composition Z, and the RMSEP and Compare the prediction results. This step helps verify the equivalence between the two data sources, ensuring they can be substituted for each other in different tasks.
[0035] Step S15 is used to calculate evaluation indicators, including the data system consistency correlation degree and the complementarity correlation degree index. The data system consistency correlation degree (DSA) is used to quantify the substitutability (i.e., consistency) of pyrolysis data and near-infrared spectral data in the composition prediction task. Specifically, in this example, the consistency correlation degree of reconstructing near-infrared data from pyrolysis data can be expressed as: (5) The consistency correlation of reconstructing pyrolysis data from near-infrared data can be expressed as: (6) in, Using reconstructed near-infrared spectral data Predicted components The root mean square error, Using raw near-infrared data Predicted components The root mean square error, Using reconstructed pyrolysis data Predicted components The root mean square error, Predicting composition using raw pyrolysis data The root mean square error.
[0036] Furthermore, the consistency and correlation of a data system can be expressed as: (7) The closer the value of this indicator is to 1, the stronger the consistency between the near-infrared data obtained by reconstructing pyrolysis data and the real near-infrared data in the target task in component prediction, and the stronger its substitutability.
[0037] The Complementarity Correlation Index (CCI) measures the additional information contribution from the original data source based on existing reconstructed information. In this example, when reconstructing near-infrared spectral data from pyrolysis data, the original near-infrared spectral data provides information about... Additional information can be represented as: (8) When reconstructing pyrolysis data using near-infrared spectroscopy data, the original pyrolysis data provides information about... Additional information can be represented as: (9) in, joint variables With target chemical components Mutual information between them To reconstruct near-infrared spectral data With target chemical components Mutual information between them joint variables With target chemical components Mutual information between them To reconstruct pyrolysis data With target chemical components The complementarity correlation index (CCI) is a measure of the mutual information between the original and reconstructed data. A CCI value closer to 1 indicates stronger complementarity and greater information gain between the original and reconstructed data; a CCI value closer to 0 indicates almost no complementarity and smaller information gain.
[0038] Furthermore, we define a symmetric CCI: (10) When both modes are completely redundant When they are perfectly complementary, there is This indicator can effectively assess the complementarity of different data sources.
[0039] Step S16 is used to perform statistical tests and confidence assessments based on the evaluation indicators. In this example, confidence intervals for the Data System Consistency Association (DSA) and Complementarity Association Index (CCI) can be estimated using a bootstrap method. Then, a permutation test is used to verify whether the consistency and complementarity between different data systems are statistically significant. For the assessment of RMSEP differences, a paired resampling t-test with nonparametric tests selected according to the error distribution is used to ensure the reliability and robustness of the experimental results. Specifically, in this example, it may include: In step S40, the regression model is evaluated at the reconstruction layer and the component layer. In step S41, the confidence intervals of DSA and CCI are estimated using the bootstrap method; In step S42, stratified reports are generated for different batches and production areas to verify the applicability of the model and its consistency across batches; In step S43, segmented direct correction and subspace alignment operations are performed.
[0040] Step S40 is used to perform reconstruction layer evaluation and component layer evaluation on the regression model. Specifically, in this example, reconstruction layer evaluation can assess the reconstruction effect of the regression model by calculating the root mean square error (RMSE), mean absolute error (MAE), cosine similarity, and relative bias spectrum. Component layer evaluation can assess the model's performance in component prediction using a test set by calculating RMSEP and These are the indicators. Steps S41 and S42 are used to estimate the confidence intervals of DSA and CCI using the bootstrap method to ensure model robustness. Stratified reports are generated for different batches and production areas to verify the model's applicability and consistency across batches. Step S43 is used for calibration transfer. Calibration transfer across instruments or operating conditions adopts piecewise direct correction (PDS) and subspace alignment strategies to further improve the model's generalization ability.
[0041] On the other hand, the present invention also provides a multimodal evaluation system for tobacco leaf quality, the system including a processor configured to perform any of the methods described above.
[0042] In another aspect, the present invention also provides a computer-readable storage medium storing instructions that, when executed by a processor, implement any of the methods described above.
[0043] The beneficial effects of this invention are: This invention provides a scientific basis for technology selection and process control by quantitatively evaluating the substitutability (DSA) and complementarity (CCI) of multi-source data. When the DSA is close to 1, it supports replacing high-cost methods (such as near-infrared) with low-cost detection methods (such as pyrolysis), significantly reducing calibration and equipment usage costs while ensuring prediction accuracy. At the same time, through complementarity identification, it guides data fusion strategies to achieve the best balance between detection efficiency and prediction performance.
[0044] The embodiments of this invention employ a smoothing constraint, leave-one-batch validation, and bootstrap confidence interval evaluation mechanism to effectively suppress batch effects and data bias, ensuring the statistical reliability of model evaluation conclusions. Combined with piecewise direct correction (PDS) and subspace alignment strategies, calibration migration across devices and operating conditions is achieved, guaranteeing stable model performance and long-term availability under different production lines and instrument environments.
[0045] The embodiments of this invention support flexible replacement and combination of data preprocessing, regression models, and regularization methods, possessing "plug-and-play" engineering adaptability. It can simultaneously perform multi-task joint modeling for various chemical components and integrate prior knowledge such as pyrolysis kinetics to enhance the interpretability and generalization ability of the model, providing a systematic solution for multimodal quality monitoring and intelligent formulation in industries such as tobacco.
[0046] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0047] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0048] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0049] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0050] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0051] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0052] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0053] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0054] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A multimodal evaluation method for tobacco leaf quality, characterized in that, The evaluation methods include: Acquire near-infrared spectral data and pyrolysis data of tobacco leaves; Reconstruct near-infrared spectral data based on the pyrolysis data to obtain reconstructed near-infrared spectral data; A chemical composition prediction model was constructed based on the near-infrared spectral data. Based on the reconstructed near-infrared spectral data, predict the chemical composition to obtain the predicted values of the chemical composition of the reconstructed data. A near-infrared data-pyrolysis data inverse regression model was constructed, and the predicted chemical composition values were validated. Calculate evaluation indicators, including the consistency correlation index and the complementarity correlation index of the data system; Based on the evaluation indicators, statistical tests and confidence assessments are performed.
2. The evaluation method according to claim 1, characterized in that, Reconstructing near-infrared spectral data from the pyrolysis data to obtain reconstructed near-infrared spectral data includes: Construct a regression model based on formula (1). ,(1) in, To reconstruct near-infrared spectral data, For pyrolysis data, For regression coefficients, For the intercept term; The regression model was trained using the pyrolysis data and near-infrared spectral data. Based on the trained regression model, the pyrolysis data is converted into reconstructed near-infrared spectral data.
3. The evaluation method according to claim 1, characterized in that, The chemical composition prediction model constructed based on the near-infrared spectral data includes: A chemical composition prediction model is constructed based on formula (2). ,(2) in, These are predicted values for chemical composition. Near-infrared spectral data, For regression coefficients, For the intercept term; Cross-validation was performed using independent datasets to obtain the optimal model parameters.
4. The evaluation method according to claim 1, characterized in that, Based on the reconstructed near-infrared spectral data, the predicted chemical composition is obtained by including: Based on formula (3), a regression model for predicting reconstructed data components is established. ,(3) in, To reconstruct the predicted values of chemical composition from the data, To reconstruct near-infrared spectral data, For regression coefficients, For the intercept term; Cross-validation was performed using independent datasets to obtain the optimal model parameters; The reconstructed near-infrared spectral data is fed into the reconstructed data component prediction regression model to obtain the corresponding predicted values of the reconstructed data chemical composition.
5. The evaluation method according to claim 1, characterized in that, The construction of a near-infrared data-pyrolysis data inverse regression model and the validation of the predicted chemical composition values include: Construct a back regression model based on formula (4). ,(4) in, To reverse the pyrolysis data, Near-infrared spectral data, For regression coefficients, For the intercept term; Chemical composition is predicted based on the reverse reconstructed pyrolysis data and the original pyrolysis data, respectively, and the accuracy of the prediction is evaluated by calculating the root mean square error and the coefficient of determination.
6. The evaluation method according to claim 1, characterized in that, The evaluation indicators, including the consistency and complementarity indices of the data system, are calculated as follows: The consistency correlation degree of the data system is calculated according to formulas (5) to (7). ,(5) ,(6) ,(7) in, For the consistency and correlation of the data system, To reconstruct the consistency correlation of near-infrared data using pyrolysis data, To reconstruct the consistency correlation of pyrolysis data using near-infrared data, Using reconstructed near-infrared spectral data Predicted components The root mean square error, Using raw near-infrared data Predicted components The root mean square error, Using reconstructed pyrolysis data Predicted components The root mean square error, Predicting composition using raw pyrolysis data The root mean square error.
7. The evaluation method according to claim 1, characterized in that, The evaluation indicators, including the consistency and complementarity indices of the data system, are calculated as follows: The complementarity correlation index is calculated according to formulas (8) to (10). ,(8) ,(9) ,(10) in, It is a complementarity correlation index. In the case of reconstructing near-infrared spectral data from pyrolysis data, the original near-infrared spectral data provides information about... Additional information, In the case of reconstructing pyrolysis data using near-infrared spectroscopy data, the original pyrolysis data provides information about... Additional information, joint variables With target chemical components Mutual information between them To reconstruct near-infrared spectral data With target chemical components Mutual information between them joint variables With target chemical components Mutual information between them To reconstruct pyrolysis data With target chemical components Mutual information between them.
8. The evaluation method according to claim 1, characterized in that, Based on the aforementioned evaluation indicators, statistical tests and confidence assessments include: The regression model was evaluated at both the reconstruction layer and the component layer. The confidence intervals for DSA and CCI were estimated using the bootstrap method; Stratified reports are generated for different batches and production areas to verify the applicability of the model and its consistency across batches; Perform segmented direct correction operations and subspace alignment operations.
9. A multimodal evaluation system for tobacco leaf quality, characterized in that, The system includes a processor configured to perform the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 8.