Tobacco leaf coding rate reduction and information theory evaluation method and system and storage medium
By using a multimodal self-consistent coding rate reduction model, tobacco leaf NIR and TGA data are mapped to a unified low-dimensional space, which solves the problems of high-dimensional noise interference and lack of unified representation. This enables high-precision prediction of tobacco leaf chemical composition and modal redundancy complementarity analysis, improving prediction accuracy and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies for the fusion and complementarity analysis of NIR-pyrolysis multimodal data of tobacco leaves suffer from severe high-dimensional noise interference, lack of a unified low-dimensional characterization space, and lack of information theory measurement guided by continuous labels, making it difficult to achieve high-precision prediction of tobacco chemical composition and quantitative characterization of modal redundancy and complementarity.
A multimodal self-consistent coding rate reduction model is adopted to map the NIR and TGA data of tobacco leaves to a unified low-dimensional latent representation space. Through the closed-loop structure of semi-reconstruction-recoding and the self-consistent loss constraint of the latent space, coding rate estimation and mutual information analysis are performed. A total loss function is constructed for training. The modal redundancy and complementarity are characterized by cooperative gain, mutual information decomposition index and coding rate reduction difference index.
It significantly alleviates the curse of dimensionality and noise interference, improves the stability and interpretability of information theory estimation, enhances prediction accuracy and model robustness to missing and noisy modes, and provides quantifiable quantitative analysis of redundancy and complementarity.
Smart Images

Figure CN121808673A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tobacco industry technology, specifically to a method, system, and storage medium for reducing the coding rate of multimodal tobacco leaves and evaluating information theory. Background Technology
[0002] With the increasing demands for product quality consistency and refined process control in the cigarette industry, rapid and accurate evaluation technology for tobacco raw materials has become a core issue of concern in the industry. Near-infrared spectroscopy (NIR), as a non-destructive, rapid, and online detection method, has been widely used for the quantitative analysis of static chemical components such as nicotine, sugar, and moisture in tobacco leaves. However, tobacco leaves undergo complex pyrolysis processes during actual combustion and cigarette processing, generating various volatile and semi-volatile products. These dynamic processes have a significant impact on the flavor and taste of the smoke, and the static composition information characterized by NIR spectroscopy alone cannot fully reflect the quality characteristics of tobacco leaves. Pyrolysis analysis (such as thermogravimetric / pyrolysis-mass spectrometry or infrared spectroscopy) can record the mass loss, exothermic process, and product release behavior of tobacco leaf samples under programmed temperature conditions, revealing the physical properties and chemical transformation characteristics of tobacco leaves during heating and combustion from the perspective of combustion process and dynamic products. Compared to NIR, pyrolysis (TGA) data is closer to actual combustion scenarios, but it is often more dimensional, more complex in structure, and more affected by test conditions and noise factors.
[0003] In recent years, multimodal data fusion methods have been introduced into fields such as tobacco and food, aiming to improve the accuracy of quantitative analysis and quality evaluation by jointly utilizing multi-source information. Common practices include: constructing a multiple regression model by simply concatenating different modal data at the feature layer, or weighted fusion of prediction results from different modalities at the decision layer; some studies have also attempted to use unidirectional cross-modal prediction (e.g., using NIR to predict pyrolysis curves) to qualitatively assess the degree of correlation between modalities. However, existing methods still have several limitations. Both NIR and pyrolysis data are high-dimensional signals, and directly estimating similarity, error, or mutual information in the original observation space is susceptible to the curse of dimensionality, sensitive to noise and local perturbations, and difficult to obtain robust and interpretable modal correlation measures. Existing methods are mostly based on linear or shallow feature transformations, failing to map NIR and pyrolysis data to a unified low-dimensional latent space and perform aligned representation, resulting in a lack of common characterization basis for subsequent analysis of modal redundancy and modal complementarity. Most methods treat chemical composition only as a regression target, failing to explicitly constrain, from an information theory or geometric perspective, that samples with similar labels should be similarly distributed in the representation space. This results in models capable of numerical predictions but struggling to explain which information in the latent space truly aligns with the labels. Regarding the crucial issue of the complementarity and redundancy between NIR and pyrolysis modes in tasks, current methods largely rely on empirical judgment or simple performance comparisons, lacking a systematic analytical framework based on latent representation, mutual information, and coding rate.
[0004] From the perspective of representation learning and information theory, high-dimensional observation spaces often contain a large amount of task-irrelevant noise and redundancy, while the structures truly relevant to the prediction task are often hidden in lower-dimensional latent spaces. If, through appropriate coding models, NIR and pyrolysis data can be jointly mapped to a unified low-dimensional representation, and the coding rate, mutual information, and cooperative gain can be estimated in this space, it will help suppress high-dimensional noise, improve statistical stability, focus on the essential structures related to chemical composition prediction, and then quantitatively analyze the redundancy and complementarity relationships between the two modes under a unified representation.
[0005] In summary, existing technologies for the fusion and complementarity analysis of multimodal NIR-pyrolysis data of tobacco leaves still suffer from problems such as severe high-dimensional noise interference, lack of a unified low-dimensional representation space, and lack of information theory metrics guided by continuous labels. There is an urgent need to propose a method and system for reducing the coding rate and evaluating the information theory of multimodal data of tobacco leaves, so as to achieve high-precision prediction of the chemical composition of tobacco leaves and to quantitatively characterize the redundancy and complementarity between different modes. Summary of the Invention
[0006] The purpose of this invention is to provide a method, system, and storage medium for reducing the coding rate of tobacco multimodal data and evaluating information theory, in order to solve the technical problems in the prior art of tobacco NIR-pyrolysis multimodal data fusion and complementarity analysis, such as severe high-dimensional noise interference, lack of a unified low-dimensional representation space, and lack of information theory measurement guided by continuous labels.
[0007] To achieve the above objectives, embodiments of the present invention provide a method for reducing the multimodal coding rate of tobacco leaves and evaluating information theory, including: Acquire tobacco leaf sample data and preprocess it; A multimodal self-consistent coding rate reduction model was constructed and trained using preprocessed sample data; The trained multimodal self-consistent coding rate reduction model is used to obtain fusion features, and modal redundancy and complementarity analysis is performed.
[0008] Optionally, constructing a multimodal self-consistent coding rate reduction model and training it using preprocessed sample data includes: Based on the preprocessed tobacco sample data, multimodal fusion encoding is performed to obtain the potential representation vector; Based on the latent representation vector, perform modality decoding and chemical composition prediction training; The multimodal self-consistent coding rate reduction model is trained with closed-loop self-consistent constraints. The coding rate of the aforementioned multimodal self-consistent coding rate reduction model is reduced. The total loss function is constructed to train the multimodal self-consistent coding rate reduction model.
[0009] Optionally, training for modality decoding and chemical composition prediction based on the latent representation vector includes: Based on the latent representation vector, a heat dissipation mode decoder is used to obtain the reconstructed heat dissipation vector; Based on the latent representation vector, a near-infrared modal decoder is used to obtain the reconstructed near-infrared vector; Based on the potential representation vector, a chemical composition prediction vector is obtained using the chemical composition prediction branch; Based on the predicted chemical composition vector, the model is trained using supervised regression loss; Based on the reconstructed thermal vector and the reconstructed near-infrared vector, the model is trained using modal reconstruction loss.
[0010] Optionally, training the multimodal self-consistent coding rate reduction model with closed-loop self-consistent constraints includes: For each sample, a semi-reconstructed input is constructed and re-encoded according to formulas (1) and (2). (1) (2) in, This is the raw near-infrared data. This is the raw pyrolysis data. Near-infrared data reconstructed by the decoder, For the pyrolysis data reconstructed by the decoder, for and The potential representation of fusion encoding, for and The potential representation of fusion coding; Self-consistent constraint training is performed according to formula (3). (3) in, For the loss of spatial self-consistency, For the latent representation vector, This represents the number of samples.
[0011] Optionally, reducing the coding rate of the multimodal self-consistent coding rate reduction model includes: The reduction in coding rate is obtained from formulas (4) to (6). (4) (5) (6) in, The reduction is approximately due to the coding rate. For unlabeled coding rate, Weighted coding rate for labels, It is the identity matrix. For hyperparameters, This is the label similarity weight matrix.
[0012] Optionally, training the multimodal self-consistent coding rate reduction model by constructing the total loss function includes: Construct the total loss function according to formula (7). (7) in, For the total loss function, To monitor the regression loss, To monitor the regression loss weights, For the loss of spatial self-consistency, For spatial self-consistency, the loss weight is... For modal reconstruction loss, For modal reconstruction loss weights, The reduction is approximately due to the coding rate. The weighting is reduced by the coding rate. The square of the L2 norm of all parameters in the model. This is the weight decay coefficient.
[0013] Optionally, the fusion features are obtained using the trained multimodal self-consistent coding rate reduction model, and modal redundancy and complementarity analysis is performed, including: Near-infrared single-mode encoders and pyrolysis single-mode encoders were constructed respectively, and near-infrared single-mode codes and pyrolysis single-mode codes were obtained; The near-infrared single-mode encoder and the pyrolysis single-mode encoder are trained using the fusion features; A chemical composition prediction network is constructed based on the near-infrared single-mode code, pyrolysis single-mode code, and fusion features. Based on the chemical composition prediction network, predicted values based on pyrolysis features, predicted values based on near-infrared features, and predicted values based on fusion features are obtained respectively. Calculate the cooperative gain according to formulas (8) to (11). (8) (9) (10) (11) in, For synergistic gain, These are predicted values based on pyrolysis characteristics. These are predicted values based on near-infrared features. These are predicted values based on fusion features. This represents the true value of the chemical composition. for and The root mean square error between them for and The root mean square error between them for and The root mean square error between them.
[0014] Optionally, the fusion features are obtained using the trained multimodal self-consistent coding rate reduction model, and modal redundancy and complementarity analysis is performed, including: Calculate the redundancy rate and complementarity index based on partial information decomposition; Calculate the complementarity index based on mutual information decomposition; Calculate the complementary gain based on coding rate minus difference.
[0015] On the other hand, the present invention also provides a tobacco leaf multimodal coding rate reduction and information theory evaluation system, the evaluation system including a processor configured to perform any of the methods described above.
[0016] In another aspect, the present invention also provides a computer-readable storage medium storing instructions that, when executed by a processor, implement any of the methods described above.
[0017] The beneficial effects of this invention are: The embodiments of the present invention use a multimodal fusion encoder to map high-dimensional NIR and TGA data to a unified low-dimensional latent representation space, and perform prediction modeling, coding rate estimation and mutual information analysis in this space, which significantly alleviates the problems of dimensionality curse and noise interference, and improves the stability and interpretability of information theory estimation.
[0018] The embodiments of the present invention employ a semi-reconstruction-recoding closed-loop structure and a latent spatial self-consistent loss to constrain the latent representation to remain consistent under different modal combination inputs, effectively avoiding the dominant effect of a single modality in representation, making multimodal fusion more sufficient, and helping to improve prediction accuracy and the robustness of the model to missing modes and noisy modes.
[0019] The coding rate is reduced by approximately [percentage missing] by introducing continuous tag guidance in the embodiments of the present invention. This extends the idea of reducing the coding rate, which was originally applied to discrete categories, to scenarios with continuous response variables such as chemical composition. This makes the learning process of latent representation both compressible and aligned with the geometric structure of the label, providing a clear information theory explanation for the model.
[0020] The embodiments of the present invention comprehensively utilize cooperative gain, conditional mutual information, mutual information decomposition indices (including redundancy rate and cooperative rate based on PID, complementarity index CI based on mutual information decomposition), and complementary gain based on coding rate reduction difference in the latent space. These indicators systematically characterize the redundancy and complementarity of NIR and TGA modes from both predictive performance and information theory perspectives, providing quantifiable decision-making basis for instrument configuration optimization, detection scheme design, and cost-performance trade-offs.
[0021] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description
[0022] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings: Figure 1A flowchart of a method for reducing the multimodal coding rate of tobacco leaves and evaluating information theory according to an embodiment of the present invention; Figure 2 A flowchart illustrating a method for acquiring and preprocessing tobacco sample data according to an embodiment of the present invention; Figure 3 A flowchart illustrating a method for constructing and training a multimodal self-consistent coding rate reduction model according to an embodiment of the present invention; Figure 4 This is a flowchart of a method for training modal decoding and chemical composition prediction according to an embodiment of the present invention. Detailed Implementation
[0023] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.
[0024] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application all comply with relevant laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.
[0025] like Figure 1 The diagram shows a flowchart of a tobacco multimodal coding rate reduction and information theory evaluation method according to an embodiment of the present invention. Figure 1 The evaluation method includes the following steps: In step S10, tobacco sample data is acquired and preprocessed; In step S11, a multimodal self-consistent coding rate reduction model is constructed and trained using preprocessed sample data; In step S12, the fusion features are obtained using the trained multimodal self-consistent coding rate reduction model, and modal redundancy and complementarity analysis is performed.
[0026] In such Figure 1In the illustrated method for reducing the multimodal coding rate of tobacco leaves and evaluating it using information theory, step S10 is used to acquire and preprocess tobacco leaf sample data. The sample data can originate from multiple domestic tobacco-producing areas, including major flue-cured tobacco producing areas such as Yunnan, Guizhou, and Sichuan. In this example, 499 tobacco leaf samples were collected, of which 399 samples were randomly selected as the modeling set (i.e., training and validation set); the remaining 100 samples were used as the independent test set. Specifically, within the modeling set, 5-fold cross-validation was used to select the network structure and hyperparameters. Then, the final model was retrained on all 399 samples, and finally, the model performance was evaluated on the 100 independent test samples. In this embodiment, the specific method for acquiring and preprocessing tobacco leaf sample data in step S10 can be of various forms known to those skilled in the art. In one example of this invention, step S10 may include, for example... Figure 2 The steps shown are described in this. Figure 2 In this context, step S10 may include: In step S20, near-infrared spectral data, pyrolysis data, and chemical composition indicators are collected for each tobacco leaf sample; In step S21, the near-infrared spectral data is preprocessed, including baseline correction, noise reduction, scattering correction, and feature normalization. In step S22, the pyrolysis data is preprocessed, including baseline correction, interpolation and alignment, normalization, optional peak feature extraction, and feature standardization.
[0027] In such Figure 2 In the method shown, step S20 is used to collect near-infrared spectral data (NIR), pyrolysis data (TGA), and chemical composition indicators for each tobacco leaf sample. Specifically, in this example, the spectral vector dimension of each sample of near-infrared spectral data is 1609 dimensions, i.e. The spectral acquisition wavenumber range was 10000-3800 cm⁻¹, and the spectral resolution was approximately 4 cm⁻¹. The data were acquired using a standard laboratory Fourier transform near-infrared spectrometer.
[0028] In this example, the pyrolysis signal vector of each sample of the pyrolysis data has a dimension of 8001, that is... The results are obtained by interpolating the thermogravimetric mass-temperature curve to a uniform temperature grid (e.g., 60–800 °C, with a step size of 0.1 °C). Thermogravimetric analysis can be performed using a Discovery series thermogravimetric analyzer (TA Instruments, USA), with a nitrogen atmosphere and a flow rate of 60.0 mL / min. The heating program is as follows: increase from room temperature to 100 °C at a rate of 30 °C / min and hold for 5 min to completely remove free water, then increase from 100 °C to 800 °C at a rate of 10 °C / min, and finally allow to cool naturally to approximately 60 °C.
[0029] In this example, multiple chemical components were measured for each sample, including but not limited to nine indicators such as total sugar, reducing sugar, nicotine, chlorine, potassium, total nitrogen, pH value, chlorogenic acid, and starch. These nine indicators were then combined to form a label vector. Stacking yields a label matrix In this embodiment, n=499.
[0030] Step S21 is used to preprocess the near-infrared spectral data. Specifically, in this example, baseline correction can be performed by fitting the baseline of each spectrum using a first-order polynomial, and then subtracting this baseline from the original spectrum to eliminate baseline drift and tilt, allowing spectral features to be compared on the same horizontal benchmark. Noise reduction can be performed by using a moving window and a second-order polynomial for local fitting to filter out high-frequency random noise, while simultaneously calculating the second derivative spectrum. Scattering correction can be performed by using a standard normal variable transformation to eliminate the multiplicative effect caused by scattering within each spectrum. Feature normalization can be performed by first truncating the bands, removing the low signal-to-noise ratio segments at both ends according to the noise level, for example, truncating the range of 9000–4000 cm⁻¹. Then, zero-mean, unit-variance normalization is performed on each wavelength dimension according to the modeling set samples. After preprocessing, the near-infrared data matrix is denoted as... ,in Or extract the dimensions after truncation.
[0031] Step S22 is used to preprocess the pyrolysis data. Specifically, in this example, baseline correction can be performed by using a low-order polynomial to fit the flat, non-pyrolysis reaction region of the curve to obtain a baseline. This baseline is then subtracted from the original TGA data. Interpolation and alignment can be performed by interpolating the original TGA curve to a uniform temperature grid [60, 800]℃ with a step size of 0.1℃ to obtain a vector of fixed length 8001 dimensions. Normalization can be performed by converting the mass loss curve to relative mass or relative mass loss percentage with an initial mass of 100%. Optional peak feature extraction can be performed by identifying peaks and recording peak position, peak height, and peak area based on the first derivative curve, but in this preferred embodiment, the entire curve vector is directly used as input. Feature standardization can be performed by standardizing each temperature point dimension with zero mean and unit variance according to the modeling set samples. After preprocessing, the pyrolysis data matrix is denoted as... .
[0032] The label vector consisting of nine indicators is standardized according to the modeling set samples (subtract the mean and divide by the standard deviation) to enhance the comparability of loss magnitudes between different components.
[0033] In step S11, a multimodal self-consistent coding rate reduction model is constructed and trained using preprocessed sample data. In this embodiment, the specific method for constructing and training the multimodal self-consistent coding rate reduction model in step S11 can be of various forms known to those skilled in the art. In one example of the present invention, step S11 may include, for example... Figure 3 The steps shown are described in this. Figure 3 In this context, step S11 may include: In step S30, multimodal fusion encoding is performed based on the preprocessed tobacco sample data to obtain the potential representation vector; In step S31, modality decoding and chemical composition prediction training are performed based on the latent representation vector; In step S32, the multimodal self-consistent coding rate reduction model is trained with closed-loop self-consistent constraints. In step S33, the coding rate of the multimodal self-consistent coding rate reduction model is reduced. In step S34, a total loss function is constructed to train the multimodal self-consistent coding rate reduction model.
[0034] In such Figure 3 In the method shown, step S30 is used for multimodal fusion coding to obtain the latent representation vector. In this example, a dual-branch multimodal fusion encoder is employed. The network comprises an NIR branch coding network, a TGA branch coding network, and a fusion layer and a shared latent layer. The NIR branch coding network structure includes an input layer, a first fully connected layer, a first activation layer, a second fully connected layer, a second activation layer, a third fully connected layer, and a third activation layer. The input layer provides a 1609-dimensional vector for each sample. The first fully connected layer reduces the dimension from 1609 to 512, the second fully connected layer further compresses the dimension from 512 to 128, and the third fully connected layer reduces the dimension to 64. The output of the NIR branch coding network is denoted as... The TGA branch-coding network structure includes an input layer, a fourth fully connected layer, a fourth activation layer, a fifth fully connected layer, a fifth activation layer, a sixth fully connected layer, and a sixth activation layer. The input layer provides an 8001-dimensional vector for each sample. The fourth fully connected layer reduces the dimension from 8001 to 1024, the fifth fully connected layer reduces the dimension from 1024 to 256, and the sixth fully connected layer reduces the dimension from 256 to 64. The output of the TGA branch-coding network is denoted as... The fusion layer and the shared latent layer concatenate the outputs of the two branches to obtain... By sharing the fusion layer and the latent layer, the dimension is reduced to 32. Therefore, this embodiment selects the latent dimension. The latent vector of a single sample is Stacking .
[0035] Step S31 is used to perform modality decoding and chemical composition prediction training based on the latent representation vector. In this embodiment, the specific method for performing modality decoding and chemical composition prediction training in step S31 can be of various forms known to those skilled in the art. In one example of the present invention, step S31 may include, for example... Figure 4 The steps shown are described in this. Figure 4 In this context, step S31 may include: In step S40, the reconstructed heat decomposition vector is obtained using a heat decomposition mode decoder based on the latent representation vector; In step S41, a near-infrared modal decoder is used to obtain the reconstructed near-infrared vector based on the latent representation vector; In step S42, based on the potential representation vector, the chemical composition prediction branch is used to obtain the chemical composition prediction vector; In step S43, the model is trained using supervised regression loss based on the chemical composition prediction vector; In step S44, the model is trained using modal reconstruction loss based on the reconstructed thermal vector and the reconstructed near-infrared vector.
[0036] In such Figure 4 In the method shown, step S40 is used to employ a thermal mode decoder. Low-dimensional latent representation vector Expanding back to the original data space. Specifically, in this example, the thermal modal decoder structure is symmetrical to but in reverse of the encoder, taking a 32-dimensional latent representation vector as input. Hidden layer 1 extends the 32-dimensional structure to 128-dimensionality through a fully connected layer and introduces non-linearity using the ReLU activation function. Hidden layer 2 extends the 128-dimensional structure to 512-dimensionality through a fully connected layer, using ReLU. The output layer directly maps the 512-dimensional structure to the target dimension of 8001 dimensions through a fully connected layer, using linear activation. The output reconstructs the heatmap vector. Step S41 is used to employ a near-infrared mode decoder. Low-dimensional latent representation vector Expand back to the original data space. Specifically, in this example, the near-infrared mode decoder... Structure and Similar, but with a different output dimension to accommodate the characteristics of near-infrared data. Input is a 32-dimensional latent representation vector. The two hidden layers expand the dimensions to 128 and 256 respectively through their fully connected layers, and the output layer expands the dimensions to 1609, outputting a reconstructed near-infrared vector. Step S42 is used to predict the branch using chemical composition. Obtain the chemical composition prediction vector. The structure is a smaller MLP, with a 32-dimensional latent representation vector as input. The dimension is expanded to 64 dimensions through the fully connected layer of the hidden layer, and the dimension is reduced to 9 dimensions through the output layer, corresponding to the content of 9 chemical components that need to be predicted.
[0037] Step S43 is used to train the model using supervised regression loss based on the chemical composition prediction vector. Specifically, in this example, the supervised regression loss can be as shown in formula (12): (12) in, To monitor the regression loss, This is a true chemical ingredient label matrix. for The obtained predicted chemical composition label matrix, The square of the Frobenius norm is calculated by summing the squares of the differences between all elements of two matrices. This represents the number of samples.
[0038] Step S44 is used to train the model using modal reconstruction loss based on the reconstructed thermal vector and the reconstructed near-infrared vector. Specifically, in this example, the modal reconstruction loss can be as shown in formula (13): (13) in, For modal reconstruction loss, This indicates the reconstruction error of the TGA data. This indicates the reconstruction error of NIR data; This is the raw near-infrared data. This is the raw pyrolysis data. Near-infrared data reconstructed by the decoder, The pyrolysis data reconstructed by the decoder.
[0039] Through steps S43 and S44, It must contain enough information to accurately reconstruct the original data, preventing... Focusing solely on chemical composition prediction while discarding other important structural information of the data ensures the potential representation vector... Information integrity.
[0040] Step S32 is used to perform closed-loop self-consistency constraint training on the multimodal self-consistent coding rate reduction model. Specifically, in this example, the specific method for performing closed-loop self-consistency constraint training in step S32 may include the following steps: In step S50, a semi-reconstructed input is constructed and re-encoded for each sample according to formulas (1) and (2). (1) (2) in, for and The potential representation of fusion encoding, for and The potential representation of fusion coding; In step S51, self-consistent constraint training is performed according to formula (3). (3) in, For the loss of spatial self-consistency, For the latent representation vector, This represents the number of samples.
[0041] Step S50 is used to construct two sets of mixed inputs: one set using reconstructed TGA data and original NIR data, and the other set using original TGA data and reconstructed NIR data. These are then compared in step S51. , and The consistency of these three potential representations.
[0042] Step S33 is used to reduce the coding rate of the multimodal self-consistent coding rate reduction model. Specifically, in this example, the specific method for coding rate reduction in step S33 can be to... The coding rate under unlabeled prior is defined as shown in formula (4): (4) in To preset hyperparameters, It is a 32-dimensional identity matrix. In this example, it is preferred. .
[0043] Set the tag similarity weight matrix ,in This is the standardized 9-dimensional component vector. In this example, The optimal value is obtained by selecting through a grid search on {0.5, 1.0, 2.0}. .
[0044] Calculate the label weighted coding rate according to formula (5): (5) Then, the reduction in coding rate is calculated using formula (6): (6) in, The reduction is approximately due to the coding rate. For unlabeled coding rate, Weighted coding rate for labels, It is the identity matrix. For hyperparameters, This is the label similarity weight matrix. The larger the value, the richer the amount of information in the latent representation that is aligned with the label structure.
[0045] Step S34 is used to construct the total loss function for training the multimodal self-consistent coding rate reduction model. Specifically, in this example, the total loss function can be constructed, for example, using formula (7): (7) in, For the total loss function, To monitor the regression loss, To monitor the regression loss weights, For the loss of spatial self-consistency, For spatial self-consistency, the loss weight is... For modal reconstruction loss, For modal reconstruction loss weights, The reduction is approximately due to the coding rate. The weighting is reduced by the coding rate. The square of the L2 norm of all parameters in the model. For the complete set of trainable parameters, This is the weight decay coefficient. In this example, the preferred parameter is... , , , , .
[0046] In this example, the training algorithm and parameter selection in step S34 could be: using the Adam optimizer; the initial learning rate is... The training parameters are decayed by 0.5 times every 50 epochs; the batch size is 32; the maximum number of training epochs is 500 epochs; the early stopping strategy is to terminate training if the total loss does not decrease significantly for 30 consecutive epochs on the validation set. For model selection, 5-fold cross-validation can be used on 399 modeling set samples to select the hyperparameter combination that best combines RMSE and R2 on the validation set. Then, the model is retrained on the entire modeling set, and the final performance is reported on the independent test set.
[0047] After completing the above training and obtaining a stable fusion latent representation Then, modal redundancy and complementarity analysis is performed in the latent space through step S12. Specifically, in this example, the specific method for performing modal redundancy and complementarity analysis includes the following steps: In step S60, a single-modal encoder is constructed and alignment training is performed; In step S61, the cooperative gain is calculated; In step S62, the redundancy rate and complementarity index based on partial information decomposition are calculated; In step S63, the complementarity index based on mutual information decomposition is calculated; In step S64, the complementarity gain based on the coding rate minus the difference is calculated.
[0048] Step S60 is used to construct a single-modal encoder and perform alignment training. Specifically, in this example, this could involve constructing a near-infrared single-modal encoder separately. and pyrolysis single-mode encoder And acquire near-infrared single-mode code and pyrolysis single-mode code While freezing the parameters of the multimodal fusion encoder, only the parameters of the single-modal encoder are trained. To align the loss. In this example, alignment training could be performed using, for example, formula (14) as the loss function: (14) Step S61 is used to calculate the cooperative gain, specifically, in this example, it can be based on... Construct a chemical composition prediction network (similar in structure to the prediction branch in step S31), and compute on the test set. , and : (9) (10) (11) in, These are predicted values based on pyrolysis characteristics. These are predicted values based on near-infrared features. These are predicted values based on fusion features. This represents the true value of the chemical composition. for and The root mean square error between them for and The root mean square error between them for and The root mean square error between them.
[0049] Define the modal cooperative gain as: (8) like This indicates that the fusion model outperforms the optimal single-mode model in terms of prediction error, demonstrating a positive synergistic and complementary effect; if If the result is negative, it indicates that the performance improvement brought about by the fusion is limited or that there is negative migration.
[0050] Step S62 is used to calculate the redundancy rate and complementarity index based on partial information decomposition. Specifically, in this example, a partial information decomposition (PID) framework can be used to decompose the two-modal latent representations. For target chemical components Total mutual information Decomposed into: Redundant information Used to represent task-related information shared between the two modalities; unique information (corresponding to the unique information of each single modality); and collaborative information. The complementary information can only be obtained through the joint use of two modes. At the implementation level, existing PID estimation methods such as Minimum Mutual Information (MMI) or BROJA optimization can be used. For continuous variable scenarios, Gaussian approximation or numerical optimization based on entropy estimation can be combined. The redundancy rate is calculated according to formulas (15) and (16). and complementarity index , (15) (16) in, Indicates about General information The amount of redundant information, The amount of collaborative information. The higher the value, the higher the redundancy of the two modalities in task-related information; The higher the value, the more significant the collaborative behavior and the stronger the complementarity.
[0051] Step S63 is used to calculate the complementarity index based on mutual information decomposition. To simplify the calculation, in this example, we use... The Kraskov–Stögbauer–Grassberger (KSG) mutual information estimation method, using a nearest neighbor number k=5, estimates the following in the latent space: , , , Define a task-related complementarity index based on mutual information decomposition: (17) in This indicates that fusion brings positive synergistic benefits. This indicates that modal redundancy is dominant.
[0052] Step S64 is used to calculate the complementary gain based on the coding rate minus the difference. Based on step S33, the single-modal latent representation is... Same definition: (18) (19) (20) ,(twenty one) Based on this, the approximate reduction in their respective coding rates can be obtained: ,(twenty two) ,(twenty three) ,(twenty four) And define the complementarity gain index: (25) when This indicates that the fusion latent representation surpasses any single modality in terms of compressibility and label alignment, demonstrating complementary gains at the information theory level.
[0053] On the other hand, the present invention also provides a tobacco multimodal coding rate reduction and information theory evaluation system, the evaluation system including a processor configured to perform any of the methods described in the tobacco multimodal coding rate reduction and information theory evaluation method.
[0054] In another aspect, the present invention also provides a computer-readable storage medium storing instructions that, when executed by a processor, implement any of the methods described in the tobacco multimodal coding rate reduction and information theory evaluation methods.
[0055] The beneficial effects of this invention are: The embodiments of the present invention use a multimodal fusion encoder to map high-dimensional NIR and TGA data to a unified low-dimensional latent representation space, and perform prediction modeling, coding rate estimation and mutual information analysis in this space, which significantly alleviates the problems of dimensionality curse and noise interference, and improves the stability and interpretability of information theory estimation.
[0056] The embodiments of the present invention employ a semi-reconstruction-recoding closed-loop structure and a latent spatial self-consistent loss to constrain the latent representation to remain consistent under different modal combination inputs, effectively avoiding the dominant effect of a single modality in representation, making multimodal fusion more sufficient, and helping to improve prediction accuracy and the robustness of the model to missing modes and noisy modes.
[0057] The coding rate is reduced by approximately [percentage missing] by introducing continuous tag guidance in the embodiments of the present invention. This extends the idea of reducing the coding rate, which was originally applied to discrete categories, to scenarios with continuous response variables such as chemical composition. This makes the learning process of latent representation both compressible and aligned with the geometric structure of the label, providing a clear information theory explanation for the model.
[0058] The embodiments of the present invention comprehensively utilize cooperative gain, conditional mutual information, mutual information decomposition indices (including redundancy rate and cooperative rate based on PID, complementarity index CI based on mutual information decomposition), and complementary gain based on coding rate reduction difference in the latent space. These indicators systematically characterize the redundancy and complementarity of NIR and TGA modes from both predictive performance and information theory perspectives, providing quantifiable decision-making basis for instrument configuration optimization, detection scheme design, and cost-performance trade-offs.
[0059] The embodiments of this invention are based on a general machine learning and information theory framework, and can be seamlessly integrated with existing near-infrared spectrometers, pyrolysis analyzers and data acquisition systems. It is not only applicable to tobacco quality evaluation, but can also be smoothly extended to multimodal quality evaluation scenarios of other agricultural products, food and complex industrial raw materials, and has good engineering feasibility and promotion value.
[0060] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0061] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0062] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0063] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0064] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0065] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0066] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0067] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0068] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for reducing the multimodal coding rate of tobacco leaves and evaluating it using information theory, characterized in that, The evaluation methods include: Acquire tobacco leaf sample data and preprocess it; A multimodal self-consistent coding rate reduction model was constructed and trained using preprocessed sample data; The trained multimodal self-consistent coding rate reduction model is used to obtain fusion features, and modal redundancy and complementarity analysis is performed.
2. The evaluation method according to claim 1, characterized in that, Constructing a multimodal self-consistent coding rate reduction model and training it using preprocessed sample data includes: Based on the preprocessed tobacco sample data, multimodal fusion encoding is performed to obtain the potential representation vector; Based on the latent representation vector, perform modality decoding and chemical composition prediction training; The multimodal self-consistent coding rate reduction model is trained with closed-loop self-consistent constraints. The coding rate of the aforementioned multimodal self-consistent coding rate reduction model is reduced. The total loss function is constructed to train the multimodal self-consistent coding rate reduction model.
3. The evaluation method according to claim 2, characterized in that, Based on the latent representation vector, the training for modality decoding and chemical composition prediction includes: Based on the latent representation vector, a heat dissipation mode decoder is used to obtain the reconstructed heat dissipation vector; Based on the latent representation vector, a near-infrared modal decoder is used to obtain the reconstructed near-infrared vector; Based on the potential representation vector, a chemical composition prediction vector is obtained using the chemical composition prediction branch; Based on the predicted chemical composition vector, the model is trained using supervised regression loss; Based on the reconstructed thermal vector and the reconstructed near-infrared vector, the model is trained using modal reconstruction loss.
4. The evaluation method according to claim 2, characterized in that, The closed-loop self-consistent constraint training of the multimodal self-consistent coding rate reduction model includes: For each sample, a semi-reconstructed input is constructed and re-encoded according to formulas (1) and (2). ,(1) ,(2) in, This is the raw near-infrared data. This is the raw pyrolysis data. Near-infrared data reconstructed by the decoder, For the pyrolysis data reconstructed by the decoder, for and The potential representation of fusion encoding, for and The potential representation of fusion coding; Self-consistent constraint training is performed according to formula (3). ,(3) in, For the loss of spatial self-consistency, For the latent representation vector, This represents the number of samples.
5. The evaluation method according to claim 2, characterized in that, The coding rate reduction of the aforementioned multimodal self-consistent coding rate reduction model includes: The reduction in coding rate is obtained from formulas (4) to (6). ,(4) ,(5) ,(6) in, The reduction is approximately due to the coding rate. For unlabeled coding rate, Weighted coding rate for labels, It is the identity matrix. For hyperparameters, This is the label similarity weight matrix.
6. The evaluation method according to claim 2, characterized in that, Constructing the total loss function to train the multimodal self-consistent coding rate reduction model includes: Construct the total loss function according to formula (7). ,(7) in, For the total loss function, To monitor the regression loss, To monitor the regression loss weights, For the loss of spatial self-consistency, For spatial self-consistency, the loss weight is... For modal reconstruction loss, For modal reconstruction loss weights, The reduction is approximately due to the coding rate. The weighting is reduced by the coding rate. The square of the L2 norm of all parameters in the model. This is the weight decay coefficient.
7. The evaluation method according to claim 1, characterized in that, The trained multimodal self-consistent coding rate reduction model is used to obtain fusion features, and modal redundancy and complementarity analysis is performed, including: Near-infrared single-mode encoders and pyrolysis single-mode encoders were constructed respectively, and near-infrared single-mode codes and pyrolysis single-mode codes were obtained; The near-infrared single-mode encoder and the pyrolysis single-mode encoder are trained using the fusion features; A chemical composition prediction network is constructed based on the near-infrared single-mode code, pyrolysis single-mode code, and fusion features. Based on the chemical composition prediction network, predicted values based on pyrolysis features, predicted values based on near-infrared features, and predicted values based on fusion features are obtained respectively. Calculate the cooperative gain according to formulas (8) to (11). ,(8) ,(9) ,(10) ,(11) in, For synergistic gain, These are predicted values based on pyrolysis characteristics. These are predicted values based on near-infrared features. These are predicted values based on fusion features. This represents the true value of the chemical composition. for and The root mean square error between them for and The root mean square error between them for and The root mean square error between them.
8. The evaluation method according to claim 7, characterized in that, The trained multimodal self-consistent coding rate reduction model is used to obtain fusion features, and modal redundancy and complementarity analysis is performed, including: Calculate the redundancy rate and complementarity index based on partial information decomposition; Calculate the complementarity index based on mutual information decomposition; Calculate the complementary gain based on coding rate minus difference.
9. A multimodal coding rate reduction and information theory evaluation system for tobacco leaves, characterized in that, The evaluation system includes a processor configured to perform the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 8.