Model training method, model training device, tobacco flavor type classification method, tobacco flavor type classification device and tobacco flavor type classification system
By constructing a multimodal deep neural network model based on near-infrared spectroscopy and chemical composition, the problem of strong subjectivity and instability of the classification of smoky fragrance types in the prior art is solved, and efficient and accurate classification of scented fragrance types is achieved.
Patent Information
- Application Number
- CN202510216063.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, the classification of flavour type of tablets mainly relies on the manual suction judgment of formulaters, and there are problems such as strong subjectivity, high working intensity, and difficulty in quantification, resulting in inconsistency and instability of flavour type classification.
By obtaining the near-infrared spectra and chemical composition of different fragrance-type tobacco leaf samples, pretreatment and feature extraction, a tobacco leaf fragrance classification model based on near-infrared spectra and chemical composition was constructed, and a multimodal deep neural network was used for training to achieve automatic classification of fragrance.
It improves the scientificity and accuracy of fragrance classification, reduces manual misjudgment, ensures the consistency of the quality of leaf group formulas, and improves the efficiency and reliability of fragrance classification.
Smart Images

Figure CN120105103A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of tobacco leaf processing, and in particular to a model training method and device, and a tobacco leaf flavor classification method, device and system. Background Art
[0002] Aroma type refers to the overall aroma type or flavor style of the smoke during the burning process of tobacco leaves. The formation of aroma type is closely related to the ecological environment, internal chemical composition, cultivation management and processing technology of tobacco leaves.
[0003] The classification of tobacco leaf flavors is crucial to cigarette production. Accurate classification of tobacco leaf flavors can ensure the clear and stable flavor characteristics of tobacco leaves from different production areas, and manufacturers can more accurately select and adjust leaf group formulas to improve the flavor stability of leaf groups. Summary of the invention
[0004] The inventors found through research that the classification of tobacco flavors in related art mainly relies on manual puffing judgment by formulators, which has the following disadvantages: strong subjectivity. Manual puffing judgment relies on the formulator's personal senses and experience. Different people may have different perceptions and evaluation criteria for flavors, which is prone to individual subjective bias, resulting in inconsistent and unstable flavor classification.
[0005] In view of at least one of the above technical problems, the present disclosure provides a model training method and device, a tobacco flavor classification method, device and system, which can quantitatively evaluate tobacco from multiple angles and effectively improve the scientificity and accuracy of flavor classification.
[0006] According to one aspect of the present disclosure, a model training method is provided, comprising:
[0007] Obtaining near infrared spectra and chemical compositions of tobacco leaf samples with different aroma types, and labeling the tobacco leaf samples with different aroma types;
[0008] Preprocessing the near infrared spectra and chemical components of the tobacco leaf samples with different aroma types to form training data;
[0009] Constructing the tobacco leaf aroma classification model, wherein the tobacco leaf aroma classification model is a tobacco leaf aroma classification model based on near infrared spectroscopy and chemical composition;
[0010] The training data is used to train the tobacco leaf aroma classification model, wherein the trained tobacco leaf aroma classification model is used for tobacco leaf aroma classification.
[0011] In some embodiments of the present disclosure, the tobacco leaf aroma classification model includes a near infrared spectroscopy feature extraction module, a chemical component feature extraction module, a feature fusion module and an aroma classification module, and the training of the tobacco leaf aroma classification model using the training data includes:
[0012] Divide the training data into training set, validation set and test set;
[0013] The near infrared spectrum and chemical composition data in the training set are processed through various modules of the tobacco leaf aroma classification model to obtain the output result of the tobacco leaf aroma classification model, thereby realizing the training of the tobacco leaf aroma classification model.
[0014] In some embodiments of the present disclosure, the near infrared spectrum and chemical composition data in the training set are processed by various modules of the tobacco flavor classification model to obtain the output result of the tobacco flavor classification model, which includes:
[0015] The near infrared spectrum feature extraction module is used to extract features from the near infrared spectrum data of the tobacco leaf samples with different aroma types;
[0016] Using the chemical component feature extraction module to perform feature extraction on the chemical component data of the tobacco leaf samples with different aroma types;
[0017] The feature fusion module is used to splice and fuse the extracted near-infrared spectrum features and chemical composition features;
[0018] The aroma classification module is used to classify the aroma of tobacco leaf samples according to the spliced and fused features.
[0019] In some embodiments of the present disclosure, the near infrared spectral feature extraction module is used to extract features from the near infrared spectral data of the different flavor tobacco leaf samples, including:
[0020] After adding position coding to the near infrared spectrum data, a first near infrared spectrum feature is obtained;
[0021] Performing a linear transformation on the first near infrared spectral feature to obtain a query matrix, a key matrix and a value matrix;
[0022] Determine the attention weight by calculating the correlation between the query matrix and the key matrix;
[0023] Performing weighted summation on the value matrix using the attention weights to obtain a weighted output value;
[0024] The weighted output values are residually connected and normalized to obtain a second near-infrared spectrum feature.
[0025] In some embodiments of the present disclosure, the step of extracting features of chemical components of tobacco leaf samples with different aroma types using the chemical component feature extraction module includes:
[0026] Multilayer perceptron was used to extract features of the chemical components of tobacco leaf samples.
[0027] In some embodiments of the present disclosure, the using the feature fusion module to splice and fuse the extracted near-infrared spectral features and chemical composition features includes:
[0028] Vector splicing is used to perform feature fusion on the near infrared spectral features output by the near infrared spectral feature extraction module and the chemical composition features output by the chemical composition feature extraction module.
[0029] In some embodiments of the present disclosure, the preprocessing of the near infrared spectra and chemical components of tobacco leaf samples with different flavors includes:
[0030] Performing feature scaling on the chemical component data of the tobacco leaf samples with different flavors to unify the features of different magnitudes and dimensions into a comparable range;
[0031] Selecting standard normal transformation to preprocess the spectral data of the tobacco leaf samples with different aromas;
[0032] One-dimensional convolution is used to extract local features from the near-infrared spectra of the different flavor tobacco samples.
[0033] In some embodiments of the present disclosure, the feature scaling of the chemical composition data of tobacco leaf samples with different flavors includes:
[0034] determining a median and an interquartile range for a data set of the chemical composition data;
[0035] Determine a data difference according to the chemical composition data and the median;
[0036] The scaled feature data is determined according to the ratio of the data difference value to the interquartile range.
[0037] In some embodiments of the present disclosure, the flavor categories of the tobacco leaves include at least one of light sweet flavor, honey sweet flavor, mellow sweet flavor, caramel sweet flavor, caramel sweet and mellow sweet flavor, light sweet sweet flavor, honey sweet caramel flavor, woody honey sweet flavor, Zimbabwe type, Brazilian type, Argentine type, Zambian type, and American type.
[0038] According to another aspect of the present disclosure, a method for classifying tobacco leaf aroma types is provided, comprising:
[0039] Obtain near infrared spectra and chemical compositions of tobacco leaves to be classified;
[0040] Preprocessing the near infrared spectrum and chemical composition of the tobacco leaves to be classified;
[0041] The trained tobacco leaf aroma classification model is used to classify the tobacco leaves to be classified according to the pre-processed near infrared spectrum and chemical composition.
[0042] In some embodiments of the present disclosure, the preprocessing of the near infrared spectrum and chemical composition of the tobacco leaves to be classified includes:
[0043] Performing feature scaling on the chemical component data of the tobacco leaves to be classified, unifying features of different magnitudes and dimensions into a comparable range;
[0044] Selecting standard normal transformation to preprocess the spectral data of the tobacco leaves to be classified;
[0045] One-dimensional convolution is used to extract local features from the near-infrared spectrum of the tobacco leaves to be classified.
[0046] In some embodiments of the present disclosure, the aroma classification of the tobacco leaves to be classified includes:
[0047] Extracting features from the near-infrared spectrum data of the tobacco leaves to be classified;
[0048] Performing feature extraction on the chemical composition data of the tobacco leaves to be classified;
[0049] The extracted near-infrared spectral features and chemical composition features are spliced and fused;
[0050] The aroma types of tobacco leaf samples are classified according to the characteristics after splicing and fusion.
[0051] In some embodiments of the present disclosure, the feature extraction of the near infrared spectrum data of the tobacco leaves to be classified includes:
[0052] After adding position coding to the near infrared spectrum data, a first near infrared spectrum feature is obtained;
[0053] Performing a linear transformation on the first near infrared spectral feature to obtain a query matrix, a key matrix and a value matrix;
[0054] Determine the attention weight by calculating the correlation between the query matrix and the key matrix;
[0055] Performing weighted summation on the value matrix using the attention weights to obtain a weighted output value;
[0056] The weighted output values are residually connected and normalized to obtain a second near-infrared spectrum feature.
[0057] According to another aspect of the present disclosure, a model training device is provided, comprising:
[0058] A first data acquisition unit is configured to acquire near infrared spectra and chemical components of tobacco leaf samples with different aroma types, and to label the tobacco leaf samples with different aroma types with aroma types;
[0059] A first preprocessing unit is configured to preprocess the near infrared spectra and chemical components of the tobacco leaf samples with different flavors to form training data;
[0060] A model building unit, configured to build the tobacco aroma classification model;
[0061] The model training unit is configured to use the training data to train the tobacco leaf flavor classification model, wherein the trained tobacco leaf flavor classification model is used for tobacco leaf flavor classification.
[0062] According to another aspect of the present disclosure, a model training device is provided, comprising:
[0063] a memory configured to store instructions; and
[0064] The processor is configured to execute the instructions so that the model training device implements the model training method as described in any of the above embodiments.
[0065] According to another aspect of the present disclosure, there is provided a tobacco leaf aroma classification device, comprising:
[0066] A second data acquisition unit is configured to acquire the near infrared spectrum and chemical composition of the tobacco leaves to be classified;
[0067] A second preprocessing unit is configured to preprocess the near infrared spectrum and chemical composition of the tobacco leaves to be classified;
[0068] The aroma classification unit is configured to use a trained tobacco aroma classification model to classify the aroma of the tobacco to be classified according to the pre-processed near infrared spectrum and chemical composition.
[0069] According to another aspect of the present disclosure, there is provided a tobacco leaf aroma classification device, comprising:
[0070] a memory configured to store instructions; and
[0071] The processor is configured to execute the instructions so that the tobacco aroma classification device implements the tobacco aroma classification method as described in any of the above embodiments.
[0072] According to another aspect of the present disclosure, a tobacco flavor classification system is provided, comprising a model training device as described in any of the above embodiments, and a tobacco flavor classification device as described in any of the above embodiments.
[0073] According to another aspect of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the tobacco leaf aroma classification method as described in any of the above embodiments is implemented.
[0074] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the model training method as described in any of the above embodiments and / or the tobacco leaf aroma classification method as described in any of the above embodiments is implemented.
[0075] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the model training method as described in any of the above embodiments and / or the tobacco leaf aroma classification method as described in any of the above embodiments is implemented.
[0076] The present invention can quantitatively evaluate tobacco leaves from multiple angles, effectively improving the scientificity and accuracy of aroma classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0078] Figure 1 Schematic diagram of some embodiments of the tobacco leaf aroma classification method disclosed in the present invention.
[0079] Figure 2 Schematic diagram of the tobacco aroma classification model in some embodiments of the present disclosure.
[0080] Figure 3 Schematic diagram of some embodiments of the tobacco leaf aroma classification method disclosed in the present invention.
[0081] Figure 4 Schematic diagram of the structure of some embodiments of the model training device disclosed in the present invention.
[0082] Figure 5 Schematic diagram of the structure of other embodiments of the model training device disclosed in the present invention.
[0083] Figure 6 Schematic diagram of the structure of some embodiments of the tobacco flavor classification device disclosed in the present invention.
[0084] Figure 7 Schematic diagram of the structure of other embodiments of the tobacco flavor classification device disclosed in the present invention.
[0085] Figure 8 Schematic diagram of the structure of some embodiments of the tobacco flavor classification system disclosed in the present invention. DETAILED DESCRIPTION
[0086] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present disclosure and its application or use. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0087] Unless specifically stated otherwise, the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure.
[0088] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0089] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered as part of the specification.
[0090] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0091] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0092] The inventors also found through research that the classification of tobacco flavors in related art mainly relies on manual puffing judgment by formulators, which has the following disadvantages:
[0093] The work intensity is high. Manual smoking requires the tobacco leaves to be made into cigarettes manually, which is a heavy burden for manual operation. In addition, the formulator needs to repeatedly evaluate the smoke to judge the flavor, which takes a lot of time and energy. Therefore, the manual smoking judgment method limits the accuracy and speed of tobacco leaf flavor classification, especially when faced with a large number of tobacco leaf samples, the efficiency problem is particularly prominent.
[0094] Difficult to quantify. The sensory evaluation of flavors is difficult to quantify and standardize. Formulators may make qualitative descriptions of flavors based on sensory experience, but lack quantitative data support, which has certain limitations for later flavor optimization and precise blending. This problem is more prominent when multiple formulators collaborate. In particular, tobacco leaves from non-core production areas may have multiple regional flavor characteristics or partial characteristics at the same time. Traditional manual classification methods can only simply merge them into a certain category. The softmax soft label of deep learning can vividly represent tobacco leaves at the edges of different style divisions.
[0095] It can be seen from this that it is very important to use unified quantitative rules to characterize the aroma of tobacco leaves.
[0096] In view of at least one of the above technical problems, the present disclosure provides a model training method and device, a tobacco flavor classification method, device and system, which are described below through embodiments.
[0097] Figure 1 Schematic diagram of some embodiments of the model training method disclosed in the present invention. Figure 1 The embodiment can be executed by the model training device disclosed in the present invention. Figure 1 As shown, Figure 1 The method of the embodiment may include at least one of steps 110 to 140 .
[0098] Step 110, obtaining near infrared spectra and chemical compositions of tobacco leaf samples with different aroma types, and labeling the tobacco leaf samples with different aroma types.
[0099] In some embodiments of the present disclosure, the tobacco leaves may be tobacco leaves in the form of sheet tobacco or leaf groups.
[0100] In some embodiments of the present disclosure, step 110 may include: collecting near-infrared spectra of tobacco leaf samples with different aroma types; and obtaining chemical components of tobacco leaf samples with different aroma types.
[0101] In some embodiments of the present disclosure, the near infrared spectrum collection method is as follows: take an appropriate amount of tobacco leaves or tobacco powder made by grinding leaf groups into a sample cup, and press the sample with a 500 g copper sample press. The NIR spectrum of the sample is collected, with a resolution of 8 cm-1, a spectrum collection range of 4,000 to 10,000 cm-1, 64 scans, and each sample is scanned twice continuously, and the spectrum data is averaged. The near infrared spectrum of the sample is formed by collecting 1557 wavelength points.
[0102] In some embodiments of the present disclosure, the chemical components of different flavor tobacco samples include: total alkaloids, reducing sugars, total sugars, total nitrogen, K, chlorine, pH, starch, dichloromethane extract, solanesol, sulfate, phosphate, Mg, Ca, neochlorogenic acid, chlorogenic acid, cryptochlorogenic acid, scopoletin, rutin, oxalic acid, malonic acid, succinic acid, malic acid, citric acid, vanillic acid, myristic acid, hexadecanoic acid, linoleic acid, oleic acid + linolenic acid, octadecanoic acid, eicosanoic acid, aspartic acid, threonine, serine, asparagine, glutamic acid, glutamine, glycine, alanine, valine, cystine, methionine A total of 70 tobacco leaf chemical components were detected, including Fru-Glu, Fru-Ile, Fru-Leu, Fru-Tyr, Fru-Phe, Fru-Trp, and neophytodienoic acid.
[0103] In some embodiments of the present disclosure, the flavor categories of the tobacco leaves may include at least one of a light sweet flavor, a honey sweet flavor, a mellow sweet flavor, a caramel sweet flavor, a caramel sweet and mellow sweet flavor, a light sweet flavor, a honey sweet caramel flavor, a woody honey sweet flavor, a Zimbabwean type, a Brazilian type, an Argentinian type, a Zambian type, and an American type.
[0104] Step 120, preprocessing the near infrared spectra and chemical components of the tobacco leaf samples with different aroma types to form training data.
[0105] In some embodiments of the present disclosure, step 120 may include at least one of steps 121 to 123 .
[0106] Step 121, chemical component pretreatment.
[0107] In some embodiments of the present disclosure, step 121 may include: performing feature scaling on the chemical composition data of the tobacco leaf samples with different flavors, and unifying features of different magnitudes and dimensions into a comparable range.
[0108] In some embodiments of the present disclosure, step 121 may include: scaling the chemical composition data to unify the features of different magnitudes and dimensions into a comparable range. The chemical composition of the tobacco leaf sample includes a total of 70 intrinsic component data, including: total alkaloids, reducing sugars, total sugars, total nitrogen, K, chlorine, pH, starch, dichloromethane extract, solanesol, sulfate, phosphate, Mg, Ca, neochlorogenic acid, chlorogenic acid, cryptochlorogenic acid, scopoletin, rutin, oxalic acid, malonic acid, succinic acid, malic acid, citric acid, vanillic acid, tetradecanoic acid, hexadecanoic acid, linoleic acid, oleic acid + linolenic acid, octadecanoic acid, eicosanoic acid, aspartic acid, threonine, serine, asparagine, glutamic acid, glutamine, glycine, alanine, valine, cystine , methionine, isoleucine, leucine, tyrosine, phenylalanine, 4-aminobutyric acid, lysine, histidine, tryptophan, arginine, proline, Glu-An, glutamate, Fru-Amb, Fru-His, Fru-Pro, Fru-Val, Fru-Thr, Fru-Gly, Fru-Ala, Fru-Asn, Fru-Asp, Fru-Gln, Fru-Glu, Fru-Ile, Fru-Leu, Fru-Tyr, Fru-Phe, Fru-Trp, neophytadiene. Since these component data are in different dimensions and magnitudes, feature scaling is required to unify the features of different magnitudes and dimensions into a comparable range for easy model processing.
[0109] In some embodiments of the present disclosure, the step of performing feature scaling on the chemical composition data may include: determining the median and interquartile range of a data set of the chemical composition data; determining a data difference based on the chemical composition data and the median; and determining scaled feature data based on a ratio of the data difference to the interquartile range.
[0110] In some embodiments of the present disclosure, the step of scaling the chemical composition data may include: given a feature data to be scaled , For the specific number of data for this feature data, calculate the median of the feature data set and the interquartile range , the scaled feature data is , .
[0111] Step 122, near infrared spectrum preprocessing.
[0112] In some embodiments of the present disclosure, step 122 may include: selecting a standard normal transformation to pre-process the spectral data of the tobacco leaf samples with different flavors.
[0113] In some embodiments of the present disclosure, step 122 may include: selecting a standard normal transformation to pre-process the collected raw spectral data. Due to factors such as background noise, stray light, and manual operation, the collected spectral data inevitably contains noise signals. The above embodiments of the present disclosure select a standard normal transformation (SNV) to pre-process the collected raw spectral data before modeling.
[0114] In some embodiments of the present disclosure, the step of selecting a standard normal transformation to pre-process the collected raw spectral data may specifically include: The first wavelength points, calculate the mean of the spectrum and standard deviation ,in, It is The spectrum is The intensity value of each wavelength, It is The mean of the spectra, It is The standard deviation of the spectrum, is the number of wavelength points in the spectrum.
[0115] Step 123, near infrared spectrum embedding.
[0116] In some embodiments of the present disclosure, step 123 may include: using one-dimensional convolution to extract local features from the near-infrared spectra of the different flavor tobacco leaf samples.
[0117] In some embodiments of the present disclosure, step 123 may include: using one-dimensional convolution to extract local features of the near-infrared spectrum. In view of the characteristics of high dimensionality, redundant band information, and susceptibility to noise interference of near-infrared spectrum data, the above embodiments of the present disclosure use one-dimensional convolution to extract local features of the near-infrared spectrum. The above embodiments of the present disclosure can reduce the dimension of the data while maintaining the main features of the data through one-dimensional convolution operations, thereby reducing the computational complexity of subsequent models. At the same time, the convolution operation helps to smooth the data and filter out certain noise signals.
[0118] In some embodiments of the present disclosure, the step of using one-dimensional convolution to extract local features of the near-infrared spectrum may specifically include: given near-infrared spectrum data , The specific number of bands for this spectrum; one-dimensional convolution kernel ,in is the size of the convolution kernel. Then Strided convolution results ,in, is the step size of the convolution operation. After the convolution operation, the near-infrared spectrum embedding can be obtained. , ∈ , is the dimension of the embedding.
[0119] Step 130, constructing the tobacco aroma classification model.
[0120] In some embodiments of the present disclosure, the tobacco leaf aroma classification model may include a near-infrared spectroscopy feature extraction module, a chemical component feature extraction module, a feature fusion module and an aroma classification module.
[0121] In some embodiments of the present disclosure, the tobacco aroma classification model can be a multimodal deep neural network model based on near-infrared spectroscopy and chemical composition.
[0122] Figure 2 Schematic diagram of tobacco leaf aroma classification model in some embodiments of the present disclosure. The tobacco leaf aroma classification model is a multimodal deep neural network model based on near infrared spectroscopy and chemical composition. Figure 2 As shown, the tobacco leaf aroma classification model may include a near infrared spectrum feature extraction module, a chemical component feature extraction module, a feature fusion module and an aroma classification module.
[0123] In some embodiments of the present disclosure, Figure 2 As shown in the figure, the tobacco aroma classification model consists of three modules: feature extraction module (near infrared spectroscopy feature extraction module and chemical composition feature extraction module), feature fusion module, and aroma classification module. In the near infrared spectroscopy feature extraction module, The Transformer encoder with stacked layers in series extracts features from the near-infrared spectrum; in the chemical component feature extraction module, the features of the chemical components are extracted by MLP (multi-layer perceptron). The feature fusion module is used to splice and fuse the features extracted by the two. The aroma classification module realizes aroma classification through a fully connected neural network and softmax.
[0124] (1) Near infrared spectral feature extraction module. Figure 2 As shown, the near infrared spectroscopy feature extraction module consists of The Transformer encoder consists of stacked layers in series, and each layer mainly includes two sublayers: self-attention mechanism and feedforward neural network. Each sublayer has a residual connection, and layer normalization is used after the output of each sublayer.
[0125] Specifically, in the self-attention mechanism sublayer, for a given near-infrared spectrum embedding , add position coding After getting ;Will After three linear transformations, we get the query matrix , key matrix Sum Matrix , the formula is expressed as ; By calculating the query matrix and key matrix The correlation between the original data and the value matrix is weighted to calculate the correlation between each band data and all other band positions (i.e., attention weight). ), the formula is expressed as , is a dot product operation, which calculates the attention weight of Q on V, through The purpose of scaling is to avoid too large a dot product, because when the dot product is too large, the gradient through softmax will be very small; through the attention weight Pair Matrix Perform weighted summation to obtain the weighted output representation of each band ; Output of the self-attention mechanism , after residual connection and layer normalization, we get .
[0126] The feedforward neural network sublayer consists of two fully connected layers and a nonlinear activation function, which is expressed as ,, is the input tensor, W1, b1, W2, b2 are learnable weight matrices and bias vectors. max(0, xW1 + b1) is an activation function, usually ReLU (Rectified Linear Unit).
[0127] Each layer of Transformer encoder consists of the same operation, after After the encoder processing of the layer, the output of the near infrared spectral feature extraction module can be obtained .
[0128] (2) Chemical composition feature extraction module. Figure 2 As shown in Figure 1, the chemical composition feature extraction module consists of a multi-layer perceptron (MLP) using a ReLU activation function.
[0129] Specifically, for the chemical composition data of sample tobacco leaves , is the number of sample tobacco leaves, and the calculation process of the multilayer perceptron is expressed as , where MLPBlock is a fully connected neural network using the ReLU activation function ( ), the formula is expressed as , Dropout is the random inactivation rate of the neural network.
[0130] (3) Feature fusion module. Figure 2 As shown, the output of the feature fusion module to the near infrared spectroscopy feature extraction module And the output of the chemical composition feature extraction module , vector concatenation is used for multimodal feature fusion, and the formula is expressed as .
[0131] (4) Fragrance classification module. Figure 2 As shown in Figure 1, the aroma classification module consists of a fully connected neural network and a softmax classifier, and the formula is as shown in formula (1), where: This is the result of fragrance classification.
[0132] (1)
[0133] Step 140: Use the training data to train the tobacco flavor classification model.
[0134] In some embodiments of the present disclosure, step 140 may include: parameter setting and iterative round number and accuracy analysis.
[0135] In some embodiments of the present disclosure, step 140 may include at least one of steps 141 to 146 .
[0136] Step 141, divide the training data set into a training set, a validation set and a test set.
[0137] In some embodiments of the present disclosure, step 141 may include: randomly dividing the data set into a training set, a validation set, and a test set in a ratio of 8:1:1.
[0138] In some embodiments of the present disclosure, step 141 may further include: selecting appropriate training parameters and training the improved Transformer-based cigarette classification model using the training set.
[0139] In some embodiments of the present disclosure, in the near infrared spectrum feature extraction module, the feature embedding dimension Set to 80, the number of Transformer encoder layers The number of Transformer encoder attention heads is set to 6, and the random inactivation rate of the self-attention layer is set to 0.1. In the chemical composition feature extraction module, the number of neurons in the input layer of the fully connected neural network is 70, and the number of neurons in the two hidden layers is 128 and 64. The model learning rate is 1e -5 The number of iterations is 20 and the size of each training set data batch is 128.
[0140] Step 142, by forward propagation, the input near infrared spectrum and chemical composition data in the training set are transmitted through each module of the tobacco leaf flavor classification model to obtain the output result of the tobacco leaf flavor classification model.
[0141] In some embodiments of the present disclosure, step 142 may include: the training process uses a forward propagation algorithm to embed the input features into The output result of the tobacco flavor classification model is calculated and obtained by passing through each layer of the model. .
[0142] In some embodiments of the present disclosure, step 142 may include at least one of steps 1421 to 1424 .
[0143] Step 1421, using the near-infrared spectrum feature extraction module to extract features from the near-infrared spectra of the tobacco leaf samples with different flavors.
[0144] In some embodiments of the present disclosure, step 1421 may include: adding position encoding to a given near-infrared spectral embedded feature to obtain a first near-infrared spectral feature; performing a linear transformation on the first near-infrared spectral feature to obtain a query matrix, a key matrix, and a value matrix; calculating the correlation between each band data and all other band positions by calculating the correlation between the query matrix and the key matrix and weighting the value matrix of the original data; performing weighted summation on the value matrix through attention weights to obtain a weighted output value for each band; performing residual connection and layer normalization on the weighted output value to obtain a second near-infrared spectral feature.
[0145] In some embodiments of the present disclosure, step 1421 may include: using a self-attention mechanism sublayer to add position encoding to a given near-infrared spectrum embedding feature to obtain a first near-infrared spectrum feature; performing a linear transformation on the first near-infrared spectrum feature to obtain a query matrix, a key matrix, and a value matrix; calculating the correlation between each band data and all other band positions by calculating the correlation between the query matrix and the key matrix and weighting the value matrix of the original data; performing weighted summation on the value matrix through attention weights to obtain a weighted output value for each band; performing residual connection and layer normalization on the weighted output value to obtain a second near-infrared spectrum feature.
[0146] Step 1422: Use the chemical component feature extraction module to extract features of the chemical components of the tobacco leaf samples with different aroma types.
[0147] In some embodiments of the present disclosure, step 1422 may include: using a multi-layer perceptron to extract features of the chemical components of the tobacco leaf sample.
[0148] Step 1423, using the feature fusion module to splice and fuse the extracted near-infrared spectral features and chemical composition features.
[0149] In some embodiments of the present disclosure, step 1423 may include: performing multimodal feature fusion using vector splicing on the near-infrared spectral features output by the near-infrared spectral feature extraction module and the chemical composition features output by the chemical composition feature extraction module.
[0150] Step 1424, using the aroma classification module to classify the aroma of the tobacco leaf sample according to the spliced and fused features.
[0151] Step 143, calculating the error ( ),in, is the predicted value, i.e., the output result of the tobacco leaf aroma classification model; is the actual value.
[0152] Step 144, back-propagating from the output layer of the tobacco flavor classification model to the input layer, using back-propagation to calculate the gradient of each node; using a stochastic gradient descent optimization algorithm to update the weights and biases of the tobacco flavor classification model.
[0153] Step 145, perform iterative training, and after completing the iteration, obtain the tobacco leaf and leaf group flavor classification model .
[0154] In some embodiments of the present disclosure, the iterative training includes: repeating the process from step 142 to step 144 (forward propagation, error calculation, back propagation, parameter update) until a stopping criterion is met (for example, a preset maximum number of iterations is reached or the error reaches a preset minimum value).
[0155] Step 146, accuracy analysis.
[0156] In some embodiments of the present disclosure, step 146 may include: after the model training is completed, use the test set to test the model. The performance of the classification model is evaluated using the weighted F1 value (weight-F1), which indicates the accuracy of the classification model.
[0157] In some embodiments of the present disclosure, the weighted F1 value of the aroma classification of tobacco leaves and leaf groups reaches 0.9188.
[0158] Figure 3 Schematic diagram of some embodiments of the tobacco leaf aroma classification method disclosed in the present invention. Figure 3 The embodiment can be performed by the tobacco leaf aroma classification device disclosed in the present invention. Figure 3 As shown, Figure 3The method of the embodiment may include at least one of steps 310 to 330 .
[0159] Step 310, obtaining the near infrared spectrum and chemical composition of the tobacco leaves to be classified.
[0160] In some embodiments of the present disclosure, step 310 may include: given a cigarette sheet to be classified , near-infrared equipment was used to obtain the near-infrared spectrum of tobacco leaves and determine 70 chemical components of tobacco leaves.
[0161] Step 320, preprocessing the near infrared spectrum and chemical composition of the tobacco leaves to be classified.
[0162] In some embodiments of the present disclosure, step 320 may include: Figure 1 The pre-processing operation of step 120 in the embodiment.
[0163] In some embodiments of the present disclosure, step 320 may include: at least one of steps 121 to 123 .
[0164] In some embodiments of the present disclosure, step 320 may include: chemical component preprocessing; near infrared spectrum preprocessing; and near infrared spectrum embedding.
[0165] In some embodiments of the present disclosure, step 320 may include: performing feature scaling on the chemical composition data of the tobacco leaves to be classified to unify features of different magnitudes and dimensions into a comparable range; selecting a standard normal transformation to preprocess the spectral data of the tobacco leaves to be classified; and using one-dimensional convolution to extract local features from the near-infrared spectrum of the tobacco leaves to be classified.
[0166] In some embodiments of the present disclosure, the step of performing feature scaling on the chemical composition data of the tobacco leaves to be classified may include: determining the median and interquartile range of a data set of the chemical composition data; determining a data difference based on the chemical composition data and the median; and determining the scaled feature data based on the ratio of the data difference to the interquartile range.
[0167] Step 330, using the trained tobacco leaf aroma classification model, and classifying the tobacco leaves to be classified according to their aroma types based on the pre-processed near infrared spectra and chemical components.
[0168] In some embodiments of the present disclosure, the tobacco leaf aroma classification model is a tobacco leaf aroma classification model trained using the model training method described in any of the above embodiments.
[0169] In some embodiments of the present disclosure, step 330 may include: extracting features from the near-infrared spectral data of the tobacco leaves to be classified; extracting features from the chemical composition data of the tobacco leaves to be classified; splicing and fusing the extracted near-infrared spectral features and chemical composition features; and classifying the tobacco leaf samples by aroma type based on the spliced and fused features.
[0170] In some embodiments of the present disclosure, the step of extracting features from the near-infrared spectral data of the tobacco leaves to be classified may include: adding position coding to the near-infrared spectral data to obtain a first near-infrared spectral feature; performing a linear transformation on the first near-infrared spectral feature to obtain a query matrix, a key matrix, and a value matrix; determining an attention weight by calculating the correlation between the query matrix and the key matrix; performing a weighted summation on the value matrix using the attention weight to obtain a weighted output value; and performing a residual connection and normalization on the weighted output value to obtain a second near-infrared spectral feature.
[0171] In some embodiments of the present disclosure, the step of using the chemical component feature extraction module to extract features of the chemical components of the tobacco leaves to be classified may include: using a multi-layer perceptron to extract features of the chemical components of the tobacco leaf samples.
[0172] In some embodiments of the present disclosure, the step of using the feature fusion module to splice and fuse the extracted near-infrared spectral features and chemical composition features may include: using vector splicing to perform feature fusion on the near-infrared spectral features output by the near-infrared spectral feature extraction module and the chemical composition features output by the chemical composition feature extraction module.
[0173] In some embodiments of the present disclosure, steps 320 and 330 may include: after the near infrared spectrum and the 70 chemical components are pre-processed in step 120, they are sent to the image processing system constructed in step 130. Figure 2 The tobacco aroma classification model shown , get tobacco leaves The aroma prediction results Combined with the softmax value, a concrete representation of the aroma of tobacco leaves P can be achieved. The above embodiments of the present disclosure can more accurately describe the aroma of tobacco leaves P, providing strong support for formulators to accurately use tobacco leaves P and quantify the aroma of leaf groups.
[0174] The above embodiments of the present disclosure provide a multimodal classification method for the aroma of tobacco leaves and leaf groups based on near-infrared spectroscopy and chemical composition. The above embodiments of the present disclosure can achieve accurate classification of the aroma of tobacco leaves, and the accurate classification of the aroma of tobacco leaves can ensure the clear and stable aroma characteristics of tobacco leaves from different production areas, so that production companies can more accurately select and formulate leaf group formulas and improve the flavor stability of leaf groups. At the same time, the aroma zoning of the above embodiments of the present disclosure provides a theoretical basis for the innovation of cigarette flavors, and supports companies in developing cigarette products that meet the tastes of different consumers.
[0175] The above-mentioned embodiments of the present disclosure can assist formulators in classifying the flavors of raw tobacco leaves and leaf group formulas through deep learning algorithms, thereby effectively improving the efficiency and reliability of flavor classification, reducing human misjudgment, and ensuring the consistency of the quality of leaf group formulas.
[0176] The inventors have found through research that both near-infrared spectra and compound content are of great significance to the classification of fragrance types. Therefore, the above embodiments of the present disclosure fully consider the combined use of the two types of features, namely, near-infrared spectra and compound content.
[0177] The above-mentioned embodiments of the present disclosure propose a method for classifying aroma types based on a multimodal deep neural network model of near-infrared spectroscopy and chemical composition. The method of the above-mentioned embodiments of the present disclosure collects near-infrared spectral data of tobacco leaves and their chemical composition content, and uses deep learning technology to perform multimodal modeling of near-infrared spectra and chemical composition, thereby realizing automatic grading and classification of aroma types of tobacco leaves and leaf groups. Compared with the machine learning algorithm of the related technology that uses a single type of features, the multimodal method of the above-mentioned embodiments of the present disclosure can quantitatively evaluate tobacco leaves from multiple angles, thereby effectively improving the scientificity and accuracy of aroma classification.
[0178] Figure 4 Schematic diagram of the structure of some embodiments of the model training device disclosed in the present invention. Figure 4 As shown, the model training device of the present disclosure may include a first data acquisition unit 41, a first preprocessing unit 42, a model construction unit 43 and a model training unit 44.
[0179] The first data acquisition unit 41 is configured to acquire near infrared spectra and chemical components of tobacco leaf samples with different aroma types, and to label the tobacco leaf samples with different aroma types.
[0180] In some embodiments of the present disclosure, the flavor categories of the tobacco leaves include at least one of light sweet flavor, honey sweet flavor, mellow sweet flavor, caramel sweet flavor, caramel sweet and mellow sweet flavor, light sweet sweet flavor, honey sweet caramel flavor, woody honey sweet flavor, Zimbabwe type, Brazilian type, Argentine type, Zambian type, and American type.
[0181] The first preprocessing unit 42 is configured to preprocess the near infrared spectra and chemical components of the tobacco leaf samples with different flavors to form training data.
[0182] In some embodiments of the present disclosure, the first preprocessing unit 42 can be configured to perform feature scaling on the chemical composition data of the different flavor tobacco samples, unify features of different magnitudes and dimensions into a comparable range; select a standard normal transformation to preprocess the spectral data of the different flavor tobacco samples; and use one-dimensional convolution to extract local features of the near-infrared spectra of the different flavor tobacco samples.
[0183] In some embodiments of the present disclosure, when the first preprocessing unit 42 performs feature scaling on the chemical composition data of tobacco leaf samples of different flavors, it can be configured to determine the median and interquartile range of a data set of the chemical composition data; determine the data difference based on the chemical composition data and the median; and determine the scaled feature data based on the ratio of the data difference to the interquartile range.
[0184] The model building unit 43 is configured to build the tobacco aroma classification model.
[0185] In some embodiments of the present disclosure, the tobacco leaf aroma classification model includes a near-infrared spectral feature extraction module, a chemical component feature extraction module, a feature fusion module and an aroma classification module.
[0186] The model training unit 44 is configured to use the training data to train the tobacco leaf flavor classification model, wherein the trained tobacco leaf flavor classification model is used for tobacco leaf flavor classification.
[0187] In some embodiments of the present disclosure, the model training unit 44 can be configured to divide the training data into a training set, a validation set, and a test set; process the near-infrared spectrum and chemical composition data in the training set through various modules of the tobacco flavor classification model to obtain the output result of the tobacco flavor classification model, thereby realizing the training of the tobacco flavor classification model.
[0188] In some embodiments of the present disclosure, when the model training unit 44 processes the near-infrared spectrum and chemical composition data in the training set through the various modules of the tobacco flavor classification model to obtain the output result of the tobacco flavor classification model, it can be configured to use the near-infrared spectrum feature extraction module to extract features from the near-infrared spectrum data of the tobacco samples with different flavors; use the chemical composition feature extraction module to extract features from the chemical composition data of the tobacco samples with different flavors; use the feature fusion module to splice and fuse the extracted near-infrared spectrum features and chemical composition features; and use the flavor classification module to classify the tobacco samples by flavor based on the spliced and fused features.
[0189] In some embodiments of the present disclosure, when the model training unit 44 uses the near-infrared spectral feature extraction module to extract features from the near-infrared spectral data of the different flavor tobacco leaf samples, it can be configured to obtain a first near-infrared spectral feature after adding position coding to the near-infrared spectral data; perform a linear transformation on the first near-infrared spectral feature to obtain a query matrix, a key matrix and a value matrix; determine the attention weight by calculating the correlation between the query matrix and the key matrix; perform a weighted summation on the value matrix using the attention weight to obtain a weighted output value; perform residual connection and normalization on the weighted output value to obtain a second near-infrared spectral feature.
[0190] In some embodiments of the present disclosure, when the model training unit 44 uses the chemical component feature extraction module to extract the chemical components of the tobacco leaf samples with different aromas, it can be configured to use a multi-layer perceptron to extract the chemical components of the tobacco leaf samples.
[0191] In some embodiments of the present disclosure, when the model training unit 44 uses the feature fusion module to splice and fuse the extracted near-infrared spectral features and chemical composition features, it can be configured to use vector splicing to perform feature fusion on the near-infrared spectral features output by the near-infrared spectral feature extraction module and the chemical composition features output by the chemical composition feature extraction module.
[0192] In some embodiments of the present disclosure, the model training device of the present disclosure may be configured to implement the model training method as described in any of the above embodiments.
[0193] Figure 5 Schematic diagram of the structure of some other embodiments of the model training device disclosed in the present invention. Figure 5 As shown, the model training device disclosed herein may include a memory 51 and a processor 52 .
[0194] The memory 51 is used to store instructions, the processor 52 is coupled to the memory 51, and the processor 52 is configured to execute the model training method involved in the above embodiment based on the instructions stored in the memory.
[0195] like Figure 5 As shown, the model training device also includes a communication interface 53 for information exchange with other devices. At the same time, the model training device also includes a bus 54, through which the processor 52, the communication interface 53, and the memory 51 communicate with each other.
[0196] The memory 51 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory. The memory 51 may also be a memory array. The memory 51 may also be divided into blocks, and the blocks may be combined into virtual volumes according to certain rules.
[0197] In addition, the processor 52 may be a central processing unit (CPU), or may be an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present disclosure.
[0198] Figure 6 Schematic diagram of the structure of some embodiments of the tobacco flavor classification device disclosed in the present invention. Figure 6 As shown, the tobacco flavor classification device disclosed in the present invention may include a second data acquisition unit 61, a second preprocessing unit 62 and a flavor classification unit 63.
[0199] The second data acquisition unit 61 is configured to acquire the near infrared spectrum and chemical composition of the tobacco leaves to be classified.
[0200] The second preprocessing unit 62 is configured to preprocess the near infrared spectrum and chemical composition of the tobacco leaves to be classified.
[0201] In some embodiments of the present disclosure, the second preprocessing unit 62 can be configured to perform feature scaling on the chemical composition data of the tobacco leaves to be classified, unifying features of different magnitudes and dimensions into a comparable range; selecting a standard normal transformation to preprocess the spectral data of the tobacco leaves to be classified; and using one-dimensional convolution to extract local features from the near-infrared spectra of the tobacco leaves to be classified.
[0202] The aroma classification unit 63 is configured to use the trained tobacco aroma classification model to classify the tobacco to be classified according to the pre-processed near infrared spectrum and chemical composition.
[0203] In some embodiments of the present disclosure, the aroma classification unit 63 can be configured to perform feature extraction on the near-infrared spectral data of the tobacco leaves to be classified; perform feature extraction on the chemical composition data of the tobacco leaves to be classified; perform splicing and fusion of the extracted near-infrared spectral features and chemical composition features; and perform aroma classification of the tobacco leaf samples based on the spliced and fused features.
[0204] In some embodiments of the present disclosure, when the aroma classification unit 63 performs feature extraction on the near-infrared spectral data of the tobacco leaves to be classified, it can be configured to obtain a first near-infrared spectral feature after adding position coding to the near-infrared spectral data; perform linear transformation on the first near-infrared spectral feature to obtain a query matrix, a key matrix and a value matrix; determine the attention weight by calculating the correlation between the query matrix and the key matrix; perform weighted summation on the value matrix using the attention weight to obtain a weighted output value; perform residual connection and normalization on the weighted output value to obtain a second near-infrared spectral feature.
[0205] In some embodiments of the present disclosure, the tobacco aroma classification device of the present disclosure may be configured to implement the tobacco aroma classification method as described in any of the above embodiments.
[0206] The above-mentioned embodiments of the present disclosure provide a flavor classification device based on a multimodal deep neural network model of near-infrared spectroscopy and chemical composition, thereby the above-mentioned embodiments of the present disclosure can improve the speed and accuracy of flavor classification of tobacco leaves and leaf groups.
[0207] Figure 7 Schematic diagram of the structure of some other embodiments of the tobacco flavor classification device disclosed in the present invention. Figure 7 As shown, the structure and function of the tobacco flavor classification device disclosed in the present invention are similar to Figure 5 The structures and functions of the model training devices in the embodiments are the same or similar. Figure 7 Tobacco aroma classification device of the embodiment and Figure 5 The model training device of the embodiment differs only in that: Figure 7 The processor 52 of the embodiment is configured to execute the tobacco aroma classification method involved in the above embodiment based on the instructions stored in the memory.
[0208] Figure 8 Schematic diagram of the structure of some embodiments of the tobacco aroma classification system disclosed in the present invention. Figure 8 As shown, the tobacco leaf aroma classification system disclosed in the present invention may include a near infrared device 81, a chemical composition determination device 82, a model training device 83 and a tobacco leaf aroma classification device 84.
[0209] The near infrared device 81 can be configured to obtain the near infrared spectrum of the tobacco leaves.
[0210] The chemical composition measuring device 82 may be configured to measure the chemical composition of tobacco leaves.
[0211] In some embodiments of the present disclosure, the chemical component determination device 82 may be configured to determine 70 chemical components of tobacco leaves.
[0212] In some embodiments of the present disclosure, the model training device 83 may be a model training device as described in any of the above embodiments.
[0213] In some embodiments of the present disclosure, the tobacco flavor classification device 84 may be a tobacco flavor classification device as described in any of the above embodiments.
[0214] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the method for classifying tobacco leaf aroma types as described in any of the above embodiments is implemented.
[0215] According to another aspect of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the tobacco leaf aroma classification method as described in any of the above embodiments is implemented.
[0216] The computer-readable storage medium of the present disclosure may be implemented as a non-transitory computer-readable storage medium.
[0217] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, devices, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable non-transient storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0218] The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0219] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0220] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0221] The model training device, first data acquisition unit, first preprocessing unit, model building unit, model training unit, tobacco flavor classification device, second data acquisition unit, second preprocessing unit, flavor classification unit, near-infrared spectral feature extraction module, chemical component feature extraction module, feature fusion module, flavor classification module, Transformer encoder, self-attention mechanism sublayer, feedforward neural network sublayer, multilayer perceptron, fully connected neural network and softmax classifier described above can be implemented as a general-purpose processor, programmable logic controller, digital signal processor (DSP), application-specific integrated circuit (ASIC), field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component or any appropriate combination thereof for performing the functions described in the present disclosure.
[0222] Those skilled in the art will appreciate that all or part of the steps of the above-described embodiment method of the present disclosure may be accomplished by hardware, and the hardware may be implemented as a general-purpose processor, a programmable logic controller, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, or any appropriate combination thereof for executing the method described in the present disclosure.
[0223] So far, the present disclosure has been described in detail. In order to avoid obscuring the concept of the present disclosure, some details known in the art are not described. Based on the above description, those skilled in the art can fully understand how to implement the technical solution disclosed here.
[0224] A person of ordinary skill in the art will appreciate that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by instructing the relevant hardware through a program, and the program may be stored in a non-transitory computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk, or an optical disk, etc.
[0225] The description of the present disclosure is given for the purpose of illustration and description, and is not intended to be exhaustive or to limit the present disclosure to the disclosed form. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments are selected and described in order to better illustrate the principles and practical applications of the present disclosure, and to enable those of ordinary skill in the art to understand the present disclosure and thereby design various embodiments with various modifications suitable for specific uses.
Claims
1. A model training method, comprising: Obtaining near infrared spectra and chemical compositions of tobacco leaf samples with different aroma types, and labeling the tobacco leaf samples with different aroma types; Preprocessing the near infrared spectra and chemical components of the tobacco leaf samples with different aroma types to form training data; Constructing the tobacco leaf aroma classification model, wherein the tobacco leaf aroma classification model is a tobacco leaf aroma classification model based on near infrared spectroscopy and chemical composition; The training data is used to train the tobacco leaf aroma classification model, wherein the trained tobacco leaf aroma classification model is used for tobacco leaf aroma classification.
2. The model training method according to claim 1, wherein: The tobacco leaf aroma classification model comprises a near infrared spectrum feature extraction module, a chemical component feature extraction module, a feature fusion module and an aroma classification module, and the training of the tobacco leaf aroma classification model using the training data comprises: Divide the training data into training set, validation set and test set; The near infrared spectrum and chemical composition data in the training set are processed through various modules of the tobacco leaf aroma classification model to obtain the output result of the tobacco leaf aroma classification model, thereby realizing the training of the tobacco leaf aroma classification model.
3. The model training method according to claim 2, wherein: The near infrared spectrum and chemical composition data in the training set are processed through various modules of the tobacco leaf aroma classification model to obtain the output result of the tobacco leaf aroma classification model, including: The near infrared spectrum feature extraction module is used to extract features from the near infrared spectrum data of the tobacco leaf samples with different aroma types; Using the chemical component feature extraction module to perform feature extraction on the chemical component data of the tobacco leaf samples with different aroma types; The feature fusion module is used to splice and fuse the extracted near-infrared spectrum features and chemical composition features; The aroma classification module is used to classify the aroma of tobacco leaf samples according to the spliced and fused features.
4. The model training method according to claim 3, wherein: The method of using the near infrared spectrum feature extraction module to extract features from the near infrared spectrum data of the tobacco leaf samples with different aroma types includes: After adding position coding to the near infrared spectrum data, a first near infrared spectrum feature is obtained; Performing a linear transformation on the first near infrared spectral feature to obtain a query matrix, a key matrix and a value matrix; Determine the attention weight by calculating the correlation between the query matrix and the key matrix; Performing weighted summation on the value matrix using the attention weights to obtain a weighted output value; The weighted output values are residually connected and normalized to obtain a second near-infrared spectrum feature.
5. The model training method according to claim 3 or 4, wherein: The step of extracting the chemical components of the tobacco leaf samples with different aroma types by using the chemical component feature extraction module comprises: Multilayer perceptron was used to extract features of the chemical components of tobacco leaf samples.
6. The model training method according to claim 3 or 4, wherein: The step of using the feature fusion module to splice and fuse the extracted near infrared spectral features and chemical composition features includes: Vector splicing is used to perform feature fusion on the near infrared spectral features output by the near infrared spectral feature extraction module and the chemical composition features output by the chemical composition feature extraction module.
7. The model training method according to any one of claims 1 to 4, wherein: The preprocessing of the near infrared spectra and chemical components of tobacco leaf samples with different aroma types includes: Performing feature scaling on the chemical component data of the tobacco leaf samples with different flavors to unify the features of different magnitudes and dimensions into a comparable range; Selecting standard normal transformation to preprocess the spectral data of the tobacco leaf samples with different aromas; One-dimensional convolution is used to extract local features from the near-infrared spectra of the different flavor tobacco samples.
8. The model training method according to claim 7, wherein: The feature scaling of the chemical composition data of tobacco leaf samples with different aroma types includes: determining a median and an interquartile range for a data set of the chemical composition data; Determining a data difference according to the chemical composition data and the median; The scaled feature data is determined according to the ratio of the data difference value to the interquartile range.
9. The model training method according to any one of claims 1 to 4, wherein: The flavor categories of the tobacco leaves include at least one of fresh sweet flavor, honey sweet flavor, mellow sweet flavor, burnt sweet flavor, burnt sweet and mellow sweet flavor, fresh sweet sweet flavor, honey sweet burnt flavor, woody honey sweet flavor, Zimbabwe type, Brazilian type, Argentine type, Zambian type, and American type.
10. A method for classifying tobacco leaf aroma types, comprising: Obtain near infrared spectra and chemical compositions of tobacco leaves to be classified; Preprocessing the near infrared spectrum and chemical composition of the tobacco leaves to be classified; The trained tobacco leaf aroma classification model is used to classify the tobacco leaves to be classified according to the pre-processed near infrared spectrum and chemical composition.
11. The tobacco leaf aroma classification method according to claim 10, wherein: The preprocessing of the near infrared spectrum and chemical composition of the tobacco leaves to be classified comprises: Performing feature scaling on the chemical component data of the tobacco leaves to be classified, unifying features of different magnitudes and dimensions into a comparable range; Selecting standard normal transformation to preprocess the spectral data of the tobacco leaves to be classified; One-dimensional convolution is used to extract local features from the near-infrared spectrum of the tobacco leaves to be classified.
12. The tobacco leaf aroma classification method according to claim 10 or 11, wherein: The aroma classification of the tobacco leaves to be classified comprises: Extracting features from the near-infrared spectrum data of the tobacco leaves to be classified; Performing feature extraction on the chemical composition data of the tobacco leaves to be classified; The extracted near-infrared spectral features and chemical composition features are spliced and fused; The aroma types of tobacco leaf samples are classified according to the characteristics after splicing and fusion.
13. The tobacco leaf aroma classification method according to claim 10 or 11, wherein: The feature extraction of the near infrared spectrum data of the tobacco leaves to be classified comprises: After adding position coding to the near infrared spectrum data, a first near infrared spectrum feature is obtained; Performing a linear transformation on the first near infrared spectral feature to obtain a query matrix, a key matrix and a value matrix; Determine the attention weight by calculating the correlation between the query matrix and the key matrix; Performing weighted summation on the value matrix using the attention weights to obtain a weighted output value; The weighted output values are residually connected and normalized to obtain a second near-infrared spectrum feature.
14. A model training device, comprising: A first data acquisition unit is configured to acquire near infrared spectra and chemical components of tobacco leaf samples with different aroma types, and to label the tobacco leaf samples with different aroma types with aroma types; A first preprocessing unit is configured to preprocess the near infrared spectra and chemical components of the tobacco leaf samples with different flavors to form training data; A model building unit, configured to build the tobacco aroma classification model; The model training unit is configured to use the training data to train the tobacco leaf flavor classification model, wherein the trained tobacco leaf flavor classification model is used for tobacco leaf flavor classification.
15. A model training device, comprising: a memory configured to store instructions; and The processor is configured to execute the instructions so that the model training device implements the model training method as described in any one of claims 1-9.
16. A tobacco leaf aroma classification device, comprising: A second data acquisition unit is configured to acquire the near infrared spectrum and chemical composition of the tobacco leaves to be classified; A second preprocessing unit is configured to preprocess the near infrared spectrum and chemical composition of the tobacco leaves to be classified; The aroma classification unit is configured to use a trained tobacco aroma classification model to classify the aroma of the tobacco to be classified according to the pre-processed near infrared spectrum and chemical composition.
17. A tobacco aroma classification device, comprising: a memory configured to store instructions; and The processor is configured to execute the instructions so that the tobacco aroma classification device implements the tobacco aroma classification method as described in any one of claims 10-13.
18. A tobacco leaf aroma classification system, comprising the model training device according to claim 14 or 15, and the tobacco leaf aroma classification device according to claim 16 or 17.
19. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer instructions, which, when executed by a processor, implement the model training method according to any one of claims 1 to 9 and / or the tobacco leaf aroma classification method according to any one of claims 10 to 13.
20. A computer program product comprising a computer program, wherein: When the computer program is executed by a processor, the model training method according to any one of claims 1 to 9 and / or the tobacco flavor classification method according to any one of claims 10 to 13 are implemented.