Tobacco leaf quality prediction method based on PSO improved Bi-LSTM model and multi-head self-attention
By using the PSO-optimized Bi-LSTM model and the multi-head self-attention mechanism in the tobacco leaf quality prediction, the problem of low model training efficiency and difficulty in capturing complex features in the existing technology is solved, and higher prediction accuracy and model stability are achieved.
Patent Information
- Application Number
- CN202510264556.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-20
AI Technical Summary
In the prediction of tobacco leaf quality, the problem of low model training efficiency, high computing resource consumption, easy overfitting, and difficulty in effectively capturing the subtle relationship between complex characteristics and influencing factors in tobacco leaf quality data.
A two-way long and short time-calling network (Bi-LSTM) model based on particle swarm optimization (PSO) is adopted, and combined with the multi-head self-attention mechanism, the model parameters are optimized and complex timing features are extracted to improve the accuracy and reliability of tobacco leaf quality prediction.
The accuracy of tobacco leaf quality prediction and the stability of the model are significantly improved, the response ability to complex leaf-beating and re-baking process parameters are enhanced, and the generalization ability and robustness of the model are improved.
Smart Images

Figure CN120180904A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of tobacco leaf quality prediction, and particularly to a tobacco leaf quality prediction method based on an improved Bi-LSTM model with PSO and multi-head self-attention. Background Art
[0002] As the core raw material of the tobacco industry, the quality of tobacco leaves plays a crucial role in the quality of tobacco products. In the continuous development process of the tobacco industry, the influencing factors of tobacco leaf quality are complex and diverse, covering multiple links such as planting environment, variety characteristics, cultivation management measures, and processing technology.
[0003] In the process of tobacco leaf processing, the threshing and redrying process is particularly crucial, which involves many key time-series parameters, such as temperature, humidity, time, etc. The dynamic changes of these parameters during the processing have a profound impact on the final tobacco leaf quality. At the same time, the image features of tobacco leaves after threshing, such as leaf texture, shape, color, etc., can also reflect the tobacco leaf quality to a certain extent. However, traditional tobacco leaf quality evaluation and prediction methods mainly rely on empirical judgment and simple statistical analysis methods. Empirical judgment is highly subjective and has great limitations, making it difficult to cope with complex actual situations. Simple statistical analysis methods are unable to effectively mine deep-seated laws and correlation relationships when dealing with tobacco leaf quality data with significant non-linearity, time-series, and complexity characteristics, resulting in low prediction accuracy and reliability.
[0004] In recent years, deep learning technology has shown great potential in data modeling and prediction tasks in various fields. However, existing deep learning-based tobacco leaf quality prediction methods still have many problems. For example, the model training efficiency is low, consuming a large amount of computing resources and time costs, and it is prone to an increased risk of overfitting due to too long training time. At the same time, the ability to extract complex data features in tobacco leaf quality data is limited, making it difficult to comprehensively capture and utilize the subtle relationships between various influencing factors, thereby affecting the performance of the prediction model.
[0005] To overcome these limitations, there is an urgent need for an innovative tobacco leaf quality prediction method that can fully integrate multi-source data throughout the tobacco leaf production process, deeply mine data features, improve model training efficiency and prediction accuracy, and thus provide strong technical support for the refined management of tobacco leaf production, the improvement of tobacco product quality, and the sustainable development of the industry.
[0006] The present invention aims to solve the above problems existing in the prior art and proposes a tobacco leaf quality prediction method based on an improved Bi-LSTM model with PSO and multi-head self-attention mechanism. Summary of the Invention
[0007] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is how to improve the accuracy and reliability of tobacco leaf quality prediction.
[0008] To achieve the above object, the present invention provides a tobacco leaf quality prediction method based on an improved Bi-LSTM model with PSO and a multi-head self-attention mechanism. By introducing the particle swarm optimization (PSO) algorithm to optimize the parameters of the Bi-LSTM model, this method can effectively improve the training efficiency and prediction accuracy of the model. At the same time, the multi-head self-attention mechanism enables the model to adaptively focus on the important features of each time step in the input data, thereby enhancing the learning ability for complex time series patterns in tobacco leaf quality prediction.
[0009] Further, the method includes the following steps:
[0010] Step 1: Real-time collect the key time series parameter data, corresponding tobacco leaf re-drying quality index data, and tobacco leaf image data after threshing during the tobacco leaf threshing and re-drying process, integrate the data to establish a database, and preprocess the data in the database to obtain a normalized sample set;
[0011] Step 2: Construct a quality prediction model; first, complete the construction of the Bi-LSTM network architecture. To make up for the complexity of tobacco leaf quality data and the deficiencies of traditional models, then incorporate the multi-head attention mechanism, and use the output of the Bi-LSTM network as the input of the multi-head self-attention mechanism to enhance the Bi-LSTM's ability to capture the correlation relationship of data features and provide richer feature information for subsequent tobacco leaf quality prediction;
[0012] When optimizing the Bi-LSTM network in Step 3, specifically use the particle swarm optimization algorithm to perform hyperparameter tuning, realize the dynamic optimization of hyperparameters adaptively, and determine the optimal parameter combination to improve the accuracy and stability of the model;
[0013] Step 4: Use the sample set to train the model to obtain the best model performance, input the real-time collected threshing and re-drying process parameter data into the trained model, and output the quality prediction result of the tobacco leaf.
[0014] Further, in Step 1, collect tobacco leaf images from multiple perspectives and regions after the threshing process, including the front and back of the tobacco leaf and images of different parts, to comprehensively reflect the appearance characteristics of the tobacco leaf; real-time collect process parameter data such as temperature, humidity, and wind speed during the re-drying process to ensure the accuracy and continuity of the data. At the same time, record time series information such as the processing time of each stage during the re-drying process for subsequent analysis of the time series characteristics of the data; the matrix X of the database can be expressed as:
[0015]
[0016] Among them, m is the number of collected samples, which represents the data of each tobacco leaf threshing and redrying process; n is the number of process parameters related to tobacco leaf quality, including key timing parameters such as temperature, humidity, time, wind speed, and pressure during the tobacco leaf threshing and redrying process; x ij is the value of the j-th process parameter in the i-th sample (i = 1, 2,..., m; j = 1, 2,..., n); k is the number of features extracted from the tobacco leaf image after threshing; i pq represents the q-th image feature corresponding to the p-th sample; y i represents the tobacco leaf quality index value corresponding to the i-th sample, such as the large and medium slice rate, the stem content in the leaf, nicotine content, sugar content, etc.
[0017] Furthermore, in step 1, when preprocessing the data in the database to obtain a sample set, the preprocessing of data cleaning and normalization includes:
[0018] Clean the collected data, identify and correct outliers by setting a threshold range. The method is to calculate the mean μ and standard deviation σ of each process parameter, and regard the data outside the μ ± kσ (k takes 2 or 3) threshold range as outliers; for the outliers, they can be corrected or deleted according to the data distribution and actual process knowledge; for missing values, interpolation methods based on time series can be used for filling to ensure data continuity;
[0019] Normalize the cleaned data to ensure that the numerical scales of the data are consistent and prevent the data after weight operation from causing bias in training; normalize each process parameter and tobacco leaf quality parameter to the [0, 1] interval to eliminate the differences in dimension and numerical range between different parameters, which is convenient for subsequent model training and analysis; use a logistic regression model with L1 regularization to perform feature screening on the normalized database, and use the screened features as the sample set;
[0020] In this method, the logistic regression method with L1 regularization applicable to high-dimensional data sets is selected. It generates a sparse weight matrix by sparsifying the coefficients of the model, so that the model focuses on the features with non-zero coefficient values, and thus achieves the purpose of feature selection. It is constructed using the Kears module in Python; its working principle is as follows:
[0021]
[0022] Among them, |w i | represents the L1 norm of the weight vector of the model, that is, the absolute value of each weight, w i represents the weight vector, and n is the number of weights;
[0023] Integrate the data to establish a database, preprocess the data in the database, and obtain a normalized sample set.
[0024] Furthermore, step 2 also includes the following steps:
[0025] Step 2.1: Determine the network architecture parameters, determine the structures of the input layer, hidden layer, and output layer of the Bi-LSTM network. The number of input layer nodes is determined according to the number of data features related to tobacco leaf quality collected;
[0026] Set the hidden layer structure of the Bi-LSTM, including the number of hidden units in the forward LSTM and backward LSTM. The selection of the number of hidden units affects the model's learning ability for data features, which is determined through experiments and tuning. First, set it to 64, and then adjust it according to the model's performance on the validation set;
[0027] Configure the output layer. The number of output layer nodes depends on the number of tobacco leaf quality indicators to be predicted;
[0028] Step 2.2: Construct the LSTM unit and define two parameters, the weight matrix and bias vector, in the LSTM unit;
[0029] For the forward LSTM, at each time step t, the dimension of the received input data is the same as the number of input layer nodes, and the dimension of the hidden state at the previous time step is the number of hidden units;
[0030] The forget gate determines which part of the previous cell state C t+1 needs to be retained or forgotten. The calculation formula of the forget gate is as follows:
[0031] z f =σ(W f ·h t 1 ,x t )+b f )
[0032] where W f is the weight matrix of the forget gate, whose dimension is (number of hidden units + number of input layer nodes) × number of hidden units, and b f is the bias vector of the forget gate, whose dimension is the number of hidden units, and σ is the Sigmoid activation function;
[0033] The calculation formula of the input gate is as follows:
[0034] z i =σ(W i ·h t 1 ,x t )+b i )
[0035] where W i and b i are respectively the weight matrix and bias vector of the input gate, and their dimensions are the same as those of the forget gate;
[0036] Simultaneously calculate the candidate value C for updating the cell state t =tanh(W C ·[h t-1 ,x t +b C ), where W C and b C are the corresponding weight matrix and bias vector;
[0037] The formula for updating the cell state is as follows:
[0038]
[0039] The formula for the output gate is as follows:
[0040] z o =σ(W o ·h t 1 ,x t +b o )
[0041] where W o and b o are the weight matrix and bias vector of the output gate, and finally the hidden state h at the current time step is obtained t =z o ·tanh(C t );
[0042] The calculation process of the backward LSTM is similar to that of the forward one, but the data processing direction is opposite, starting from the end of the sequence and calculating towards the start direction; the backward LSTM receives the input data x t and the hidden state h of the previous time step t+1 , and randomly initializes the vector, and calculates according to the above calculation steps of the forget gate, input gate, cell state update and output gate to obtain the hidden state of the backward LSTM at each time step;
[0043] Step 2.3, The forward LSTM processes data from the start of the sequence, and the backward LSTM processes from the end. The weighted combination of their outputs gives the hidden state output of the Bi-LSTM.
[0044] Furthermore, in the said Step 2, on the basis of integrating the multi-head attention mechanism into the constructed Bi-LSTM network architecture, taking the multiple outputs of the model as the input of the multi-head self-attention, the implementation steps include:
[0045] Step S1, Perform feature encoding on it. The output feature vectors of the Bi-LSTM are and Map it to different subspaces through linear transformation to obtain the query matrix, key matrix, and value matrix. The calculation formula for the linear transformation is as follows:
[0046] q i = W q f i
[0047] k i = W k f i
[0048] v i = W v f i
[0049] Among them, W q 、W k and W v are learnable weight matrices. From this, the query matrix Q = [q 1 ,…,q n , key matrix K = [k 1 ,…,k n , and value matrix V = [v 1 ,…,v n can be obtained;
[0050] Step S2: To obtain the attention distribution, calculate the similarity between the query vector q and the key vector k. The calculation expression is as follows:
[0051]
[0052] By normalizing the similarity scores, the weight distribution of each query vector relative to all key vectors is obtained. The weight reflects the degree of attention between different elements. Multiply the weight distribution by the corresponding numerical vector and perform a weighted sum on the results to obtain the output result
[0053]
[0054] The calculation process of the self-attention mechanism can be expressed as
[0055]
[0056] In the formula, d k represents the dimension of the key matrix K. To further enhance the expression ability of the model, by adding multiple independent self-attention heads from different perspectives, the temporal features are mapped to different sub-representation spaces to guide the model to learn the correlation information of different sub-representation spaces to obtain richer representations. The calculation formula for multi-head self-attention:
[0057] MHSA Concatenate(SAM1,SAM2,…,SAM n )
[0058] Step S3: After normalization, multiply by the value vector and sum with weights to obtain the output of each attention head. Finally, concatenate the outputs of multiple heads to get the result of the multi-head self-attention mechanism, enhancing the ability to capture the correlation relationships of data features.
[0059] Furthermore, in step 2, a fully connected layer is connected after the output layer of the multi-head self-attention mechanism. The role of the fully connected layer is to further integrate and transform the features output by the multi-head self-attention mechanism and map them to a space more suitable for the prediction task.
[0060] A SoftMax layer is connected after the fully connected layer. The fully connected layer is used to integrate features, and the SoftMax layer maps the output to the prediction of tobacco leaf quality categories to construct a tobacco leaf quality prediction model.
[0061] Furthermore, step 3 further includes the following steps:
[0062] Step 3.1: Initialize the particle swarm and set parameters, determine the particle swarm size. Each particle represents a set of parameters to be optimized in the Bi-LSTM model, including randomly generating the position and velocity of the particle. Set parameters such as the particle swarm size, maximum number of iterations, inertia weight, cognitive factor, and social factor.
[0063] Step 3.2: Calculate the fitness value of each particle. Assign the parameters represented by the particle to the Bi-LSTM model and use the training set data to train the model. During the training process, calculate the loss of the model on the training set according to the set loss function. After training, evaluate the model performance using the validation set data, and use this performance metric as the fitness value of the particle.
[0064] Step 3.3: Evaluate the model performance, and comprehensively evaluate the model using three indicators: root mean square error RMSE, mean absolute error MAE, and coefficient of determination R 2 . RMSE is used to measure the deviation degree between the predicted value and the actual value, MAE reflects the average size of the prediction error, and R 2 evaluates the goodness of fit of the model to the data. Through the analysis of these indicators, intuitively display the prediction accuracy and reliability of the model, providing a basis for the improvement and application of the model.
[0065] Furthermore, in step 3.2, according to the current position, velocity, individual best position, and global best position of the particle, use the following formula to update the velocity and position:
[0066] Velocity update formula:
[0067] Position update formula:
[0068] Where and are the velocity and position of the j-th particle in the d-th dimension in the i-th iteration respectively, ω is the inertia weight, c1 and c2 are the cognitive factor and social factor, r1 and r2 are random numbers between [0, 1], is the individual optimal position of the i-th particle, is the global optimal position.
[0069] Furthermore, in step 3.2, the particle updates its velocity and position according to the following method:
[0070] Velocity update: The velocity is adjusted by comprehensively considering the particle's own historical optimal position and the global optimal position. The inertial part keeps the particle with a certain motion inertia; the cognitive part makes the particle approach its own historical best position; the social part makes the particle approach the historical best position of the whole group. In this way, the particle continuously adjusts its direction and velocity in the search space to find a better parameter combination.
[0071] Position update: The position of the particle is adjusted according to the updated velocity, which is also the value of the parameter.
[0072] By continuously iterating the above process, the particle swarm can search in the hyperparameter space, avoiding the blindness and inefficiency of manual tuning or simple random search. Finally, a set of optimal hyperparameter combinations is found, thereby improving the performance of the Bi-LSTM model, such as improving the training efficiency and avoiding falling into local optimal solutions.
[0073] The mean square error loss function expression set for calculating the particle fitness value is as follows:
[0074]
[0075] where m is the number of training samples, y j is the actual tobacco leaf quality value, is the model prediction value.
[0076] Iterative optimization and termination condition judgment, repeat the above steps of calculating the fitness value, updating the velocity and position until the maximum number of iterations is reached, and output the global optimal position as the optimized Bi-LSTM model parameters. In a preferred embodiment of the present invention,
[0077] By combining PSO to improve Bi-LSTM and the multi-head self-attention mechanism, the present invention can accurately predict the key indicators of tobacco leaf quality and improve the response ability to complex tobacco leaf threshing and redrying process parameters. This method not only optimizes the accuracy of tobacco leaf quality prediction, but also enhances the robustness of the model in complex environments and improves its generalization ability.
[0078] The following will further illustrate the concept, specific structure and technical effects of the present invention in conjunction with the accompanying drawings to fully understand the purpose, features and effects of the present invention. Brief Description of the Drawings
[0079] Figure 1 is a flowchart of a preferred embodiment of the present invention. Detailed Embodiments
[0080] The following introduces multiple preferred embodiments of the present invention with reference to the accompanying drawings of the specification to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.
[0081] In the accompanying drawings, components with the same structure are denoted by the same numeral labels, and components with similar structures or functions are denoted by similar numeral labels. The size and thickness of each component shown in the drawings are arbitrarily shown, and the present invention does not limit the size and thickness of each component. In order to make the illustration clearer, the thickness of some components in the drawings is appropriately exaggerated.
[0082] As Figure 1 shown, the method of this embodiment includes the following steps:
[0083] Step 1: Real-time collect the key time-series parameter data, the corresponding tobacco leaf redrying quality index data, and the tobacco leaf image data after threshing during the tobacco leaf threshing and redrying process, establish a database, integrate the data to establish a database, and preprocess the data in the database to obtain a normalized sample set;
[0084] Step 2: Construct a quality prediction model. First, complete the construction of the Bi-LSTM network architecture. In order to make up for the complexity of tobacco leaf quality data and the deficiencies of traditional models, then integrate the multi-head attention mechanism to enhance the ability to capture feature correlation relationships. Take the output of the Bi-LSTM network as the input of the multi-head self-attention mechanism to enhance the ability of Bi-LSTM to capture feature correlation relationships in the data and provide richer feature information for subsequent tobacco leaf quality prediction;
[0085] When optimizing the Bi-LSTM network, specifically use the particle swarm optimization algorithm to perform hyperparameter tuning, realize the dynamic optimization of hyperparameters adaptively, and determine the optimal parameter combination to improve the accuracy and stability of the model;
[0086] Through the above steps, the model is trained using the said sample set to obtain the best model performance, and the data of the threshing and redrying process parameters collected in real time are input into the trained model to output the quality prediction result of the tobacco leaves.
[0087] In step 1 of this embodiment, the data collection of the tobacco leaf images after threshing, the parameters of the redrying process, and the tobacco leaf quality indicators is completed. After the threshing process is completed, tobacco leaf images from multiple perspectives and regions are collected, including the front and back of the tobacco leaves and images of different parts, comprehensively reflecting the appearance characteristics of the tobacco leaves; the process parameter data such as temperature, humidity, and wind speed during the redrying process are collected in real time to ensure the accuracy and continuity of the data. At the same time, the time series information such as the processing time at each stage during the redrying process is recorded for subsequent analysis of the time series characteristics of the data. The database matrix X can be expressed as:
[0088]
[0089] where m is the number of collected samples, which represents the data of each tobacco leaf threshing and redrying process; n is the number of process parameters related to the tobacco leaf quality, such as key time series parameters such as temperature, humidity, time, wind speed, and pressure during the threshing and redrying process; x ij is the j-th process parameter value in the i-th sample (i = 1, 2,..., m; j = 1, 2,..., n); k is the number of features extracted from the tobacco leaf images after threshing; i pq represents the q-th image feature corresponding to the p-th sample; y i represents the tobacco leaf quality indicator value corresponding to the i-th sample, such as the large and medium leaf rate, the stem content in the leaf, nicotine content, sugar content, etc.
[0090] In step 1, when preprocessing the data in the database to obtain the sample set, the preprocessing of data cleaning and normalization includes:
[0091] Clean the collected data, identify and correct outliers by setting a threshold range. The method is to calculate the mean μ and standard deviation σ of each process parameter, and regard the data outside the threshold range of μ ± kσ (k takes 2 or 3) as outliers.
[0092] For outliers, they can be corrected or deleted according to the data distribution and actual process knowledge; for missing values, interpolation methods based on time series can be used for filling to ensure the continuity of the data.
[0093] The cleaned data is normalized to ensure that the numerical scales of the data are consistent and prevent the data after weight operation from causing bias in training.
[0094] Normalize each process parameter and tobacco leaf quality parameter to the interval [0, 1] to eliminate the differences in dimension and numerical range between different parameters, facilitating subsequent model training and analysis.
[0095] Use a logistic regression model with L1 regularization to perform feature screening on the normalized database, and use the screened features as the sample set.
[0096] Considering that initial screening is required before high-dimensional features with complex data types are input into the model, in this embodiment, the L1-regularized logistic regression method applicable to high-dimensional data sets is selected. It generates a sparse weight matrix by sparsifying the coefficients of the model, so that the model focuses on the features whose coefficients are non-zero values, thereby achieving the purpose of feature selection.
[0097] Build using the Kears module in Python. Its working principle is as follows:
[0098]
[0099] Among them, |w i | represents the L1 norm of the weight vector of the model, that is, the absolute value of each weight, w i represents the weight vector, and n is the number of weights.
[0100] Integrate data to establish a database, preprocess the data in the database, and obtain a normalized sample set.
[0101] Step 2 of this embodiment specifically includes, for constructing a Bi-LSTM model:
[0102] First, determine the network architecture parameters:
[0103] Determine the structures of the input layer, hidden layer, and output layer of the Bi-LSTM network. The number of input layer nodes is determined according to the number of data features related to tobacco leaf quality collected;
[0104] Set the hidden layer structure of the Bi-LSTM, including the number of hidden units of the forward LSTM and the backward LSTM; the selection of the number of hidden units will affect the model's learning ability for data features, and it is generally determined through experiments and tuning. First, set it to 64, and then adjust it according to the performance of the model on the validation set;
[0105] Configure the output layer. The number of output layer nodes depends on the number of tobacco leaf quality indicators to be predicted;
[0106] Secondly, construct the LSTM unit, define parameters such as the weight matrix and bias vector in the LSTM unit, and these parameters will be adjusted in the subsequent model training and optimization process;
[0107] For the forward LSTM, at each time step t, it receives the input data (with the same dimension as the number of input layer nodes) and the hidden state of the previous time step (with the dimension of the number of hidden units).
[0108] The forget gate determines which part of the previous cell state C t+1 needs to be retained or forgotten. The calculation formula of the forget gate is as follows:
[0109] z f = σ(W f ·h t 1 , x t ) + b f )
[0110] where W f is the weight matrix of the forget gate, with the dimension [(number of hidden units + number of input layer nodes) × number of hidden units], b f is the bias vector of the forget gate, with the dimension of the number of hidden units, and σ is the Sigmoid activation function;
[0111] The calculation formula of the input gate is as follows:
[0112] z i = σ(W i ·h t 1 , x t ) + b i )
[0113] where W i and b i are the weight matrix and bias vector of the input gate respectively, with the same dimension as the forget gate;
[0114] At the same time, calculate the candidate value C for updating the cell state t = tanh(W C ·[h t-1 , x t ) + b C ), W C and b C are the corresponding weight matrix and bias vector;
[0115] The calculation formula for updating the cell state is as follows:
[0116]
[0117] The calculation formula of the output gate is as follows:
[0118] z o = σ(W o ·h t 1 , x t ) + b o )
[0119] where W o and b o are the weight matrix and bias vector of the output gate, and finally the hidden state h at the current time step is obtained t = z o ·tanh(C t );
[0120] The calculation process of the backward LSTM is similar to that of the forward one, but the data processing direction is opposite, starting from the end of the sequence and calculating in the starting direction. The backward LSTM receives the input data x t and the hidden state h of the previous time step t+1 , and randomly initializes the vector, and calculates according to the above calculation steps of the forget gate, input gate, cell state update and output gate to obtain the hidden state of the backward LSTM at each time step.
[0121] Finally, the forward LSTM processes data from the start of the sequence, and the backward LSTM processes from the end. The weighted combination of their outputs gives the hidden state output of the Bi-LSTM.
[0122] In step 2, the multi-head attention mechanism is incorporated into the constructed Bi-LSTM network architecture, and multiple outputs of the model are used as the input of the multi-head self-attention. The implementation steps include:
[0123] First, perform feature encoding on it. The output feature vectors of the Bi-LSTM are and Map them to different subspaces through linear transformation to obtain the query matrix, key matrix, and value matrix. The linear transformation calculation formula is as follows:
[0124] q i = W q f i
[0125] k i = W k f i
[0126] v i = W v f i
[0127] where W q , W k and W v are learnable weight matrices, and thus the query matrix Q = [q 1 , …, q n , key matrix K = [k 1 , …, k nThe sum matrix V = [v 1 , …, v n . Next, in order to obtain the attention distribution, the similarity between the query vector q and the key vector k is calculated, and the calculation expression is as follows:
[0128]
[0129] By normalizing the similarity scores, the weight distribution of each query vector relative to all key vectors is obtained. These weights reflect the degree of attention between different elements. Multiply these weight distributions by the corresponding value vectors and perform a weighted sum on the results to obtain the output result
[0130]
[0131] The calculation process of the self-attention mechanism can be expressed as
[0132]
[0133] In the formula, d k represents the dimension of the key matrix K. To further enhance the expressive power of the model, by adding multiple independent self-attention heads from different perspectives, the time series features are mapped to different sub-representation spaces to guide the model to learn the correlation information in different sub-representation spaces to obtain richer representations. The calculation formula of the multi-head self-attention:
[0134] MHSA Concatenate(SAM1, SAM2, …, SAM n )
[0135] After normalization, it is multiplied by the value vector and weighted and summed to obtain the output of each attention head. Finally, the outputs of multiple heads are concatenated to obtain the result of the multi-head self-attention mechanism, enhancing the ability to capture the correlation relationship of data features.
[0136] Connect a fully connected layer after the output layer of the multi-head self-attention mechanism. The role of the fully connected layer is to further integrate and transform the features output by the multi-head self-attention mechanism and map them to a space more suitable for the prediction task;
[0137] Connect a SoftMax layer after the fully connected layer. The fully connected layer is used to integrate features, and the SoftMax layer maps the output to the prediction of the tobacco leaf quality category to construct a tobacco leaf quality prediction model;
[0138] By constructing the model through the above steps, it combines the processing ability of Bi-LSTM for time series data and the ability of the multi-head self-attention mechanism to capture the correlation relationship of data features, making up for the complexity of tobacco leaf quality data and the deficiencies of traditional models, laying a foundation for subsequent use of PSO to optimize model parameters and predict tobacco leaf quality.
[0139] Step 3 of this embodiment for optimizing the model parameters using PSO includes:
[0140] First, initialize the particle swarm and set the parameters. Determine the particle swarm size. Each particle represents a set of parameters to be optimized in the Bi-LSTM model, including randomly generating the position and velocity of the particle. Set parameters such as the particle swarm size, maximum number of iterations, inertia weight, cognitive factor, and social factor.
[0141] Then, calculate the fitness value of each particle. Assign the parameters represented by the particle to the Bi-LSTM model, and use the training set data for model training. During the training process, calculate the loss of the model on the training set according to the set loss function. After training, evaluate the model performance using the validation set data, and this performance metric is used as the fitness value of the particle.
[0142] Update the velocity and position of the particle. According to the current position, velocity, individual best position, and global best position of the particle, use the following formulas to update the velocity and position:
[0143] Velocity update formula:
[0144] Position update formula:
[0145] Where and are the velocity and position of the th particle in the d - th dimension in the i - th iteration respectively. ω is the inertia weight, c1 and c2 are the cognitive factor and social factor, r1 and r2 are random numbers between [0, 1], is the individual best position of the i - th particle, is the global best position.
[0146] In each iteration, the particle updates its velocity and position in the following way:
[0147] Velocity update: Comprehensively consider the particle's own historical best position and the global best position to adjust the velocity. The inertia part keeps the particle with a certain motion inertia; the cognitive part makes the particle approach its own historical best position; the social part makes the particle approach the historical best position of the entire group. In this way, the particle continuously adjusts its direction and velocity in the search space to find a better parameter combination.
[0148] Position update: Adjust the position of the particle according to the updated velocity, which is also the value assignment of the parameter.
[0149] By continuously iterating the above process, the particle swarm can search in the hyperparameter space, avoiding the blindness and inefficiency of manual hyperparameter tuning or simple random search. Eventually, an optimal combination of hyperparameters is found, thereby improving the performance of the Bi-LSTM model, such as enhancing the training efficiency and avoiding getting trapped in local optimal solutions.
[0150] The mean squared error loss function expression for calculating the particle fitness value is as follows:
[0151]
[0152] where m is the number of training samples, y j is the actual tobacco leaf quality value, is the model prediction value.
[0153] Iterative optimization and termination condition judgment: Repeat the above steps of calculating the fitness value, updating the velocity, and updating the position until the maximum number of iterations is reached, and output the global optimal position as the optimized parameters of the Bi-LSTM model.
[0154] After training the life prediction model using the sample set, calculate the model evaluation metrics. If the model evaluation metrics do not meet the requirements, then adjust the specified parameters of the tobacco leaf quality prediction model and then use the sample set to train the prediction model again until the model evaluation metrics meet the requirements.
[0155] Finally, evaluate the model performance, and comprehensively evaluate the model using metrics such as root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R 2 ). RMSE is used to measure the deviation degree between the predicted value and the actual value, MAE reflects the average size of the prediction error, and R 2 evaluates the goodness of fit of the model to the data. Through the analysis of these metrics, intuitively display the prediction accuracy and reliability of the model, providing a basis for the improvement and application of the model.
[0156] In summary, the present invention proposes a tobacco leaf quality prediction method based on an improved Bi-LSTM model with PSO and multi-head self-attention. This method includes: collecting key time-series data during the tobacco leaf threshing and redrying process, establishing a complete data set and preprocessing it to generate a standardized sample set; constructing a Bi-LSTM model combined with the multi-head self-attention mechanism, using the multi-head self-attention mechanism to extract key feature information in the threshing and redrying parameters, and using it as the input of the Bi-LSTM network to capture long time-series features; adaptively optimizing key hyperparameters such as the learning rate and the number of hidden layer neurons of Bi-LSTM through the particle swarm optimization algorithm to ensure the accuracy and convergence of the model. Finally, input the time-series data of the threshing and redrying process into the trained model to output the tobacco leaf quality prediction result. The present invention enhances the generalization ability and stability of the model through PSO dynamic tuning and the multi-head self-attention mechanism, and significantly improves the prediction accuracy of tobacco leaf quality.
[0157] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.
[0158] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.
Claims
1. A tobacco leaf quality prediction method based on PSO improved Bi-LSTM model and multi-head self-attention relates to the field of tobacco leaf quality prediction. This method optimizes the parameters of the Bi-LSTM model by introducing a particle swarm optimization PSO algorithm, which can effectively improve the training efficiency and prediction accuracy of the model. At the same time, the multi-head self-attention mechanism enables the model to adaptively focus on the important features of each time step in the input data, thereby enhancing the learning ability of complex time series patterns in tobacco leaf quality prediction.
2. The tobacco leaf quality prediction method based on PSO improved Bi-LSTM model and multi-head self-attention as claimed in claim 1 is characterized in that: The method comprises the following steps: Step 1: real-time collection of key time series parameter data during tobacco leaf threshing and redrying, corresponding tobacco leaf redrying quality index data, and tobacco leaf image data after threshing, integrating the data to establish a database, and preprocessing the data in the database to obtain a standardized sample set; Step 2: Build a quality prediction model. First, complete the construction of the Bi-LSTM network architecture. In order to make up for the complexity of tobacco leaf quality data and the shortcomings of traditional models, then integrate the multi-head attention mechanism, use the output of the Bi-LSTM network as the input of the multi-head self-attention mechanism, enhance the Bi-LSTM's ability to capture the correlation between data features, and provide richer feature information for subsequent tobacco leaf quality prediction. Step 3: When optimizing the Bi-LSTM network, the particle swarm optimization algorithm is used to adjust the hyperparameters, realize dynamic optimization of hyperparameter adaptation, and determine the optimal parameter combination to improve the accuracy and stability of the model; Step 4: Use the sample set to train the model to obtain the best model performance, input the real-time collected leaf threshing and redrying process parameter data into the trained model, and output the tobacco leaf quality prediction result.
3. The tobacco leaf quality prediction method based on PSO improved Bi-LSTM model and multi-head self-attention as claimed in claim 2 is characterized in that: In step 1, after the threshing process is completed, tobacco leaf images from multiple perspectives and multiple regions are collected, including images of the front side, back side and different parts of the tobacco leaves, to fully reflect the appearance characteristics of the tobacco leaves; The process parameter data such as temperature, humidity, wind speed, etc. during the re-roasting process are collected in real time to ensure the accuracy and continuity of the data. At the same time, the time series information such as the processing time of each stage in the re-roasting process is recorded to facilitate the subsequent analysis of the time series characteristics of the data. The matrix X of the database can be expressed as: Where m is the number of samples collected, which represents the data of each tobacco leaf threshing and redrying process; n is the number of process parameters related to tobacco leaf quality, including the key timing parameters of temperature, humidity, time, wind speed, and pressure during the threshing and redrying process; x ij is the jth process parameter value in the i-th sample (i=1,2,…,m; j=1,2,…,n); k is the number of features extracted from the tobacco leaf image after threshing; i pq represents the qth image feature corresponding to the pth sample; y i It represents the tobacco leaf quality index value corresponding to the i-th sample, such as the rate of large and medium leaves, the rate of stems in leaves, nicotine content, sugar content, etc.
4. The tobacco leaf quality prediction method based on PSO improved Bi-LSTM model and multi-head self-attention as claimed in claim 3 is characterized in that: In step 1, when the data in the database is preprocessed to obtain a sample set, the data cleaning and normalization preprocessing includes: Clean the collected data, identify and correct outliers by setting a threshold range, by calculating the mean μ and standard deviation σ of each process parameter, and treating data beyond the threshold range of μ±kσ (k is 2 or 3) as outliers; the outliers can be corrected or deleted according to the distribution of the data and actual process knowledge; missing values can be filled with time series-based interpolation methods to ensure data continuity; The cleaned data was normalized to ensure that the numerical scale of the data was consistent and to prevent the data after weight operation from causing deviation in training; each process parameter and tobacco leaf quality parameter was normalized to the interval [0,1] to eliminate the differences in dimensions and numerical ranges between different parameters, which facilitated the training and analysis of subsequent models; the L1 regularized logistic regression model was used to screen features of the normalized database, and the screened features were used as the sample set; This method uses the L1 regularized logistic regression method suitable for high-dimensional data sets. It produces a sparse weight matrix by sparsely calculating the coefficients of the model, so that the model focuses on the features whose coefficients are non-zero values, thereby achieving the purpose of feature selection. It is constructed using the Kears module in Python; its working principle is as follows: Among them, |w i | represents the L1 norm of the model's weight vector, that is, the absolute value of each weight, w i represents the weight vector, n is the number of weights; The data are integrated to establish a database, and the data in the database are preprocessed to obtain a standardized sample set.
5. The tobacco leaf quality prediction method based on PSO improved Bi-LSTM model and multi-head self-attention as claimed in claim 4 is characterized in that: The step 2 also includes the following steps: Step 2.1, determine the network architecture parameters, determine the input layer, hidden layer and output layer structure of the Bi-LSTM network, and the number of input layer nodes is determined according to the number of data features collected related to tobacco leaf quality; Set the hidden layer structure of Bi-LSTM, including the number of hidden units of forward LSTM and backward LSTM. The number of hidden units will affect the model's ability to learn data features. It is determined through experiments and tuning. It is first set to 64, and then adjusted according to the performance of the model on the validation set. Configure the output layer. The number of nodes in the output layer depends on the number of tobacco leaf quality indicators to be predicted. Step 2.2: Construct an LSTM unit and define two parameters: the weight matrix and the bias vector in the LSTM unit. For the forward LSTM, at each time step t, the dimension of the received input data is the same as the number of nodes in the input layer, and the dimension of the hidden state of the previous time step is the number of hidden units; The forget gate determines the cell state C at the previous moment t+1 Some information in needs to be retained or forgotten. The calculation formula of the forget gate is as follows: z f =σ(W f ·[h t-1 ,x t ]+b f ) Among them, W f is the weight matrix of the forget gate, whose dimension is (number of hidden units + number of input layer nodes) × number of hidden units, b f The bias vector of the forget gate, whose dimension is the number of hidden units, and σ is the Sigmoid activation function; The calculation formula of the input gate is as follows: z i =σ(W i ·[h t-1 ,x t ]+b i ) Among them, W i and b i They are the weight matrix and bias vector of the input gate, respectively, and their dimensions are consistent with those of the forget gate; At the same time, the candidate value C for updating the cell state is calculated t =tanh(W C ·[h t-1 ,x t ]+b C ),W C and b C are the corresponding weight matrices and bias vectors; The calculation formula for updating the cell state is as follows: C t =z f ·C t-1 +z i ·C t 。; The calculation formula of the output gate is as follows: z o =σ(W o ·[h t-1 ,x t ]+b o ) Where W o and b o is the weight matrix and bias vector of the output gate, and finally the hidden state h of the current time step is obtained t =z o tanh(C t ); The calculation process of backward LSTM is similar to that of forward LSTM, but the data processing direction is opposite, starting from the end of the sequence to the beginning. Backward LSTM receives input data x at each time step. t and the hidden state h at the next time step t+1 , and randomly initialize the vector, and calculate according to the above calculation steps of forget gate, input gate, cell state update and output gate to obtain the hidden state of backward LSTM at each time step; Step 2.3: The forward LSTM processes data from the beginning of the sequence, and the backward LSTM processes it from the end. The weighted combination of the two outputs is used to obtain the hidden state output of the Bi-LSTM.
6. The tobacco leaf quality prediction method based on PSO improved Bi-LSTM model and multi-head self-attention as claimed in claim 5, characterized in that: In step 2, the multi-head attention mechanism is integrated into the constructed Bi-LSTM network architecture, and multiple outputs of the model are used as inputs of the multi-head self-attention. The implementation steps include: Step S1: perform feature encoding. The output feature vector of Bi-LSTM is: and It is mapped to different subspaces through linear transformation to obtain the query matrix, key matrix and value matrix, where the linear transformation calculation formula is as follows: q i =W q f i k i =W k f i v i =W v f i Among them, W q , W k and W v is a learnable weight matrix, from which we can get the query matrix Q = [q 1 ,…,q n ], key matrix K = [k 1 ,…,k n ] and value matrix V = [v 1 ,…,v n ]; Step S2: In order to obtain the attention distribution, the similarity between the query vector q and the key vector k is calculated. The calculation expression is as follows: By normalizing the similarity scores, we can obtain the weight distribution of each query vector relative to all key vectors. The weight reflects the degree of attention between different elements. We multiply the weight distribution with the corresponding numerical vector and perform weighted summation on the results to obtain the output result. The calculation process of the self-attention mechanism can be expressed as Where, d k Represents the dimension of the key matrix K. To further enhance the expressiveness of the model, multiple independent self-attention heads are added from different angles to map the temporal features to different sub-representation spaces, so as to guide the model to learn the associated information of different sub-representation spaces to obtain richer representations. The calculation formula of multi-head self-attention is: MHSA=Concatenate(SAM1,SAM2,…,SAM n ) Step S3: After normalization, the weighted sum is multiplied with the value vector to obtain the output of each attention head, and finally multiple head outputs are concatenated to obtain the result of the multi-head self-attention mechanism, thereby enhancing the ability to capture the correlation between data features.
7. The tobacco leaf quality prediction method based on PSO improved Bi-LSTM model and multi-head self-attention as claimed in claim 6, characterized in that: In step 2, a fully connected layer is connected after the output layer of the multi-head self-attention mechanism. The function of the fully connected layer is to further integrate and transform the features output by the multi-head self-attention mechanism and map them to a space that is more suitable for the prediction task; The SoftMax layer is connected after the fully connected layer. The fully connected layer is used to integrate features. The SoftMax layer maps the output to the prediction of the tobacco leaf quality category to build a tobacco leaf quality prediction model.
8. The tobacco leaf quality prediction method based on PSO improved Bi-LSTM model and multi-head self-attention as claimed in claim 7, characterized in that: The step 3 also includes the following steps: Step 3.1 is to initialize the particle swarm and set the parameters. Determine the size of the particle swarm. Each particle represents a set of parameters to be optimized in the Bi-LSTM model, including the position and speed of randomly generated particles. Set the size of the particle swarm, the maximum number of iterations, the inertia weight, the cognitive factor, and the social factor. Step 3.2 is to calculate the fitness value of each particle, assign the parameters represented by the particle to the Bi-LSTM model, and use the training set data to train the model. During the training process, the loss of the model on the training set is calculated according to the set loss function. After the training is completed, the model performance is evaluated with the validation set data, and the performance index is used as the fitness value of the particle; Step 3.3: Evaluate the model performance using the root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R). 2 The three indicators are used to comprehensively evaluate the model; RMSE is used to measure the degree of deviation between the predicted value and the actual value, MAE reflects the average size of the prediction error, and R 2 The goodness of fit of the model to the data is evaluated; through the analysis of these indicators, the prediction accuracy and reliability of the model are intuitively displayed, providing a basis for the improvement and application of the model.
9. The tobacco leaf quality prediction method based on PSO improved Bi-LSTM model and multi-head self-attention as claimed in claim 8, characterized in that: In step 3.2, according to the current position, speed, individual optimal position and global optimal position of the particle, the speed and position are updated using the following formula: Speed update formula: Position update formula: in and are the velocity and position of the d-th particle in the ith iteration, ω is the inertia weight, c1 and c2 are cognitive factors and social factors, r1 and r2 are random numbers between [0,1], is the individual optimal position of the ith particle, is the global optimal position.
10. The tobacco leaf quality prediction method based on PSO improved Bi-LSTM model and multi-head self-attention as claimed in claim 9, characterized in that: In step 3.2, the particle updates its velocity and position according to the following method: Speed update: adjust the speed by considering the particle's own historical optimal position and the global optimal position; inertia part Keep particles in a certain inertia; cognitive part Make particles move closer to their best historical position; social part Make the particles approach the best historical position of the entire group; in this way, the particles continuously adjust their direction and speed in the search space to find a better combination of parameters; Position update: adjust the particle position according to the updated speed, which is also the value of the parameter; By continuously iterating the above process, the particle swarm can search in the hyperparameter space, avoiding the blindness and inefficiency of manual parameter adjustment or simple random search; and finally find a set of optimal hyperparameter combinations, thereby improving the performance of the Bi-LSTM model, such as improving training efficiency and avoiding falling into local optimal solutions. The mean square error loss function expression set for calculating the particle fitness value is as follows: Among them, m is the number of training samples, y j is the actual tobacco leaf quality value, is the model prediction value; Iterate optimization and determine the termination condition, repeat the above steps of calculating fitness value, updating speed and position until the maximum number of iterations is reached, and output the global optimal position as the optimized Bi-LSTM model parameters.
Citation Information
Cited By
Process parameter prediction control method and system in tobacco leaf feeding process, medium and terminal
CN120802644A