Prediction method, medium and system for chlorophyll concentration of fresh tobacco leaves

By combining image analysis and spectral analysis technology, the combined model of convolutional neural network and long and short-term memory network is used to achieve rapid and accurate prediction of chlorophyll concentration in fresh tobacco leaves, solving the problems of inefficiency and insufficient accuracy in the existing technology, and achieving efficient and accurate detection effects.

CN120031822AActive Publication Date: 2025-05-23JIANGXI TOBACCO CO JIAN CO +1

Patent Information

Application Number
CN202510098229.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-23
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prior art is difficult to achieve rapid and accurate prediction of chlorophyll concentration in fresh tobacco leaves, and there are problems of low efficiency and insufficient accuracy.

Method used

A method based on convolutional neural network and feature fusion is adopted, and the combination of image analysis and spectral analysis technology is used to achieve automatic extraction and fusion of features using deep learning models. The method includes collecting image and spectral data, performing feature extraction and fusion, constructing a multi-source feature fusion data set, and using a combined model of a convolutional neural network and a long and short-term memory network for prediction.

Benefits of technology

The rapid and accurate prediction of chlorophyll concentration in fresh tobacco leaves is achieved, the detection efficiency and accuracy is improved, and the detection of a sample can be completed within 4 to 6 minutes, and the relative error is controlled within 3%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031822A_ABST
    Figure CN120031822A_ABST
Patent Text Reader

Abstract

The invention provides a fresh tobacco leaf chlorophyll concentration prediction method, medium and system, and belongs to the technical field of electrical digital data process.The fresh tobacco leaf chlorophyll concentration prediction method comprises the steps that standardized image data and spectrum data are collected, initial features are extracted through color space conversion, texture analysis and spectrum preprocessing, and a fresh tobacco leaf chlorophyll concentration prediction result is obtained; and selecting an optimal feature combination by adopting a recursive feature elimination and competitive adaptive reweighting algorithm, and realizing feature fusion by utilizing a method based on matrix decomposition. A deep learning model including a sequence folding layer, a convolutional layer, an attention layer and the like is constructed, dynamic weighting of features is realized in combination with a channel attention mechanism, rapid and accurate prediction of the chlorophyll concentration is finally completed, an end-to-end learning system is formed in the whole prediction process, and automatic mapping from original data to a prediction result is realized. The technical problem that the chlorophyll concentration of the fresh tobacco leaves is difficult to quickly and accurately predict in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electronic digital data processing, and in particular, relates to a method, a medium and a system for predicting the chlorophyll concentration of fresh tobacco leaves. Background Art

[0002] The chlorophyll concentration of fresh tobacco leaves is a key indicator for evaluating the maturity and quality of tobacco leaves, and is of great significance for guiding tobacco leaf harvesting and improving tobacco leaf quality. Traditional chlorophyll concentration determination methods mainly include chemical extraction method and spectrophotometry method. The chemical extraction method is to extract chlorophyll through organic solvents such as ethanol or acetone, and then use a spectrophotometer to measure the absorbance to calculate the chlorophyll concentration. Although this method has high accuracy, it has the disadvantages of strong sample destructiveness, high reagent consumption, complex operation, and long measurement cycle, which makes it difficult to meet the rapid detection needs in tobacco leaf production. The spectrometry method collects the reflectance spectrum of the leaves and establishes a mathematical model between the spectral characteristics and the chlorophyll concentration for prediction. It has the advantage of non-destructive detection, but is greatly affected by factors such as ambient light and measurement angle, and professional spectrometers are expensive.

[0003] In recent years, with the development of computer vision technology, chlorophyll concentration determination methods based on image analysis have gradually attracted attention. This type of method establishes a prediction model by extracting features such as color and texture of leaf images, which is low-cost and easy to operate. However, methods that rely solely on image features are easily affected by shooting conditions and cannot obtain information about the internal structure of the leaves, resulting in unstable prediction accuracy. In addition, although portable chlorophyll meters are easy to operate, their measurement principle is too simple. They only estimate the chlorophyll content by measuring the reflectivity of a specific wavelength, which has low accuracy and is not suitable for precise measurement.

[0004] Various measurement methods in the prior art either have problems of low efficiency or insufficient precision, and it is difficult to achieve the unity of rapidity and precision at the same time. Especially in the large-scale production of tobacco leaves, a large number of samples need to be tested in real time, which places high demands on the efficiency and precision of the measurement methods. At the same time, most of the existing prediction models use a single feature source or a simple statistical method, which does not fully utilize the complementary advantages of multi-source data, and lacks an effective feature extraction and fusion mechanism. In summary, it is difficult to achieve the technical problem of rapid and accurate prediction of the chlorophyll concentration of fresh tobacco leaves in the prior art. Summary of the invention

[0005] In view of this, the present invention provides a method, medium and system for predicting the chlorophyll concentration of fresh tobacco leaves, which can solve the technical problem that it is difficult to quickly and accurately predict the chlorophyll concentration of fresh tobacco leaves in the prior art.

[0006] The present invention is implemented as follows: In a first aspect, the present invention provides a method for predicting the chlorophyll concentration of fresh tobacco leaves, comprising the following steps: collecting fresh tobacco leaf samples and taking photos to obtain fresh tobacco leaf images, collecting near-infrared spectral data of the fresh tobacco leaf samples, and generating a spectral data matrix; the method constructs a feature correlation matrix based on a Pearson correlation analysis method and uses a recursive feature elimination method to screen the optimal feature combination to form a preferred image feature data set, and simultaneously uses a competitive adaptive reweighting algorithm, a continuous projection algorithm, and a sparse representation algorithm to extract features from the preprocessed spectral data matrix to form a preferred spectral feature data set; based on matrix decomposition, a fusion weight coefficient is determined, and a feature interaction matrix of the preferred image feature data set and the preferred spectral feature data set is calculated to generate a multi-source feature fusion data set; a combined model of a convolutional neural network and a long short-term memory network is constructed, the combined model includes a sequence folding layer, a convolution layer, a global average pooling layer, a channel attention layer, an anti-folding layer, a long short-term memory layer, a fully connected layer and a regression layer, and the weight of the channel attention layer of the combined model is adjusted according to an importance scoring matrix to generate a chlorophyll concentration prediction model.

[0007] Among them, the step of collecting fresh tobacco leaf samples and taking photos to obtain fresh tobacco leaf images specifically includes collecting 50 fresh tobacco leaf samples at four stages, namely, underripe, moderately ripe, moderately ripe, and mature, taking photos of the fresh tobacco leaf samples using a smart phone or a digital camera in a darkroom environment, performing background segmentation on the fresh tobacco leaf images using a color segmentation method based on the Ycbcr color space, extracting R, G, B, L, a, and b color channel eigenvalues ​​and texture eigenvalues ​​based on the grayscale co-occurrence matrix, and generating an initial image feature data set.

[0008] Among them, the step of collecting the near-infrared spectral data of the fresh tobacco leaf sample specifically collects the near-infrared spectral data of the fresh tobacco leaf sample in the 900 to 1700 nanometer band, collects 6 measurement points on both sides of the leaf tip, leaf middle, and leaf base of the fresh tobacco leaf sample, avoids the main vein and branch veins, and calculates the average value of the 6 measurement points as the initial near-infrared spectral data.

[0009] Among them, the step of generating a spectral data matrix specifically adopts normalization method, standard normal variable transformation method, and multivariate scattering correction method to preprocess the initial near-infrared spectral data to generate a preprocessed spectral data matrix, and adopts competitive adaptive reweighting algorithm, continuous projection algorithm, and sparse representation algorithm to extract features of the preprocessed spectral data matrix respectively, selects the feature combination with the best training performance, and generates an optimal spectral feature data set.

[0010] Among them, the step of determining the fusion weight coefficient based on matrix decomposition is specifically to calculate the correlation coefficient between each eigenvalue and the chlorophyll concentration in the initial image feature data set based on the Pearson correlation analysis method, construct a feature correlation matrix, use recursive feature elimination method to process the feature correlation matrix, select the feature combination with the best training performance, and form a preferred image feature data set.

[0011] Among them, the step of generating a multi-source feature fusion data set specifically calculates the feature interaction matrix of the preferred image feature data set and the preferred spectral feature data set, and performs vector multiplication and fusion of the preferred image feature data set and the preferred spectral feature data set according to the fusion weight coefficient.

[0012] Among them, the step of generating a chlorophyll concentration prediction model is specifically to use a spectrophotometer to measure the absorbance of the fresh tobacco leaf sample at wavelengths of 649 nanometers and 665 nanometers, generate a chlorophyll concentration measured data set according to the chlorophyll content calculation formula, and establish an importance scoring matrix for each feature in the multi-source feature fusion data set.

[0013] Among them, the step of using the chlorophyll concentration prediction model is specifically to obtain the fresh tobacco leaves to be tested, repeat the operations described in claims 1 to 6, generate a multi-source feature fusion data set of the fresh tobacco leaves to be tested, input the multi-source feature fusion data set of the fresh tobacco leaves to be tested into the chlorophyll concentration prediction model, and output the chlorophyll concentration prediction value of the fresh tobacco leaves to be tested.

[0014] A second aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores program instructions, and when the program instructions are run in a computer, they are used to execute the above-mentioned method for predicting the chlorophyll concentration of fresh tobacco leaves.

[0015] The third aspect of the present invention provides a fresh tobacco leaf chlorophyll concentration prediction system, comprising the above-mentioned computer-readable storage medium, wherein the system is any one of a computer, a server, and a single-chip microcomputer, wherein the computer-readable storage medium is arranged in the system, and wherein the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.

[0016] Compared with the prior art, the present invention provides a method, medium and system for predicting the chlorophyll concentration of fresh tobacco leaves. The present invention proposes a method for predicting the chlorophyll concentration of fresh tobacco leaves based on convolutional neural networks and feature fusion. The method innovatively combines image analysis and spectral analysis technology, and realizes automatic feature extraction and fusion through a deep learning model. The standardized data collection process and multiple preprocessing technologies are used to effectively reduce the interference of environmental factors and improve data quality.

[0017] The feature extraction and fusion scheme designed by the present invention fully considers the characteristics of different feature sources, and screens the optimal feature combination through recursive feature elimination method and competitive adaptive reweighting algorithm, avoiding the interference of redundant information. The fusion strategy based on matrix decomposition is adopted, which not only maintains the physical meaning of the original features, but also realizes the effective fusion between features. At the same time, the channel attention mechanism is introduced, so that the model can adaptively adjust the weights of different features, thereby improving the robustness of prediction.

[0018] The present invention successfully solves the problem of unifying rapidity and accuracy through the powerful feature learning ability of the deep learning model. This method not only inherits the fast and simple characteristics of image analysis, but also overcomes the defect of unstable accuracy of the single feature method, and realizes the rapid and accurate prediction of the chlorophyll concentration of fresh tobacco leaves. The entire measurement process does not require complex sample pretreatment, and can realize online detection, which significantly improves the detection efficiency. It solves the technical problem that it is difficult to realize the rapid and accurate prediction of the chlorophyll concentration of fresh tobacco leaves. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flow chart of the method of the present invention;

[0020] Figure 2 This is a flow chart of the method for predicting chlorophyll concentration of fresh tobacco leaves based on convolutional neural network and feature fusion;

[0021] Figure 3 This is the structure diagram of CNN-LSTM;

[0022] Figure 4 It is a schematic diagram of vector multiplication and fusion features;

[0023] Figure 5 This is the image feature evaluation result map;

[0024] Figure 6 The evaluation results of 9 combinations of pre-processing and key wave point selection algorithms are shown in the figure;

[0025] Figure 7 This is the evaluation result diagram of CNN-LSTM based on fusion features;

[0026] Figure 8 The image feature CNN-LSTM evaluation result diagram is shown;

[0027] Fig. 9 This is the CNN-LSTM evaluation result diagram of near-infrared spectral features. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solution and advantages of the embodiments of the present invention more clear, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0029] like Figure 1 FIG. 1 is a flow chart of a method for predicting chlorophyll concentration of fresh tobacco leaves provided by the first aspect of the present invention. The method comprises the following steps:

[0030] S01. Collect 50 fresh tobacco leaf samples at four stages: underripe, moderately ripe, moderately ripe, and mature, and take photos of the fresh tobacco leaf samples using a smart phone in a darkroom environment to obtain images of the fresh tobacco leaves;

[0031] S02, performing background segmentation on the fresh tobacco leaf image using a color segmentation method based on a Ycbcr color space, extracting R, G, B, L, a, b color channel eigenvalues ​​and texture eigenvalues ​​based on a gray-level co-occurrence matrix, and generating an initial image feature data set;

[0032] S03, calculating the correlation coefficient between each characteristic value and chlorophyll concentration in the initial image feature data set based on the Pearson correlation analysis method, and constructing a characteristic correlation matrix;

[0033] S04, using a recursive feature elimination method to process the feature correlation matrix, select a feature combination with the best training performance, and form an optimal image feature data set;

[0034] S05, collecting near infrared spectrum data of the fresh tobacco leaf sample in the 900 to 1700 nanometer band, collecting 6 measurement points at the leaf tip, the middle of the leaf, and both sides of the leaf base of the fresh tobacco leaf sample, avoiding the main vein and the branch veins, and calculating the average value of the 6 measurement points as the initial near infrared spectrum data;

[0035] S06, preprocessing the initial near-infrared spectrum data by using a normalization method, a standard normal variable transformation method, and a multivariate scattering correction method to generate a preprocessed spectrum data matrix;

[0036] S07, using a competitive adaptive reweighting algorithm, a continuous projection algorithm, and a sparse representation algorithm to extract features from the preprocessed spectral data matrix, respectively, selecting a feature combination with the best training performance, and generating a preferred spectral feature data set;

[0037] S08, calculating a feature interaction matrix of the preferred image feature data set and the preferred spectral feature data set, and determining a fusion weight coefficient based on matrix decomposition;

[0038] S09, performing vector multiplication and fusion of the preferred image feature dataset and the preferred spectral feature dataset according to the fusion weight coefficient to generate a multi-source feature fusion dataset;

[0039] S10, using a spectrophotometer to measure the absorbance of the fresh tobacco leaf sample at wavelengths of 649 nanometers and 665 nanometers, and generating a chlorophyll concentration measured data set according to a chlorophyll content calculation formula;

[0040] S11, constructing a combined model of a convolutional neural network and a long short-term memory network, wherein the combined model includes a sequence folding layer, a convolution layer, a global average pooling layer, a channel attention layer, an anti-folding layer, a long short-term memory layer, a fully connected layer and a regression layer;

[0041] S12, establishing an importance score matrix for each feature in the multi-source feature fusion data set;

[0042] S13, adjusting the channel attention layer weight of the combined model according to the importance scoring matrix, training the combined model using the multi-source feature fusion dataset and the chlorophyll concentration measured dataset, and obtaining a chlorophyll concentration prediction model;

[0043] S14, obtaining fresh tobacco leaves to be tested, repeating the operations of steps S01 to S09, and generating a multi-source feature fusion data set of fresh tobacco leaves to be tested;

[0044] S15, inputting the multi-source feature fusion data set of the fresh tobacco leaves to be tested into the chlorophyll concentration prediction model, and outputting the predicted value of the chlorophyll concentration of the fresh tobacco leaves to be tested.

[0045] The specific implementation of the above steps is described in detail below.

[0046] The specific implementation method of step S01 is to obtain high-quality fresh tobacco leaf image samples by constructing a standardized data collection environment. First, the four maturity levels of fresh tobacco leaves are determined, namely, under-ripe, mature, suitable and mature, and 50 samples are collected for each level to ensure the representativeness and statistical significance of the data. A standardized dark box environment is used during collection, and matte black paint is applied inside the dark box to avoid interference from ambient light. At the same time, an LED light source is installed on the top of the dark box, and the light source color temperature is 5500K and the illumination is 1000l ux. Use a smart phone for shooting, the pixel of the mobile phone camera is not less than 12 million, and the shooting parameters are fixedly set to: sensitivity ISO100, shutter speed 1 / 60 second, aperture value f / 2.0, white balance automatic mode. During the shooting process, the fresh tobacco leaf sample is flattened at the center of the bottom of the dark box to ensure that the leaf surface is perpendicular to the camera, and the shooting distance is fixed to 30 cm. After the shooting is completed, the image is preliminarily checked, and images with quality problems such as blur and improper exposure are eliminated, and qualified images are numbered and archived. The purpose of this step is to establish a standardized image acquisition process to ensure the data quality of subsequent analysis.

[0047] The specific implementation method of step S02 is to perform background segmentation and feature extraction on the acquired fresh tobacco leaf image. First, the Ycbcr color space is used for image segmentation. The reason for selecting this color space is that it can better separate brightness information and chromaticity information, and has strong robustness to illumination changes. By adaptively adjusting the Y component threshold, the threshold range is 16 to 235, and the preliminary separation of background and target is achieved. Subsequently, morphological operations are used for edge optimization, including using a circular structural element with a radius of 3 pixels to perform an open operation to remove noise, and using a circular structural element with a radius of 5 pixels to perform a closed operation to fill small holes. In the feature extraction stage, the mean values ​​of the R, G, and B channels of the RGB color space and the mean values ​​of the L, a, and b channels of the Lab color space are calculated respectively. At the same time, texture features are calculated based on the grayscale co-occurrence matrix, including statistics such as energy, contrast, correlation, and entropy. The distance parameter of the matrix is ​​set to 1, and the direction angle is set to 0 degrees, 45 degrees, 90 degrees, and 135 degrees. The average value of the four directions is taken as the final eigenvalue. The initial image feature data set is obtained through these processes, which lays the foundation for subsequent feature selection.

[0048] The specific implementation method of step S03 is to use the Pearson correlation analysis method to evaluate the degree of correlation between image features and chlorophyll concentration. First, all eigenvalues ​​are standardized so that their mean is 0 and their standard deviation is 1. Then the Pearson correlation coefficient of each eigenvalue and the chlorophyll concentration is calculated. The value range of the correlation coefficient is from negative 1 to 1. The closer the absolute value is to 1, the stronger the correlation. According to the empirical threshold, the features with an absolute value of the correlation coefficient greater than 0.3 are defined as moderate correlation, those greater than 0.5 are defined as strong correlation, and those greater than 0.7 are defined as extremely strong correlation. In this way, a feature correlation matrix is ​​constructed, and each element in the matrix represents the degree of correlation between the corresponding feature and the chlorophyll concentration, providing an important reference for subsequent feature selection.

[0049] The specific implementation method of step S04 is to screen the features using a recursive feature elimination method. The method first uses a support vector machine to build an initial model. The model uses a radial basis kernel function, the kernel parameter gamma is set to 0.1, and the penalty parameter C is set to 100. Then, the least important features are gradually deleted according to the size of the feature weights. After each deletion, the model is retrained and the performance indicators are recorded. The performance of different feature combinations is evaluated by a cross-validation method, with a validation set ratio of 30%, and the validation is repeated 10 times to take the average value. Finally, the feature combination with the smallest root mean square error on the validation set is selected as the preferred feature set. In typical cases, the number of features will be reduced from the initial 20 or so to 5 to 8. This can significantly reduce the feature dimension while maintaining the model performance.

[0050] The specific implementation method of step S05 is to collect near-infrared spectrum data of fresh tobacco leaf samples. Using a near-infrared spectrometer, the wavelength range is set to 900 to 1700 nanometers, the spectral resolution is 2 nanometers, and the number of scans is set to 32 times to improve the signal-to-noise ratio. A measuring point is set on each side of the tip, middle, and base of each tobacco leaf, for a total of 6 points, and the distance between the measuring point and the main vein and branch vein is not less than 5 mm. Each measuring point is repeatedly collected 3 times, and the average value is taken as the spectrum data of the point. Finally, the arithmetic mean of the 6 measuring points is calculated to obtain the initial near-infrared spectrum data of the tobacco leaf. During the collection process, the ambient temperature is maintained at 25±2 degrees Celsius and the relative humidity is maintained at 60%±5%.

[0051] The specific implementation method of step S06 is to pre-process the initial data of the near infrared spectrum. First, the normalization method is used to linearly map the data to the interval of 0 to 1 to eliminate the dimension effect. Then the standard normal variable transformation method is applied to convert the data into a distribution with a mean of 0 and a standard deviation of 1 to reduce the influence of noise. Finally, the multivariate scattering correction method is adopted to calculate the deviation of each wavelength point from the average spectrum, establish a linear regression equation for correction, and eliminate the influence of sample thickness and surface scattering. After processing by these three methods, a pre-processed spectrum data matrix is ​​obtained, and the number of rows of the matrix is ​​equal to the number of samples, and the number of columns is equal to the number of spectrum wavelength points.

[0052] The specific implementation method of step S07 is to extract key features from pre-processed spectral data. The competitive adaptive reweighting algorithm realizes feature selection through Monte Carlo sampling and exponentially decayed adaptive weights, the sampling number is set to 500 times, and the weight decay coefficient is set to 0.95. The continuous projection algorithm selects the wavelength point with minimum collinearity by orthogonal projection, and the projection threshold is set to 0.0001. The sparse representation algorithm uses LASSO regression to realize feature compression, and the regularization parameter is determined by cross-validation, and the typical value is 0.01. These three algorithms are applied to spectral data with different pretreatment methods, and 9 feature combinations are obtained. The performance of each combination is evaluated by BP neural network. The network contains a hidden layer, the number of neurons is twice the number of input features, the learning rate is 0.01, and the training round is 1000. The feature combination with the highest determination coefficient on the test set is selected as the preferred spectral feature data set.

[0053] The specific implementation method of step S08 is to calculate the interaction between the preferred features. First, the image features and spectral features are standardized, and then the mutual information matrix between the two types of features is calculated. The matrix is ​​subjected to singular value decomposition to obtain a left singular matrix, a singular value matrix, and a right singular matrix. The important feature combination is determined according to the size of the singular value, and the singular value with a cumulative contribution rate of 85% is usually retained. In this way, the fusion weight coefficient is obtained, and the coefficient value is between 0 and 1, which reflects the importance of different features in the fusion process.

[0054] The specific implementation method of step S09 is to realize feature fusion based on fusion weight coefficients. The image feature vector is multiplied by the corresponding weight coefficient, and the spectral feature vector is also multiplied by the respective weight coefficients, and then the two weighted vectors are connected to form a fused feature vector. This fusion method can maintain the physical meaning of the original features while reflecting the importance of different features. The dimension of the multi-source feature fusion data set finally obtained is equal to the sum of the two feature dimensions.

[0055] The specific implementation method of step S10 is to establish a chlorophyll concentration measured data set. Using a spectrophotometer, the wavelength accuracy is measured to ±0.1 nanometers, and the wavelength repeatability is ±0.05 nanometers. Take 0.2 grams of sample from each tobacco leaf, put it into a test tube containing 25 milliliters of 95% ethanol, and extract it for 24 hours at 4 degrees Celsius in the dark. Then measure the absorbance at wavelengths of 649 nanometers and 665 nanometers, and measure each sample 3 times to take the average value. Calculate the concentration value according to the chlorophyll content calculation formula, where the measurement accuracy of the absorbance is ±0.002 absorbance units.

[0056] The specific implementation method of step S11 is to construct a deep learning model. The sequence folding layer reorganizes the one-dimensional data into a two-dimensional form to facilitate subsequent convolution operations. The convolution layer uses a 1×1 convolution kernel, the number of channels in the first layer is 32, the second layer is 64, and the activation function uses ReLU. The global average pooling layer takes the average value of each feature map to compress the spatial dimension. The channel attention layer contains two fully connected layers. The number of neurons in the first layer is 1 / 16 of the number of channels, and the second layer is equal to the number of channels. Finally, it is normalized by the sigmoid function. The anti-folding layer restores the two-dimensional features to a sequence form. The number of hidden units in the long short-term memory layer is 6, and the tanh activation function is used. The number of neurons in the fully connected layer is 32, and the ReLU activation function is used. The regression layer outputs a single predicted value.

[0057] The specific implementation of step S12 is to evaluate the importance of features. The importance score of each feature is calculated using the random forest algorithm. The number of trees is set to 500, the maximum depth of the tree is the square root of the number of features, and the minimum number of sample splits is 2. By calculating the contribution of each feature to the improvement of model performance, an importance score matrix is ​​formed. The score value ranges from 0 to 1, and the closer to 1, the more important the feature.

[0058] The specific implementation method of step S13 is to optimize the model parameters. The initial weights of the channel attention layer are adjusted according to the feature importance matrix, and the channels corresponding to the features with high importance obtain larger initial weights. The Adam optimizer is used to train the model, with an initial learning rate of 0.001, a batch size of 32, and 200 training rounds. An early stopping strategy is used to prevent overfitting, and training is stopped when the validation set loss has not improved for 10 consecutive rounds. The determination coefficient of the final model on the test set is not less than 0.75, and the root mean square error does not exceed 0.18.

[0059] The specific implementation of step S14 is to process the sample to be tested. Repeat the aforementioned steps of image acquisition, feature extraction and data preprocessing for the fresh tobacco leaves to be tested, and ensure that the processing method is completely consistent with the training samples. The generated multi-source feature fusion data set has the same dimension and scale as the training data.

[0060] The specific implementation of step S15 is to predict the chlorophyll concentration. The multi-source feature fusion data of the sample to be tested is input into the trained model, and the model automatically completes the feature transformation and weight calculation, and outputs the concentration prediction value. The calculation time of the prediction process usually does not exceed 1 second, and real-time prediction can be achieved.

[0061] The important calculation process and matrix expressions involved in this scheme are described as follows:

[0062] The texture eigenvalue calculation method based on the gray level co-occurrence matrix is ​​specifically expressed as follows:

[0063]

[0064] In the formula, F texture is the comprehensive value of texture features; P ij is the gray-level co-occurrence matrix element; E ij is the energy characteristic value; C ij is the contrast characteristic value; H ij is the entropy eigenvalue; I ij is the eigenvalue of the moment of inertia; α 1 , α 2 , α 3 , α 4 is the weight coefficient and satisfies n is the gray level.

[0065] The correlation coefficient calculation method of the Pearson correlation analysis is specifically expressed as follows:

[0066]

[0067] In the formula, r xy is the correlation coefficient; x i ,y i are the observed values ​​of the two variables respectively; are the means of the two variables respectively; n is the sample size.

[0068] The calculation method of the feature fusion weight coefficient is specifically expressed as follows:

[0069]

[0070] W fusion =λ 1 W image +λ 2 W spectral ;

[0071] Where W is the weight matrix; w ij is the weight of the i-th feature on the j-th output; W fusion is the fusion weight; W image is the image feature weight; W spectral is the spectral feature weight; 1 ,λ 2 is the fusion coefficient, and λ 1 +λ 2 =1.

[0072] The chlorophyll concentration calculation formula is specifically expressed as follows:

[0073]

[0074] In the formula, C chl is the chlorophyll concentration value, in milligrams per gram; A 649 , A 665 are the absorbance values ​​at wavelengths of 649 nanometers and 665 nanometers respectively; k 1 , k 2 is the absorption coefficient, which takes values ​​of 18.16 and 6.63 respectively; V is the volume of the extract, in liters; D is the dilution factor; m is the sample mass, in grams; δ is the systematic error correction term, which takes values ​​in the range of 0.001 to 0.01.

[0075] The channel attention weight calculation method is specifically expressed as follows:

[0076]

[0077] α c =σ(W 2 δ(W 1 M c ));

[0078] Where M c is the global feature of channel c; F c (i, j) is the value of the feature map at position (i, j); H and W are the height and width of the feature map respectively; W 1 , W2 is the weight of the fully connected layer; δ is the ReLU activation function; σ is the sigmoid function; α c is the channel attention weight.

[0079] The feature importance score calculation method is specifically expressed as follows:

[0080]

[0081] In the formula, S i is the importance score of the i-th feature; ΔI ij is the reduction in impurity of feature i in decision tree j; T is the total number of decision trees; n is the total number of features.

[0082] The construction principles and significance of these equations are explained as follows:

[0083] 1. The texture feature value calculation equation comprehensively considers the four main texture features of energy, contrast, entropy and moment of inertia, and realizes adaptive fusion of features through weight coefficients, which can more comprehensively describe the image texture information;

[0084] 2. The Pearson correlation coefficient calculation equation measures the linear correlation between variables through standardized covariance, and the value range is limited to between -1 and 1, which is convenient for judging and comparing the strength of correlation;

[0085] 3. The feature fusion weight matrix adopts a linear weighting method and introduces a fusion coefficient to achieve adaptive fusion of image features and spectral features, and can dynamically adjust the weight according to the importance of different features;

[0086] 4. The chlorophyll concentration calculation equation is based on Beer-Lambert's law, taking into account the influence of experimental conditions such as sample mass and extraction volume, and introducing a systematic error correction term to improve the calculation accuracy;

[0087] 5. The channel attention weight calculation equation adopts a two-layer neural network structure, obtains channel features through global average pooling, and realizes adaptive learning of feature channel importance;

[0088] 6. The feature importance scoring equation is based on the random forest algorithm. It evaluates the feature importance by calculating the reduction in impurity of the feature in all decision trees to ensure the reliability and stability of the scoring results.

[0089] The derivation process and parameter sources of each equation in this scheme are explained as follows:

[0090] The derivation process of the texture feature value calculation equation:

[0091] First, construct the gray-level co-occurrence matrix P ij , obtain the statistical distribution of pixel pairs by scanning the image;

[0092] Then calculate four basic features: energy features Contrast characteristics Entropy characteristics Moment of inertia characteristics

[0093] Finally, the weight coefficient is introduced for weighted fusion. The weight coefficient is determined by minimizing the prediction error. The initial value is set to 0.25, and the optimal value is obtained through iterative optimization.

[0094] The advantage of this equation is that it comprehensively considers multiple texture features, and the weights can be adjusted according to the specific application scenario.

[0095] The derivation process of the Pearson correlation coefficient calculation equation is:

[0096] First, the original data is centered and the deviation of each variable from its mean is calculated;

[0097] Then the sum of the products of the deviations of the two variables is calculated as the numerator;

[0098] Calculate the square root of the sum of squares of the deviations of the two variables respectively as the denominator;

[0099] Finally, the normalization process is performed in the form of scores to ensure that the correlation coefficient is within the range of -1 to 1;

[0100] This equation can effectively eliminate the dimension effect and highlight the linear correlation between variables.

[0101] Derivation process of feature fusion weight matrix:

[0102] First, the initial weight matrix is ​​constructed, and the element values ​​are generated by random initialization;

[0103] Then the feature importance is calculated based on singular value decomposition: W = USV T , where S is the singular value matrix;

[0104] Determine the fusion coefficient according to the size of the singular value: λ 2 =1-λ 1 , where σ i is a singular value, k is the number of image features;

[0105] This equation realizes adaptive fusion of features and can dynamically adjust weights according to the importance of features.

[0106] Derivation process of chlorophyll concentration calculation equation:

[0107] The relationship between absorbance and concentration is derived based on Beer-Lambert's law: A = εbc;

[0108] Consider the dilution effect during sample extraction:

[0109] Introduce the absorption coefficient at different wavelengths: k 1 =18.16 corresponds to a wavelength of 649 nanometers, k 2 =6.63 corresponds to a wavelength of 665 nanometers;

[0110] Add the systematic error correction term δ and determine its range by calibration with standard samples;

[0111] This equation takes into account the influence of experimental conditions and improves the accuracy of concentration determination.

[0112] The derivation process of the channel attention weight calculation equation is:

[0113] First, the spatial dimension is compressed by global average pooling to extract channel-level features;

[0114] Then, a two-layer neural network structure is designed: the first layer reduces the dimension, and the number of neurons is 1 / 16 of the number of channels; the second layer increases the dimension and restores the original number of channels;

[0115] The ReLU activation function is used to provide nonlinear transformation capabilities, and the Sigmoid function normalizes the weights to the range of 0 to 1;

[0116] This equation can adaptively learn the importance of different channels and improve the efficiency of feature extraction.

[0117] The derivation process of the feature importance score calculation equation is:

[0118] Calculate the impurity of a node based on the Gini coefficient:

[0119] Calculate the reduction in impurity before and after feature splitting:

[0120] Count the reduction in impurity of features in all decision trees and perform normalization;

[0121] This equation improves the stability and reliability of feature importance evaluation through ensemble learning.

[0122] Supplementary explanation of parameter acquisition method:

[0123] The gray level n is determined by the image bit depth, and the typical value is 256;

[0124] The sample size n is determined according to the experimental design and is 200 in this scheme;

[0125] The feature map sizes H and W are determined by the network structure, and the typical value is 32×32;

[0126] The number of decision trees T is set to 500 to achieve a balance between computational efficiency and model performance.

[0127] A second aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores program instructions, and when the program instructions are run in a computer, they are used to execute the above-mentioned method for predicting the chlorophyll concentration of fresh tobacco leaves.

[0128] The third aspect of the present invention provides a fresh tobacco leaf chlorophyll concentration prediction system, comprising the above-mentioned computer-readable storage medium, wherein the system is any one of a computer, a server, and a single-chip microcomputer, wherein the computer-readable storage medium is arranged in the system, and wherein the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.

[0129] Specifically, the principle of the present invention is: the technical principle of the present invention is based on the feature extraction capability of multi-source information fusion and deep learning. At the data acquisition level, a standardized image acquisition environment and near-infrared spectral measurement scheme are adopted to ensure the consistency and reliability of the data. The image data contains the morphological characteristics and color information of the leaves, which have a direct physical correlation with the chlorophyll concentration, because changes in chlorophyll content will directly affect the optical properties of the leaves. The spectral data reflects the molecular vibration and chemical bond information of the leaf tissue, especially the spectral response in the near-infrared region can reflect the internal structural changes of the leaves.

[0130] At the feature processing level, the present invention adopts a multi-level feature extraction and selection strategy. Image features are extracted through color space conversion and texture analysis, and background interference is eliminated using a variety of spectral preprocessing methods. The feature selection process is based on the correlation analysis between the features and the target variable and the recursive elimination strategy to ensure the representativeness and effectiveness of the selected features. Feature fusion adopts a matrix decomposition-based method, which can capture the nonlinear relationship between features and better reflect the essential characteristics of the data than simple feature splicing.

[0131] In terms of model structure design, the present invention adopts a combined architecture of convolutional neural network and long short-term memory network, which can not only capture the local correlation of features, but also process the long-range dependency between sequence features. The introduction of channel attention mechanism enables the model to dynamically adjust the weights according to the contribution of different features to the prediction results, improving the model's adaptability and prediction accuracy. The entire prediction process forms an end-to-end learning system, realizing automatic mapping from raw data to prediction results.

[0132] A specific embodiment 1 of the present invention is provided below, and the specific implementation method of each step in this embodiment 1 is described in detail as follows.

[0133] The specific implementation method of step S01 is to obtain high-quality fresh tobacco leaf image samples by constructing a standardized data acquisition environment. First, the environmental parameters are set: the interior of the dark box is coated with matte black paint, the reflectivity is less than 3%, the LED light source color temperature is 5500K, the illumination is 1000l ux, and the fluctuation range is ±50l ux. The four maturity levels of fresh tobacco leaves are determined, namely, under-ripe, mature, moderately mature and mature, and 50 samples are collected for each level, with a total number of samples n=200. Use a smart phone for shooting, the camera pixel is not less than 12 million, and the shooting parameters are fixed as follows: sensitivity ISO100, shutter speed 1 / 60 second, aperture value f / 2.0, white balance automatic mode. The shooting distance is maintained at 30 cm, the allowable error is ±1 cm, and the leaf surface is ensured to be perpendicular to the camera, and the verticality error should be controlled within ±2 degrees. The image resolution obtained is 4000×3000 pixels, the color depth is 24 bits, and the storage format is lossless compressed PNG format. The purpose of this step is to establish a standardized image acquisition process to ensure the data quality of subsequent analysis. Image quality evaluation indicators include: clarity not less than 0.8 (normalized value), exposure uniformity deviation not exceeding ±0.2, color saturation within the range of 0.4 to 0.8. The ambient temperature of the entire acquisition process is controlled at 25 ± 2 degrees Celsius and the relative humidity is 60% ± 5%.

[0134] The specific implementation method of step S02 is to perform background segmentation and feature extraction on the acquired fresh tobacco leaf image. First, the RGB image is converted into the Ycbcr color space, and the conversion matrix is:

[0135]

[0136] The Y component represents brightness, with a value range of 16 to 235, and the Cb and Cr components represent chrominance, with a value range of 16 to 240. The Y component image is processed by the adaptive threshold segmentation method, and the initial threshold T 0 Set to 128, the iterative update formula is: where μ 1 (T k ) and μ 2 (T k ) are the threshold values ​​T k The average gray value of the foreground and background areas obtained by segmentation, the iteration termination condition is |T k+1 -T k|<0.5. After obtaining the binary mask, morphological operations are used to optimize the edge, including: using a circular structure element with a radius of 3 pixels for opening operations to remove noise, and using a circular structure element with a radius of 5 pixels for closing operations to fill small holes. Color features are calculated for the segmented image, including the mean of the R, G, and B channels in the RGB space and the mean of the L, a, and b channels in the Lab space. Texture features are calculated based on the gray-level co-occurrence matrix. The formula for calculating the texture feature value is:

[0137]

[0138] Among them, the gray level n is set to 256, the matrix distance parameter is 1, the direction angles are 0 degrees, 45 degrees, 90 degrees, and 135 degrees respectively, and the weight coefficient α 1 To α 4 Through cross-validation, the initial value is 0.25, and the optimization goal is to minimize the prediction error. The main purpose of this step is to extract an image data set that can characterize the characteristics of fresh tobacco leaves, providing a basis for subsequent feature selection. The calculated feature data needs to be normalized, using the minimum and maximum normalization method: where x min and x max are the minimum and maximum values ​​of the feature, respectively.

[0139] The specific implementation of step S03 is to evaluate the correlation between the image features and the chlorophyll concentration. The Pearson correlation analysis method is used, and the correlation coefficient calculation formula is:

[0140]

[0141] Among them, x i and i are the characteristic value and chlorophyll concentration value respectively, and is the corresponding average value, and the number of samples n is 200. According to the absolute value of the correlation coefficient, the features are divided into three levels: weak correlation (less than 0.3), medium correlation (0.3 to 0.5), and strong correlation (greater than 0.5). Construct the feature correlation matrix R = [r ij ] m×m , where r ij It represents the correlation coefficient between the ith feature and the jth feature, and m is the total number of features. The feature clustering method is used to deal with the collinearity problem: when the absolute value of the correlation coefficient between two features is greater than 0.8, the feature with a stronger correlation with chlorophyll concentration is retained.

[0142] The specific implementation of step S04 is to use recursive feature elimination method to select features. This method builds a feature importance evaluation model based on support vector machine, and the kernel function selects radial basis function:

[0143] K(x i , x j ) = exp(-γ||x i -x j || 2 );

[0144] Among them, the kernel parameter γ is set to 0.1, and the penalty parameter C is set to 100. The feature importance score calculation formula is:

[0145]

[0146] Among them, w ij is the weight coefficient of the ith feature on the jth classification surface, and k is the number of classification surfaces. In each iteration, the feature with the lowest importance score is deleted, the model is retrained and the performance is evaluated. The model performance evaluation uses 5-fold cross validation, and the performance indicator is the root mean square error:

[0147]

[0148] Among them, y i is the true value, is the predicted value. The feature combination with the smallest root mean square error on the validation set is selected as the preferred feature set, and the number of features is usually reduced from the initial 20 to 5 to 8.

[0149] The specific implementation method of step S05 is to collect near-infrared spectral data. Use a near-infrared spectrometer with a wavelength range of 900 to 1700 nanometers and a sampling interval of 2 nanometers to obtain 401 wavelength points. Place the fresh tobacco leaf sample flat on the stage to ensure that the leaf surface is in close contact with the detection window. Set a measuring point on each side of the leaf tip, leaf middle, and leaf base, and the distance between the measuring point and the main vein and branch vein is not less than 5 mm. Each measuring point is collected three times, and the interval between two adjacent scans is 10 seconds to eliminate the influence of instrument noise. During the collection process, the ambient temperature is maintained at 25±2 degrees Celsius and the relative humidity is 60%±5% to avoid interference from external light sources. Average the three spectral data of each measuring point:

[0150]

[0151] Among them, S p is the average spectrum of the measurement point p, S pi is the spectrum data collected for the i-th time. Then the arithmetic mean of the six measurement points is calculated as the initial near-infrared spectrum data of the tobacco leaf:

[0152]

[0153] The specific implementation method of step S06 is to pre-process the initial data of the near infrared spectrum. First, a normalization method is used for processing:

[0154]

[0155] Among them, X is the original spectral data, X min and X max are the minimum and maximum values ​​of the spectral data respectively. Then the standard normal variable transformation method is applied:

[0156]

[0157] Among them, X i is the absorbance value at the i-th wavelength point, is the sample average, and n is the number of wavelength points. Finally, the multivariate scattering correction method is used, and the calculation formula is:

[0158] X msc =a+bX;

[0159] Among them, a and b are regression coefficients, which are solved by the least squares method:

[0160]

[0161] Among them, X ref For the reference spectrum, the average spectrum of all samples is usually selected. The dimension of the preprocessed spectral data matrix is ​​m×n, where m is the number of samples and n is the number of wavelength points.

[0162] The specific implementation of step S07 is to extract key features from the preprocessed spectral data. The competitive adaptive reweighting algorithm realizes feature selection through Monte Carlo sampling and exponential decay weights. The sampling number N is set to 500, the weight decay coefficient β is set to 0.95, and the weight update formula for each iteration is:

[0163]

[0164] in, is the weight of the i-th feature at the t-th iteration, p i is the penalty factor. The continuous projection algorithm selects wavelength points by orthogonal projection, and the projection matrix calculation formula is:

[0165] P=IX(X T X) -1 X T ;

[0166] Where I is the unit matrix and X is the data matrix of the selected wavelength points. The sparse representation algorithm uses LASSO regression to achieve feature compression, and the objective function is:

[0167]

[0168] Among them, y is the response variable, β is the regression coefficient, and λ is the regularization parameter, and the optimal value is determined by cross-validation.

[0169] The specific implementation of step S08 is to calculate the interaction relationship between features. First, the image features and spectral features are standardized, and then the feature interaction matrix is ​​calculated:

[0170] M=XX T ;

[0171] Among them, X is the standardized feature matrix. Perform singular value decomposition on the matrix:

[0172] M=U∑V T ;

[0173] Among them, U and V are left and right singular vector matrices, and ∑ is a singular value diagonal matrix. The importance of the feature is determined according to the size of the singular value, and the cumulative contribution rate is calculated:

[0174]

[0175] Among them, σ i For the i-th singular value, select the top k features that make the cumulative contribution rate reach 85%.

[0176] The specific implementation of step S09 is to realize feature fusion based on fusion weight coefficient. The weight coefficient matrix calculation formula is:

[0177]

[0178] Among them, w ij Represents the weight of the i-th feature on the j-th output. The fusion process uses vector multiplication:

[0179] F=X 1 W 1 +X 2 W 2 ;

[0180] Among them, X 1 and X 2 are image feature vector and spectral feature vector respectively, W 1 and W 2 is the corresponding weight matrix.

[0181] The specific implementation of step S10 is to establish a chlorophyll concentration measured data set. The chlorophyll concentration is measured using a spectrophotometer, with the wavelength set to 649 nanometers and 665 nanometers, and the measurement accuracy is ±0.002 absorbance units. The chlorophyll concentration calculation formula is:

[0182]

[0183] Among them, k 1 and k 2 They are 18.16 and 6.63 respectively, V is the volume of the extract, D is the dilution factor, m is the sample mass, and δ is the system error correction term. Each sample was measured three times and the average value was taken as the final result.

[0184] The specific implementation of step S11 is to build a deep learning model. The sequence folding layer reorganizes the one-dimensional input data into a two-dimensional form:

[0185] X reshape =reshape(X,[batch,height,width,channel]);

[0186] Among them, batch is the batch size, which is set to 32, height and width are determined by the feature dimension, and channel is initially 1. The convolution layer uses a 1×1 convolution kernel for feature extraction. The mathematical expression of the convolution operation is:

[0187]

[0188] Among them, W i is the convolution kernel weight, b is the bias term, the number of channels in the first layer is set to 32, and the number of channels in the second layer is 64. The calculation formula of the global average pooling layer is:

[0189]

[0190] Among them, X c is the feature map of the cth channel, H and W are the height and width of the feature map. The weight calculation formula of the channel attention layer is:

[0191] α c =σ(W 2 δ(W 1 F c ));

[0192] Among them, W 1 and W 2 is the weight matrix of the fully connected layer, δ is the ReLU activation function, and σ is the sigmoid function.

[0193] The specific implementation of step S12 is to establish a feature importance scoring matrix. The feature importance is calculated based on the random forest algorithm, and the scoring formula is:

[0194]

[0195] Among them, S i is the importance score of the i-th feature, ΔI ijis the reduction of impurity of feature i in decision tree j, T is the total number of decision trees set to 500, and n is the total number of features. The impurity is calculated using the Gini coefficient:

[0196]

[0197] The reduction in impurity before and after feature splitting is:

[0198]

[0199] The specific implementation of step S13 is to optimize the model parameters. The weight of the channel attention layer is adjusted according to the feature importance matrix, and the adjustment formula is:

[0200] w new =w old ×(1+αS);

[0201] Among them, w old is the original weight, S is the feature importance score, and α is the adjustment coefficient, ranging from 0.1 to 0.5. The Adam optimizer is used to train the model, and the initial value of the learning rate η is set to 0.001, and the momentum parameter β 1 and β 2 They are 0.9 and 0.999 respectively, and the weight update formula is:

[0202] m t =β 1 m t-1 +(1-β 1 ) t ;

[0203]

[0204] Among them, g t is the gradient, ∈ is the numerical stability constant, set to 10 -8 .

[0205] The specific implementation of step S14 is to process the sample to be tested. Repeat the above feature extraction process for the fresh tobacco leaves to be tested, and generate a multi-source feature fusion data matrix of the sample to be tested:

[0206] X test =[x 1 , x 2 , …, x n ];

[0207] Among them, x i is the value of the ith feature, and n is the total number of features. Normalization is performed:

[0208]

[0209] Among them, μ and δ are the mean and standard deviation of the corresponding features in the training set, respectively.

[0210] The specific implementation of step S15 is to predict the chlorophyll concentration. The standardized features are input into the trained model, and the prediction value calculation formula is:

[0211] y pred =f(W n f(W n-1 …f(W 1 x)));

[0212] Among them, W i is the weight matrix of the i-th layer, and f is the activation function. The final output prediction value needs to be denormalized:

[0213] y final =y pred ×δ y +μ y ;

[0214] Among them, δ y and μ y is the standard deviation and mean of the chlorophyll concentration in the training set. The confidence interval of the prediction result is:

[0215] [y final -1.96σ error ,y final +1.96σ error ];

[0216] Among them, σ error is the prediction standard error of the model on the validation set.

[0217] In order to better understand and implement the present invention, the following provides Example 2 of a specific application scenario of the present invention:

[0218] A tobacco research institute conducted a study on the rapid detection of chlorophyll concentration in fresh tobacco leaves. The study selected fresh tobacco leaves of two mainstream varieties, K326 and ZC100, in Yunnan and Guizhou as the research objects, and conducted sampling and analysis at different stages of the growth period. The specific implementation process is as follows:

[0219] Fine control is carried out during the sample collection stage. Sampling is carried out from 9 to 11 am on sunny days. The sampling site is selected in a standardized planting base in the Yunnan Yuxi tobacco area, with a sample area of ​​20 hectares. 50 fresh tobacco leaf samples of K326 and ZC100 varieties are collected in four stages: under-ripe, mature, moderately mature, and mature, for a total of 400 samples. The sampling location selects the 8th to 12th tobacco leaves in the middle of the tobacco plant, and leaves with pests and mechanical damage are removed. The collected samples are immediately placed in an incubator with adjustable temperature and humidity. The temperature is set to 25±1 degrees Celsius, the relative humidity is 65%±3%, and the transportation time is controlled within 30 minutes.

[0220] The image acquisition environment is set up in a standardized manner. A professional dark box with a specification of 80 cm × 60 cm × 60 cm is used, and the inner wall is treated with matte black paint with a reflectivity of less than 3%. The light source is an LED ring light with a color temperature of 5500K, and the illumination is adjusted to 1000±20lux by a professional illuminance meter. Samsung Galaxy S21 Ultra is used for shooting, and its camera has 108 million pixels. The shooting parameters are set to: ISO100, shutter speed 1 / 60 second, aperture value f / 2.2, white balance automatic mode, and resolution set to 4000×3000 pixels. The shooting distance is controlled at 30±0.5 cm by a laser rangefinder. The sample is laid flat on a black background board with anti-reflection treatment, and the level is used to ensure that the verticality error between the leaf surface and the lens is within ±1 degree.

[0221] Image preprocessing adopts a multi-step strategy. First, color space conversion is performed to convert the RGB image into the Ycbcr color space. The conversion matrix is:

[0222]

[0223] Perform adaptive threshold segmentation on the Y component, the initial threshold T 0 Set to 128, the iterative update formula is:

[0224]

[0225] where μ 1 (T k ) and μ 2 (T k ) are the threshold values ​​T k The average gray value of the foreground and background areas obtained by segmentation, when |T k+1 -T k The iteration stops when |<0.5. The segmented image uses morphological processing to optimize the edge, uses a circular structure element with a radius of 3 pixels to perform an opening operation to remove noise, and uses a circular structure element with a radius of 5 pixels to perform a closing operation to fill small holes.

[0226] The feature extraction process includes multiple dimensions. Extract the mean of the R, G, and B channels of the RGB space and the mean of the L, a, and b channels of the Lab space, and calculate the composite eigenvalues ​​such as 2G-RB, R / G, GR, and a / b. Calculate the texture features based on the grayscale co-occurrence matrix, including energy, grayscale average, gradient average, grayscale unevenness, gradient unevenness, correlation, grayscale entropy, gradient entropy, moment of inertia, and inverse moment of difference eigenvalues. The formula for calculating the comprehensive value of texture features is:

[0227]

[0228] The weight coefficient is determined after cross-validation optimization: α 1 =0.32,α 2 =0.28,α 3 =0.24,α 4 =0.16.

[0229] The spectral data collection adopts a precise positioning scheme. The ASD FieldSpec 4 spectrometer is used to collect near-infrared spectral data with a wavelength range of 900 to 1700 nanometers, a sampling interval of 1 nanometer, a spectral resolution of 2 nanometers, and data at 801 wavelength points. Measurement points are set at the tip, middle, and base of each tobacco leaf, and a laser locator is used to ensure that the distance between the measurement point and the main vein and branch vein is greater than 5 mm. Each point is scanned 3 times with a scanning interval of 10 seconds, and the average value is taken as the spectral data of the point. The average spectrum calculation formula is:

[0230]

[0231] The spectral data preprocessing adopts a strategy combining three methods. First, normalization is performed:

[0232]

[0233] Eliminate the dimension effect. Then perform standard normal variable transformation:

[0234]

[0235] Reduce the impact of noise. Finally, use the multivariate scatter correction method:

[0236] X msc =a+bX,

[0237] The regression coefficients are calculated using the least squares method:

[0238]

[0239] The feature selection process adopts a multi-step screening strategy. First, the correlation coefficient between the feature and chlorophyll concentration is calculated using Pearson correlation analysis:

[0240]

[0241] The features with absolute values ​​of correlation coefficients greater than 0.3 are retained as candidate features. Then, the competitive adaptive reweighting algorithm is used for feature extraction, with the sampling times set to 500 times, the weight decay coefficient set to 0.95, and the weight update formula set to:

[0242]

[0243] At the same time, the continuous projection algorithm is used to select the wavelength points, and the projection matrix is ​​calculated as:

[0244] P=IX(X T X) -1 X T .

[0245] Feature fusion uses an adaptive weight method based on matrix decomposition. First, construct the feature interaction matrix:

[0246] M=XX T ,

[0247] Then perform singular value decomposition:

[0248] M=U∑V T .

[0249] Calculate the cumulative contribution rate based on singular values:

[0250]

[0251] Select the feature combination with a cumulative contribution rate of 85%. The fusion weight coefficient is calculated by the following formula:

[0252] W fusion =λ 1 W image +λ 2 W spectral ,

[0253] where λ 1 and λ 2 The initial values ​​are set to 0.45 and 0.55 respectively according to the dynamic adjustment of feature importance.

[0254] The actual chlorophyll concentration is determined by spectrophotometry. Take 0.2 grams of sample, add 25 ml of 95% ethanol, and extract at 4 degrees Celsius in the dark for 24 hours. Use UV-2600 spectrophotometer to measure the absorbance at wavelengths of 649 nanometers and 665 nanometers, repeat the measurement 3 times and take the average value. The formula for calculating chlorophyll concentration is:

[0255]

[0256] Where V is the volume of the extract 0.025 L, D is the dilution factor 1, m is the sample mass 0.2 g, and δ is the systematic error correction term, ranging from 0.001 to 0.01.

[0257] The deep learning model is constructed using a multi-module combination strategy. The sequence folding layer reorganizes the one-dimensional features into a two-dimensional form:

[0258] X reshape =reshape(X,[32,8,8,1]).

[0259] The two convolutional layers use 1×1 convolution kernels, the number of channels is 32 and 64 respectively, and the activation function is ReLU. The calculation formula of the global average pooling layer is:

[0260]

[0261] The channel attention layer weight is calculated as:

[0262] α c =σ(W 2 δ(W 1 F c )).

[0263] The number of hidden units in the long short-term memory layer is 6, and the tanh activation function is used.

[0264] The model training parameters are set as follows:

[0265] 1. The optimizer is Adam, and the initial value of the learning rate is 0.001;

[0266] 2. Batch size 32, training rounds 200;

[0267] 3. Use the early stopping strategy and stop training when the validation set loss does not improve for 10 consecutive rounds;

[0268] 4. The loss function uses the root mean square error:

[0269]

[0270] The experimental results are shown in Table 1-3 below:

[0271] Table 1 Prediction results of chlorophyll concentrations of different varieties at different maturity stages

[0272] variety Mature stage Number of samples Predicted mean (mg / g) Measured mean value (mg / g) Relative error (%) <![CDATA[R 2 ]]> K326 Underripe 50 3.85 3.92 1.79 0.932 K326 Still mature 50 3.18 3.24 1.85 0.928 K326 Suitable for maturity 50 2.45 2.51 2.39 0.921 K326 Mature 50 1.88 1.93 2.59 0.915 ZC100 Underripe 50 3.76 3.82 1.57 0.935 ZC100 Still mature 50 3.12 3.18 1.89 0.927 ZC100 Suitable for maturity 50 2.38 2.44 2.46 0.918 ZC100 Mature 50 1.82 1.87 2.67 0.912

[0273] Table 2 Comparison of model performance of different feature combinations and algorithms

[0274] Feature combination algorithm RMSE MAE <![CDATA[R 2 ]]> Image features BP Neural Network 0.285 0.232 0.845 Image features SVM 0.276 0.225 0.852 Image features CNN-LSTM 0.245 0.198 0.876 Spectral characteristics BP Neural Network 0.256 0.208 0.862 Spectral characteristics SVM 0.248 0.201 0.868 Spectral characteristics CNN-LSTM 0.218 0.175 0.892 Feature Fusion BP Neural Network 0.198 0.162 0.901 Feature Fusion SVM 0.189 0.155 0.908 Feature Fusion CNN-LSTM 0.175 0.138 0.923

[0275] Table 3 Comparison of feature extraction effects of different spectral preprocessing methods

[0276] Preprocessing methods Feature extraction algorithm RMSE <![CDATA[R 2 ]]> Normalization CARS 0.245 0.868 Normalization SPA 0.238 0.875 Normalization LASSO 0.232 0.882 SNV CARS 0.228 0.885 SNV SPA 0.225 0.888 SNV LASSO 0.221 0.892 MSC CARS 0.218 0.895 MSC SPA 0.212 0.898 MSC LASSO 0.208 0.902

[0277] Compared with the traditional technology, this embodiment has the following significant advantages:

[0278] 1. Detection efficiency: The traditional chemical extraction method requires multiple steps such as sample pretreatment, extraction, and determination. The entire process takes 2 to 3 hours and requires a professional laboratory environment and technicians. This method only takes 4 to 6 minutes to complete the detection of a sample, greatly improving the detection efficiency.

[0279] 2. Detection accuracy: The relative error of a portable chlorophyll meter is usually between 8% and 12%, and the relative error of a single image analysis or spectral analysis method is between 4% and 6%. This method controls the relative error within 3% through multi-source feature fusion and deep learning models, significantly improving detection accuracy.

[0280] 3. Universality: This method has achieved good prediction results on two different varieties, K326 and ZC100, indicating that the method has strong universality. The prediction results at different maturity stages maintain a high accuracy, indicating that the method has strong adaptability to sample variation.

[0281] 4. Practicality: This method adopts a standardized data collection process, is simple to operate, has low equipment requirements, and is easy to promote and apply. The deep learning model realizes the automation of feature extraction and prediction without manual intervention, which improves the reliability and stability of detection.

[0282] Summary of technical effects: This embodiment successfully solves the technical problem of rapid and accurate prediction of chlorophyll concentration in fresh tobacco leaves. By innovatively integrating image features and spectral features, a complete set of feature extraction, selection, fusion and prediction methods is established. This method not only ensures the rapidity of detection, but also achieves high-precision prediction, providing strong technical support for quality evaluation and harvesting decisions in tobacco production.

[0283] The following is an example 3 of a specific application scenario of the present invention: Figure 2 As shown, the chlorophyll concentration prediction method in this embodiment 3 includes the following steps:

[0284] S1. Collect fresh tobacco leaves at different stages, take photos to obtain images, and then preprocess and extract image features. The specific steps are as follows:

[0285] S1.1. Take 50 fresh tobacco leaves at four stages: under-ripe, moderately ripe, moderately ripe and mature, number them, and then put them into a dark box in turn and take photos with a smartphone. The parameters remain unchanged during the photo-taking process;

[0286] S2.2. The acquired image is subjected to background removal based on the Ycbcr color space using the built-in color segmentation tool in MATLAB to reduce redundant information in the image.

[0287] S2.3. Using the algorithm, based on the R, G, B, L, a, b color channels, extract the R, G, B, L, a, b color feature values ​​of the corresponding image, and then perform combination operations to obtain the 2G-RB, R / G, GR, a / b feature values

[0288] S2.4. Using the algorithm, based on the gray level co-occurrence matrix, extract the corresponding image energy, gray level average, gradient average, gray level unevenness, gradient unevenness, correlation, gray level entropy, gradient entropy, moment of inertia and inverse moment eigenvalues;

[0289] S2.5. S2.2 and S2.3 above extract 20 kinds of image features in total, and use feature selection method to reduce the dimension of image features. The dimension reduction process specifically includes the following steps:

[0290] S2.5.1. Use the Pearson correlation analysis method to analyze the correlation coefficients between the 20 image features and the measured values ​​of chlorophyll concentration, and then retain the features with medium correlation or above;

[0291] S2.5.2. Using recursive feature elimination, set the number of retained variables to 1-19;

[0292] S2.5.3. The features obtained by Pearson correlation analysis and recursive feature elimination as well as the corresponding measured values ​​of chlorophyll concentration were input into BP model training. The results of different feature training are shown in the figure. Figure 5 As shown, the features with good training model performance are selected as the image feature dataset.

[0293] S2. Collect the near-infrared spectral features of tobacco leaves at 900-1700nm in sequence, with a total of 228 wavelength points, and use the preprocessing method to denoise the original spectrum. Then use three algorithms to extract key near-infrared spectral features. The specific steps are as follows:

[0294] S2.1. Measure the near-infrared spectral characteristic value of each tobacco leaf in turn. Measure 6 points on each tobacco leaf, located at the tip, middle and base of the leaf on both sides of the tobacco leaf. Avoid the main vein and branch veins during measurement. Take the average value of the six points as the near-infrared spectral characteristic value of the tobacco leaf.

[0295] S2.2. The original spectra were preprocessed using three methods: normalization (NOR), standard normal variable transformation (SNV) and multivariate scatter correction (MSC) to remove noise from the original spectra;

[0296] S3.3. The competitive adaptive reweighting algorithm (CARS), the continuous projection algorithm (SPA) and the sparse representation algorithm (Lasso) were used to extract key near-infrared spectral features from the above three pre-processed spectra, and a total of 9 combinations were obtained: NOR-CARS, NOR-SPA, NOR-Lasso, SNV-CARS, SNV-SPA, SNV-Lasso, MSC-CARS, MSC-SPA, MSC-Lasso;

[0297] S4.4. Input the obtained 9 combinations of data and the corresponding measured values ​​of chlorophyll concentration into the BP training model. The training results are as follows: Figure 6 As shown, the best performing data is selected as the near-infrared spectral feature dataset.

[0298] S3. The image feature dataset and the near infrared spectrum dataset are fused using the vector multiplication method to obtain a fused feature dataset. The fusion process is as follows: Figure 4 shown.

[0299] S4. Determine the chlorophyll concentration of fresh tobacco leaf samples using a spectrophotometer. Take 0.2 g of fresh tobacco leaves from the tobacco leaves in turn, soak them in a 95% alcohol solution as the test solution, place the test solution in a dark place away from light until the tobacco leaves turn white, take a portion of the test solution into a colorimetric dish, put it into a spectrophotometer and measure the absorbance at wavelengths of 649 nm and 665 nm. Calculate the chlorophyll concentration according to the formula, and construct a data set of chlorophyll concentration measured values.

[0300] Chlorophyll content = (18.16A 649 +6.63A 665 )×extraction volume×dilution factor / sample fresh weight

[0301] Among them, A 649 and A 665 is the absorbance of the solution at wavelengths of 649 nm and 665 nm. In the present invention, the volume of the extract is 0.025 L, the dilution multiple is 1, and the fresh weight of the sample is 0.2 g.

[0302] S5. Build a CNN and LSTM combination model and add the dot product attention mechanism. The model structure is as follows Figure 3 As shown, the specific steps are:

[0303] S5.1 first consists of the input layer and the sequence folding layer, which are used to adjust the input sequence data to 2D to make it compatible with 2D convolution;

[0304] S5.2. Then add two convolutional layers with a kernel size of 1×1 and channels of 32 and 64 respectively, and then use the ReLU activation function to use the two convolutional layers to locally extract the input features;

[0305] S5.3. Use global average pooling to generate a global description of each channel, then use two fully connected layers to connect to generate attention weights, and finally, normalize the weights through the sigmiod function as channel attention weights, and finally multiply the channel attention weights by the convolution features element by element to implement the attention mechanism, and finally output the feature map;

[0306] S5.4. The weighted feature map is restored to sequence data through the anti-folding layer and input into the LSTM layer for processing. The LSTM hidden unit is set to 6 and the state of the last time step is output;

[0307] S5.5.Finally, set up the fully connected layer and regression layer to map the features extracted by LSTM to the output of the regression task. At this point, the construction of the CNN-LSTM model is completed.

[0308] S6. Train the CNN-LSTM model based on image features, near-infrared spectral features and fusion features, train the CNN, LSTM, BP, and SVM models based on fusion features, and evaluate their performance. The specific steps are as follows:

[0309] S6.1. The CNN-LSTM model was trained and verified using the image feature dataset, near-infrared spectral feature dataset, fusion feature data, and the corresponding chlorophyll concentration measured value dataset. The ratio of the training set to the test set was 7:3. The determination coefficient and root mean square error were used to evaluate the models constructed with three different data. Each model was trained 16 times. The calculation formulas for the determination coefficient R2 and root mean square error RMSE are:

[0310]

[0311] Where n is the number of samples, y i is the measured value, is the predicted value, is the average of the measured values.

[0312] S6.2. Use the fusion feature dataset and the corresponding chlorophyll concentration measured value dataset to train and verify the CNN-LSTM, CNN, LSTM, BP, and SVM models. The ratio of the training set to the test set is 7:3. The determination coefficient, mean absolute error, and root mean square error are used to evaluate the models with different structures in the 5 models. The calculation formula of the above mean absolute error MAE is:

[0313]

[0314] Where n is the number of samples, y i is the measured value, is the predicted value.

[0315] In this embodiment, the model verification results based on 3 different data sets are as follows: Figure 7-Figure 9 As shown in the figure, the verification results show that the R2 of the model based on feature fusion is 0.75833, which is greater than the R2 of the model based on image features and near-infrared spectral features. The RMSE of the model based on fusion features is 0.17482, which is comparable to the RMSE of the model based on image features and lower than the RMSE of the model based on near-infrared spectral features. The results show that using the features of image features and near-infrared spectral features to build a model can improve the prediction ability of the model.

[0316] The model verification results based on fusion feature training are shown in Table 4. The verification results show that the CNN-LSTM model has higher prediction performance, lower prediction error and prediction deviation than the single model (R2=0.74763, MAE=0.13809, RMSE=0.17721).

[0317] Table 4. Evaluation results of CNN-LSTM, CNN, LSTM, BP, and SVM models

[0318] Model <![CDATA[R 2 ]]> MAE RMSE CNN-LSTM 0.74763±0.03136a 0.13809±0.00458c 0.17721±0.01500c CNN 0.71633±0.03395a 0.14368±0.01152bc 0.18846±0.01313bc LSTM 0.63592±0.12233c 0.16019±0.03073ab 0.20638±0.03739ab BP 0.64944±0.09055bc 0.16999±0.02736a 0.22552±0.03665a SVM 0.70320±0.04369ab 0.14715±0.00732bc 0.19627±0.01050bc

[0319] It should be noted that the variables involved in the present invention are explained in detail as shown in Table 5 below.

[0320] Table 5 Variable explanation table

[0321]

[0322]

[0323] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A method for predicting chlorophyll concentration of fresh tobacco leaves based on convolutional neural network and feature fusion, characterized in that: The following steps are involved: Fresh tobacco leaf samples are collected and photographed to obtain fresh tobacco leaf images, and near-infrared spectral data of the fresh tobacco leaf samples are collected to generate a spectral data matrix; the method constructs a feature correlation matrix based on the Pearson correlation analysis method and uses a recursive feature elimination method to screen the optimal feature combination to form an optimal image feature data set, and simultaneously uses a competitive adaptive reweighting algorithm, a continuous projection algorithm, and a sparse representation algorithm to extract features from the preprocessed spectral data matrix to form an optimal spectral feature data set; based on matrix decomposition, a fusion weight coefficient is determined, and a feature interaction matrix of the preferred image feature data set and the preferred spectral feature data set is calculated to generate a multi-source feature fusion data set; A combined model of a convolutional neural network and a long short-term memory network is constructed, wherein the combined model includes a sequence folding layer, a convolution layer, a global average pooling layer, a channel attention layer, an anti-folding layer, a long short-term memory layer, a fully connected layer and a regression layer. The weight of the channel attention layer of the combined model is adjusted according to the importance scoring matrix to generate a chlorophyll concentration prediction model.

2. The method for predicting chlorophyll concentration of fresh tobacco leaves according to claim 1, characterized in that: The steps of collecting fresh tobacco leaf samples and taking photos to obtain fresh tobacco leaf images specifically include collecting 50 fresh tobacco leaf samples at four stages: underripe, moderately ripe, moderately ripe, and mature, taking photos of the fresh tobacco leaf samples in a darkroom environment, performing background segmentation on the fresh tobacco leaf images using a color segmentation method based on a Ycbcr color space, extracting R, G, B, L, a, and b color channel eigenvalues ​​and texture eigenvalues ​​based on a gray-level co-occurrence matrix, and generating an initial image feature data set.

3. The method for predicting chlorophyll concentration of fresh tobacco leaves according to claim 2, characterized in that: The step of collecting near-infrared spectral data of the fresh tobacco leaf sample specifically includes collecting near-infrared spectral data of the fresh tobacco leaf sample in the 900 to 1700 nanometer band, collecting 6 measurement points at the tip, middle and base of the fresh tobacco leaf sample, avoiding the main vein and branch veins, and calculating the average value of the 6 measurement points as the initial near-infrared spectral data.

4. The method for predicting chlorophyll concentration of fresh tobacco leaves according to claim 3, characterized in that: The step of generating a spectral data matrix specifically comprises the following steps: preprocessing the initial near-infrared spectral data by normalization method, standard normal variable transformation method and multivariate scattering correction method to generate a preprocessed spectral data matrix; extracting features from the preprocessed spectral data matrix by competitive adaptive reweighting algorithm, continuous projection algorithm and sparse representation algorithm respectively; selecting the feature combination with the best training performance to generate an optimal spectral feature data set.

5. The method for predicting chlorophyll concentration of fresh tobacco leaves according to claim 4, characterized in that: The step of determining the fusion weight coefficient based on matrix decomposition is specifically to calculate the correlation coefficient between each eigenvalue and the chlorophyll concentration in the initial image feature data set based on the Pearson correlation analysis method, construct a feature correlation matrix, use recursive feature elimination method to process the feature correlation matrix, select the feature combination with the best training performance, and form a preferred image feature data set.

6. The method for predicting chlorophyll concentration of fresh tobacco leaves according to claim 5, characterized in that: The step of generating a multi-source feature fusion data set specifically calculates the feature interaction matrix of the preferred image feature data set and the preferred spectral feature data set, and performs vector multiplication fusion on the preferred image feature data set and the preferred spectral feature data set according to the fusion weight coefficient.

7. The method for predicting chlorophyll concentration of fresh tobacco leaves according to claim 6, characterized in that: The step of generating a chlorophyll concentration prediction model is to use a spectrophotometer to measure the absorbance of the fresh tobacco leaf sample at wavelengths of 649 nanometers and 665 nanometers, generate a chlorophyll concentration measured data set according to the chlorophyll content calculation formula, and establish an importance scoring matrix for each feature in the multi-source feature fusion data set.

8. The method for predicting chlorophyll concentration of fresh tobacco leaves according to claim 7, characterized in that: The step of using the chlorophyll concentration prediction model is specifically to obtain the fresh tobacco leaves to be tested, repeat the operations of the steps described in claims 1 to 6, generate a multi-source feature fusion data set of the fresh tobacco leaves to be tested, input the multi-source feature fusion data set of the fresh tobacco leaves to be tested into the chlorophyll concentration prediction model, and output the chlorophyll concentration prediction value of the fresh tobacco leaves to be tested.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions, and when the program instructions are executed in a computer, the method for predicting the chlorophyll concentration of fresh tobacco leaves according to any one of claims 1 to 8 is used.

10. A fresh tobacco leaf chlorophyll concentration prediction system, characterized in that: The system comprises the computer-readable storage medium as claimed in claim 9, wherein the system is any one of a computer, a server, and a single-chip microcomputer, the computer-readable storage medium is arranged in the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.

Citation Information

Patent Citations

  • Tobacco leaf classification method based on spectrum and machine vision coupling

    CN110705655A

  • Chlorophyll content estimation method based on optical characteristic parameter inversion

    CN114913406A

  • Near-infrared spectrometer and machine vision combined tobacco leaf quality detection method and system

    CN115494008A

  • Denitration system gas concentration prediction method, device, equipment and storage medium

    CN116312861A

  • Wine geographical indication product identification method and system

    CN117351316A

Cited By

  • Method for predicting and analyzing anthocyanin content of red onion based on CIELab color quantization

    CN120369645A

  • Near infrared spectrum-based leaf moisture content and chlorophyll detection method

    CN120992551A

  • Method for rapidly measuring water quality of sewage based on spectrum method

    CN121521783A