Tea withering intelligent control system and method based on deep learning and multi-modal fusion
By using multimodal feature extraction based on Transformer deep learning and TimesNet temporal modeling, combined with domain adaptive transfer learning and multivariate coupled intelligent decision control, the problems of insufficient feature extraction, inaccurate temporal prediction and poor cross-domain adaptability in tea withering monitoring are solved, and accurate monitoring and efficient control of the tea withering process are achieved.
Patent Information
- Application Number
- CN202511317619.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2026-01-13
AI Technical Summary
Existing tea withering monitoring technologies suffer from insufficient feature extraction, low time-series prediction accuracy, weak cross-domain generalization ability, and crude control strategies, making it difficult to meet the needs of large-scale production and standardized management in the modern tea industry.
By employing multimodal feature extraction based on Transformer deep learning, TimesNet temporal modeling, domain adaptive transfer learning, and multivariate coupled intelligent decision control, we can achieve accurate monitoring and dynamic prediction of the tea withering process.
It enhances the deep fusion capability of feature extraction, improves the accuracy of time series prediction and the cross-domain adaptability of the model, and enhances the accuracy of control strategies. The prediction accuracy is improved by more than 30%, the adaptation speed is improved by more than 5 times, and the control accuracy is improved by more than 40%.
Smart Images

Figure CN121330484A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary fields of artificial intelligence, computer vision, deep learning, and intelligent control, specifically to an intelligent control system and method for tea withering based on the Transformer deep neural network architecture, multimodal data fusion, temporal modeling analysis, and transfer learning techniques. This system achieves precise monitoring, dynamic prediction, and automatic control of the tea withering process by constructing an advanced multimodal feature extraction network, a TimesNet-based temporal prediction model, a domain adaptive transfer learning mechanism, and an intelligent decision control engine, providing technical support for the digital transformation and intelligent upgrading of traditional tea processing technology. Background Technology
[0002] Withering is a core step in tea processing, involving complex biochemical reactions including chlorophyll degradation, polyphenol oxidation, protein hydrolysis, and aroma formation. This process directly determines the quality characteristics, flavor profile, and commercial value of the tea. Traditional withering relies entirely on the sensory experience and manual operation of tea masters, resulting in significant drawbacks such as subjective judgment standards, unstable quality control, low production efficiency, and high labor costs. It is ill-suited to the demands of modern tea industry's large-scale production and standardized management.
[0003] With the rapid development of deep learning, computer vision, multimodal data fusion, and intelligent control technologies, artificial intelligence-based tea withering monitoring and control systems have gradually become a research hotspot. Existing technical solutions can be mainly categorized as follows: The first category is monitoring methods based on a single sensor, such as pure RGB image analysis systems, which can only capture surface color changes and lack in-depth perception of internal physiological and biochemical indicators; the second category is simple environmental parameter monitoring systems, which can only record external conditions such as temperature and humidity and cannot achieve accurate judgment of the withering state; the third category is classification and recognition methods based on traditional machine learning, which have limited feature extraction capabilities and insufficient generalization performance.
[0004] However, tea withering is a highly complex, multi-factor coupled process involving dynamic changes in multiple dimensions, such as leaf morphology, pigment transformation, moisture migration and distribution, and chemical composition evolution. Single-modal data cannot fully characterize the essential features of the withering state. Existing systems generally suffer from the following technical bottlenecks:
[0005] (1) Insufficient feature extraction and lack of deep fusion and collaborative analysis capabilities for multimodal data;
[0006] (2) The time series modeling capability is weak and it is impossible to accurately predict the withering process and the remaining time;
[0007] (3) Poor environmental adaptability; the model has insufficient generalization ability under different seasons, varieties and environmental conditions.
[0008] (4) The control strategy is crude and lacks a precise multivariate collaborative control mechanism. Summary of the Invention
[0009] This invention addresses the core technical bottlenecks in existing tea withering monitoring technologies, such as insufficient depth of multimodal feature fusion, low accuracy of time-series prediction, weak cross-domain generalization ability, and coarse control strategies. It proposes an intelligent tea withering control system and method based on deep learning and multimodal fusion. This system innovatively constructs four core technical modules:
[0010] (1) Multimodal deep feature extraction module, based on YOLOv8 neural network and Random Forest algorithm, realizes high-dimensional feature extraction and representation learning of RGB image and hyperspectral data;
[0011] (2) The time series modeling and analysis module based on the TimesNet architecture achieves accurate prediction of withering completion time, multi-step sequence and withering rate through multi-scale convolution and long short-term memory network;
[0012] (3) Domain adaptive transfer learning module, which adopts fusion strategy knowledge transfer and environmental domain normalization technology, significantly improves the model’s adaptability under different varieties, seasons and environmental conditions;
[0013] (4) Multivariable Coupled Intelligent Decision Control Module: Based on incremental PID control algorithm and multivariable coupled control strategy, it realizes precise coordinated adjustment of key parameters such as temperature and humidity.
[0014] The core technological innovations of this invention are as follows: Firstly, the Transformer deep attention mechanism is introduced into the fusion of multimodal data from tea withering, achieving deep collaborative learning of different modal features through self-attention and cross-attention mechanisms; secondly, a multi-scale temporal prediction network based on TimesNet is innovatively constructed, capable of simultaneously predicting single-point time, multi-step sequences, and change rates, with prediction accuracy improved by more than 30% compared to traditional methods; thirdly, a transfer learning framework for the field of tea withering is proposed for the first time, improving the model's adaptation speed in new environments by more than 5 times through the fusion of policy transfer and domain adaptation techniques; and fourthly, a multivariable coupled PID control algorithm is designed, considering the interactive effects of temperature and humidity, improving control accuracy by more than 40%.
[0015] A multimodal intelligent control system for tea withering, comprising the following claims:
[0016] 1. A multimodal intelligent control system for tea withering, characterized by comprising the following steps:
[0017] Step 1: Acquire RGB images and hyperspectral data of tea leaves during the withering process to construct a multimodal dataset;
[0018] Step 2: Construct a multimodal feature extraction module to extract color features, morphological features, spectral features, temporal features, and texture features respectively;
[0019] Step 3: Implement multimodal feature fusion using the Transformer architecture, and capture the dependencies between features through self-attention and cross-attention mechanisms;
[0020] Step 4: Construct a time-series modeling module based on TimesNet to achieve accurate prediction of withering completion time, multi-step prediction, and withering rate;
[0021] Step 5: Design a transfer learning module, including fusion policy transfer and environmental domain adaptation, to improve the system's adaptability under different conditions;
[0022] Step 6: Construct an intelligent decision control module and use a multivariable coupled PID control algorithm to achieve precise temperature and humidity regulation.
[0023] 2. The system according to claim 1, wherein the multimodal feature extraction module in step 2 specifically comprises:
[0024] A comprehensive feature extraction system based on five categories—color, morphology, spectral, temporal, and texture—was constructed. The system was designed to comprehensively capture the multi-dimensional changes in tea leaves during the withering process, providing a reliable data foundation for subsequent state classification and prediction.
[0025] (1) Color feature extraction: Statistical features are extracted from the RGB, HSV, and LAB color spaces to quantify the color changes of tea leaves during withering.
[0026]
[0027] in Represents the average values of the RGB three channels of the image, where N represents the total number of pixels and R represents the average value of the RGB three channels. i G i B i This represents the RGB value of the i-th pixel. These features can quantitatively describe the process by which tea leaves gradually change from bright green to yellowish-green and then reddish-brown.
[0028] Formula for calculating the proportion of reddened area:
[0029]
[0030] Where 1[·] is the indicator function, RedPixel(i,j) indicates whether the pixel at position (i,j) is red, and W and H are the image width and height, respectively.
[0031] (2) Morphological feature extraction: Geometric morphological parameters of tea leaves are extracted using image processing techniques to reflect the physical changes of the leaves during withering.
[0032]
[0033]
[0034]
[0035] Where A represents the blade area, P represents the perimeter, and L... major and L minor denoted by , a and b represent the lengths of the major and minor axes of the leaf, respectively, while a and b are the semi-major and semi-minor axes of the ellipse fitting. These parameters can quantify morphological changes such as curling and wrinkling of the leaves during the withering process.
[0036] (3) Spectral feature extraction: Hyperspectral imaging technology was used to capture the reflectance spectra of tea leaves at different wavelengths, and feature parameters such as vegetation index were extracted.
[0037]
[0038]
[0039] Where ρ NIR ρ RED ρ represents the reflectance in the near-infrared and red light bands, respectively. 857 ρ 1241 These spectral indices represent reflectance values at specific wavelengths. They can reflect changes in physiological and biochemical parameters such as chlorophyll content and water status.
[0040] (4) Temporal feature extraction: Construct a time series model to capture the dynamic changes of various feature parameters over time during the withering process:
[0041]
[0042] Linear trend: f(t) = at + b
[0043] Quadratic trend: f(t) = at 2 +bt+c
[0044] Where R(t), G(t), and B(t) represent the functions of the RGB three channels changing with time t, and a, b, and c are trend fitting parameters. By analyzing the temporal change trend of the color parameters, the withering process can be predicted and the completion time estimated.
[0045] (5) Texture feature extraction: Texture information on the tea surface is extracted based on the gray-level co-occurrence matrix, reflecting the microstructural changes on the leaf surface:
[0046] Contrast = ∑i,j (ij) 2 P(i,j)
[0047] Entropy = -∑ i,j P(i,j)logP(i,j)
[0048] Where P(i,j) represents the probability of a pixel pair with gray values i and j appearing in the gray-level co-occurrence matrix. Contrast reflects the sharpness and texture depth of the image, while entropy characterizes the complexity and randomness of the texture.
[0049] 3. The system according to claim 1, wherein the Transformer multimodal feature fusion in step 3 specifically comprises:
[0050] An image feature extractor based on YOLOv8, a spectral classifier using the random forest algorithm, and a hierarchical encoder using a multilayer perceptron are employed to process RGB images, hyperspectral data, and withering level information, respectively. The spectral feature encoding uses a three-layer fully connected network structure:
[0051] F spectral =ReLU(Linear3(Dropout(ReLU(Linear2(Dropout(ReLU(Linear1(S)))))))))
[0052] Where S represents the input spectral feature vector, Linear i Let represent the linear transformation of the i-th layer, ReLU be the activation function, and Dropout be the regularization layer. Through layer-by-layer feature abstraction, high-dimensional spectral data is mapped to a representation space consistent with the dimensions of other modal features.
[0053] Multimodal feature fusion employs a weighted cross-attention mechanism:
[0054] F fused =CrossModalFusion([α spectral ·F spectral ;α image ·F image ;α grade ·F grade ])
[0055] Where F image F spectral F grade Representing image, spectral, and hierarchical features respectively, α spectral α image α grade These are learnable attention weights.
[0056] To process timing information, a position encoding mechanism is introduced:
[0057]
[0058]
[0059] Where pos represents the time position, i represents the feature dimension index, and d... model This represents the dimension of the model's hidden layers. Positional encoding enables the model to understand the relative positional relationships of features at different times within the time series data.
[0060] 4. The system according to claim 1, wherein the cross-attention mechanism in step 3 specifically comprises:
[0061] Different modal features are mapped to a unified representation space through a linear transformation:
[0062] P image =W image F image
[0063] P spectral =W spectral F spectral
[0064] P temporal =W temporal F temporal
[0065] Among them W image W spectral W temporal P is a learnable projection matrix. image P spectral P temporal This represents the projected feature representation.
[0066] Formula for calculating multi-head self-attention mechanism:
[0067]
[0068] Where Q, K, and V represent the query, key, and value matrices, respectively, and d k is the dimension of the key vector.
[0069] MultiHead(Q,K,V)=Concat[head1,…,head h W O
[0070] in h represents the number of attention heads, W O This is for outputting the projection matrix.
[0071] Cross-modal feature fusion is achieved through residual connections and layer normalization:
[0072] F cross =LayerNorm(F+MultiHead(F,F,F))
[0073] F final =LayerNorm(F cross +FFN(F cross ))
[0074] Here, FFN represents a feedforward neural network, and LayerNorm is the layer normalization operation. This design ensures that features from different modalities can interact effectively while maintaining training stability.
[0075] 5. The system according to claim 1, wherein the timing modeling module in step 4 specifically comprises:
[0076] A multi-scale time-series prediction model based on the TimesNet architecture is constructed, which includes three prediction branches: withering completion time prediction, multi-step time-series prediction, and withering rate prediction.
[0077] The prediction of withering completion time uses a single-point output network:
[0078] T completion =Softplus(TimePredictor(F fused ))×TimeScale
[0079] Where F fused The fused multimodal features are represented by TimePredictor, a time prediction network, Softplus activation function to ensure positive output, and TimeScale as a time scale factor.
[0080] Multi-step time-series prediction networks are used to predict the withering state at multiple future time steps:
[0081] T multistep =ReLU(MultiStepPredictor(F fused ))
[0082] MultiStepPredictor is a multi-layer LSTM network, and the ReLU activation function is used to process the sequence output.
[0083] Wilting rate prediction networks quantify the speed of the wilting process:
[0084] R withering =Sigmoid(RatePredictor(F fused ))
[0085] RatePredictor is a rate prediction network, and the Sigmoid function limits the output to between 0 and 1, representing the relative intensity of the withering rate.
[0086] To balance multi-task learning, a composite loss function with adaptive weights is adopted:
[0087] L balance =0.4·L variance +0.3·L entropy +0.2·L min_penalty +0.1·L max_penalty
[0088] The various losses correspond to minimizing the prediction variance, entropy regularization, minimum penalty, and maximum constraint, respectively, and multi-objective optimization is achieved through dynamic weight adjustment.
[0089] 6. The system according to claim 1, wherein the transfer learning module in step 5 specifically comprises:
[0090] Design a fusion strategy transfer submodule to realize knowledge transfer between different fusion algorithms. Define four fusion strategies:
[0091]
[0092]
[0093]
[0094]
[0095] Where P img and P spec These represent image and spectral feature representations, respectively, with || denoteing the feature concatenation operation. The four strategies are based on random forest, weighted fusion, multilayer perceptron, and attention mechanism, respectively.
[0096] The environment domain adaptive submodule handles differences in various environmental conditions through parameter standardization:
[0097]
[0098]
[0099]
[0100] Where T, H, and t represent temperature, humidity, and time, respectively. By using minimum-maximum standardization, environmental parameters are mapped to the [0,1] interval, eliminating scale differences between different environmental conditions.
[0101] 7. The system according to claim 1, wherein the intelligent decision control module in step 6 specifically comprises:
[0102] Precise temperature and humidity regulation is achieved using an incremental PID control algorithm.
[0103] u(k)=u(k-1)+K p [e(k)-e(k-1)]+K i e(k)+K d [e(k)-2e(k-1)+e(k-2)]
[0104] Where u(k) is the control output of the kth control cycle, e(k) is the current error, and K p K i K d These are proportional, integral, and differential gains, respectively.
[0105] Design an adaptive parameter adjustment mechanism to dynamically adjust PID parameters based on the error magnitude and rate of change.
[0106] K p (k)=K p0 +ΔK p ·f(Error,ErrorRate)
[0107] Where K p0 Let f(·) be the initial proportional gain, f(·) be the adaptive function, and ΔK be the variable. p Adjust the range of parameters.
[0108] Multivariable coupled control takes into account the interaction between temperature and humidity:
[0109] u temp (k)=PID temp (e temp (k))+C TH ·u humidity (k-1)
[0110] u humidity (k)=PID humidity (e humidity (k))+C HT ·u temp (k-1)
[0111] Where C TH and C HT The coupling coefficient is used to quantify the interaction between temperature and humidity controls.
[0112] 8. The system according to claim 1, characterized in that it includes a wilting progress analyzer module, specifically:
[0113] Construct a multimodal feature fusion decision engine that comprehensively considers image, spectral, and temporal information:
[0114] P fusion =Softmax(αP) image +βP spectral +γP temporal )
[0115] The weighting coefficients are dynamically calculated using an attention mechanism:
[0116] α = Attention image (Context)
[0117] β = Attention spectral (Context)
[0118] γ = Attention temporal (Context)
[0119] The control strategy is generated based on the current wilting state and the target parameters:
[0120] T target (t)=Interpolate(ControlTable,t)
[0121] H target (t)=PhaseAdaptive(Grade,t,Environment)
[0122] Dynamic parameter adjustments are made based on the prediction results and actual deviations:
[0123] ΔT=K T ·(PredictedGrade-ExpectedGrade)
[0124] ΔH=K H (WitheringRate-OptimalRate)
[0125] Wilting progress is assessed by calculating the ratio of the current level to the final target:
[0126]
[0127] 9. The system according to claim 1, characterized in that, the system performance is optimized using a multi-level loss function, including:
[0128] Cross-entropy loss is used for withering rank classification:
[0129]
[0130] Where C represents the number of categories;
[0131] Weighted cross-entropy loss addresses class imbalance:
[0132]
[0133] Where w i Category weights;
[0134] Focal loss mitigates the problem of learning from difficult samples:
[0135]
[0136] Mean squared error loss is used for time series forecasting:
[0137]
[0138] Mean absolute error loss:
[0139]
[0140] Mean absolute percentage error:
[0141]
[0142] 10. The system according to claim 1, characterized in that the system integrates auxiliary modules such as data augmentation, pre-trained models, and performance evaluation:
[0143] Data augmentation employs techniques such as random cropping, color dithering, and Gaussian noise to expand the training samples;
[0144] The pre-trained model is initialized based on the ImageNet-1K dataset to improve feature extraction capabilities;
[0145] Performance evaluation uses metrics such as accuracy, precision, recall, and F1 score to quantify classification performance, and metrics such as MAE, RMSE, and MAPE to evaluate prediction accuracy.
[0146] The system achieves scalability through modular design, supporting the integration of new sensors and algorithm upgrades. Attached Figure Description
[0147] Figure 1 This is the overall system architecture diagram of the present invention, which shows the overall architecture design of the four core modules: multimodal feature extraction, temporal modeling, transfer learning, and intelligent decision control.
[0148] Figure 2 This is a diagram of the YOLOv8 image feature extraction network structure of the present invention, which shows in detail the complete architecture of the backbone network, neck network and detection head;
[0149] Figure 3 This is a detailed diagram of the C2f module structure of the present invention, showing the specific implementation of 1×1 convolution segmentation, Bottleneck residual block and feature concatenation;
[0150] Figure 4 This is a diagram of the Transformer multimodal feature fusion network structure of the present invention, which includes the complete design of encoder-decoder architecture, multi-head attention mechanism and positional encoding;
[0151] Figure 5 This is an architecture diagram of the time series modeling module of the present invention, which shows the design of a multi-branch prediction network based on TimesNet, including completion time prediction, multi-step sequence prediction and withering rate prediction.
[0152] Figure 6 This is a structural diagram of the transfer learning module of the present invention, illustrating a dual transfer learning mechanism that integrates policy transfer and environment domain adaptation;
[0153] Figure 7 This is a graph verifying the accuracy of the time prediction of the present invention, showing the comparison between the system's predicted completion time and the actual completion time. R 2 The accuracy reached 0.982, with a mean absolute error of only 1.2 hours, verifying the high precision performance of the system's prediction model.
[0154] Figure 8 This is a feature distribution comparison diagram of the present invention, showing the distribution comparison of the source domain and the target domain across 10 key feature dimensions. It visualizes the improvement in the consistency of feature distribution before and after transfer learning.
[0155] Figure 9 The image shown is a classification effect diagram of tea withering according to the present invention. It shows the classification results of tea withering level based on YOLOv8, including the image recognition effect of 6 withering levels, which intuitively demonstrates the system's accurate recognition capability for different withering stages.
[0156] Figure 10 The training convergence curve comparison chart for this invention shows the performance comparison of three algorithms—random forest, MLP, and weighted fusion—on the training and validation sets, demonstrating the superiority of the weighted fusion strategy. Detailed Implementation
[0157] The present invention will be further described below with reference to the accompanying drawings and embodiments. However, the present invention can be implemented in many different ways and should not be construed as limited to the embodiments shown; rather, these embodiments provide those skilled in the art with implementation methods that meet applicable legal requirements.
[0158] The following detailed description of a multimodal intelligent control system for tea withering according to the present invention is provided in conjunction with specific embodiments. This embodiment is based on a dataset of white tea from a complete 66-hour withering cycle, including 36,030 RGB images, hyperspectral data samples, and complete environmental parameter records.
[0159] Example 1: Deep Learning-Based Multimodal Intelligent Control System for Tea Withering
[0160] This embodiment constructs a complete system integrating four core modules: multimodal feature extraction, temporal modeling and analysis, transfer learning adaptation, and intelligent decision control. Figure 1 The overall architecture diagram is shown. The system adopts a modular design, with each module capable of operating independently or collaboratively, offering excellent scalability and maintainability.
[0161] Step 1: Multimodal data acquisition and preprocessing
[0162] (1) RGB image data acquisition: A high-resolution industrial camera was used, with a resolution of 1920×1080 pixels, a frame rate of 30fps, and an ISO of 100-400 adaptive adjustment. The image acquisition cycle was set to once every 30 minutes to ensure continuous capture of the withering process. The acquired images were preprocessed with denoising, color correction, and geometric correction to remove the effects of uneven lighting and lens distortion.
[0163] (2) Hyperspectral Data Acquisition: A hyperspectral imager was used to acquire 121-dimensional spectral features in the 397.66-1003.81 nm band. The spectral resolution was set to 2.8 nm, and the spatial resolution reached 0.5 mm × 0.5 mm. Dark current correction, whiteboard correction, and atmospheric correction were performed on the spectral data. Reflectance was calculated and spectral smoothing was performed to effectively remove noise interference.
[0164] (3) Environmental parameter recording: The temperature (20-30℃), humidity (40-85%), wind speed, air pressure and other parameters of the wilting environment are recorded synchronously to establish a database linking environmental conditions and wilting state. All data are synchronized according to a unified timestamp to ensure the spatiotemporal consistency of multimodal data.
[0165] Step 2: Construction of the multimodal feature extraction module (e.g.) Figure 2 , Figure 3 (As shown)
[0166] An image feature extractor is built based on the YOLOv8 deep neural network architecture, and a spectral classifier based on the random forest algorithm is combined to achieve comprehensive extraction of multimodal features.
[0167] (1) Image feature extraction network design:
[0168] An improved YOLOv8 classification network is used as the image feature extractor, such as... Figure 2 As shown, the network input size is set to 640×640 pixels. The backbone network uses a C2f module instead of the traditional C3 module to enhance feature extraction capabilities. The C2f module structure includes 1×1 convolutional segmentation, multiple Bottleneck residual blocks, and final concatenation and fusion, as detailed in the diagram. Figure 3 As shown.
[0169] The formula for calculating the forward propagation of a YOLOv8 network is:
[0170] F (l) =C2f (l) (F (l-1) )
[0171] Where F (l) C2f represents the feature map of the l-th layer. (l) This represents the C2f module of layer l.
[0172] The calculation process of the C2f module is defined as follows:
[0173] F split =Split(Conv 1×1 (F in ))
[0174] F out =Concat([F split [0],Bottleneck n (…Bottleneck1(F split [1]))])
[0175] Where F in For input features, F split denoted as the segmented features, and n as the number of Bottleneck blocks.
[0176] Network parameter settings: initial number of channels is 32, which is multiplied to 1024 layer by layer; SiLU activation function is used instead of ReLU to improve nonlinear expression ability; Dropout probability is set to 0.2 to prevent overfitting; initial learning rate is 0.001, and cosine annealing scheduling strategy is adopted.
[0177] The mathematical expression for the SiLU activation function is:
[0178]
[0179] Where x is the input value, this activation function has better gradient flow characteristics.
[0180] The learning rate cosine annealing scheduling formula is:
[0181]
[0182] Where η t Let η be the learning rate at step t. max =0.001 is the maximum learning rate, η min =0.00001 is the minimum learning rate, T cur T is the current step number. max This represents the total number of steps.
[0183] Color feature extraction process: The RGB image is converted to HSV and LAB color spaces, and 16-dimensional statistical features (mean, standard deviation, skewness, kurtosis, etc.) are calculated in each color space. K-means clustering algorithm (K=5) is used to analyze the distribution of major colors and calculate the proportion of each color region. A threshold segmentation algorithm is used to identify red-colored regions, and the proportion of red-colored area is calculated as an important indicator of the degree of wilting.
[0184] Morphological feature extraction algorithm: First, the Otsu automatic thresholding algorithm is used to segment the tea leaf region, and then morphological opening and closing operations are used to remove noise. The tea leaf boundary is extracted based on a contour detection algorithm, and geometric parameters such as area, perimeter, and major and minor axes are calculated. Ellipse fitting uses the least squares method to calculate eccentricity and ellipticity. Curl is calculated by the ratio of perimeter to convex hull perimeter, quantifying the degree of leaf wrinkling.
[0185] (2) Spectral feature extraction and analysis:
[0186] Principal component analysis (PCA) was used to reduce the dimensionality of the 121-dimensional spectral data, retaining principal components (approximately 30-35 dimensions) with a cumulative variance contribution rate of 95%. Spectral vegetation indices such as Normalized Difference Vegetation Index (NDVI), Normalized Difference Moisture Index (NDWI), and Carotenoid Reflectance Index (CRI) were calculated.
[0187] The mathematical expression for principal component analysis is:
[0188] X centered =X-μ
[0189]
[0190] Cv i =λ i v i
[0191] Where X is the original spectral data matrix, μ is the mean vector, C is the covariance matrix, and v i Let λ be the i-th principal component vector. i For the corresponding eigenvalues.
[0192] The formula for calculating the cumulative variance contribution rate is:
[0193]
[0194] CVR k The cumulative variance contribution rate of the first k principal components is given by d = 121, which is the original feature dimension.
[0195] A spectral classifier based on random forest is constructed, with 100 decision trees, a maximum depth limit of 20 layers, and a minimum number of split samples of 5. A bootstrap sampling strategy is employed to enhance the model's generalization ability, and the Gini impurity criterion is used for feature selection.
[0196] The formula for calculating Gini impurity in random forests is:
[0197]
[0198] Where D is the dataset, p i Let |Y| represent the proportion of category i in dataset D, and |Y| represent the total number of categories.
[0199] The formula for calculating information gain is:
[0200]
[0201] Where A is an attribute, and D v Let A be a subset of samples whose attribute A has a value of v.
[0202] (3) Temporal feature extraction and modeling:
[0203] A sliding window with a time window size of 12 was constructed to extract temporal features such as color change rate and spectral index trends. Linear regression and quadratic polynomial fitting were used to analyze the development trend of the withering process, and the goodness of fit R was calculated. 2 As an indicator of trend stability.
[0204] Goodness of fit R 2 The calculation formula is:
[0205]
[0206] Where y i This is the actual value. For predicted values, is the mean of the actual values, and n is the sample size.
[0207] Step 3: Transformer multimodal feature fusion (e.g.) Figure 4 (See detailed architecture diagram)
[0208] Design an adaptive attention fusion network to achieve deep integration of RGB images, hyperspectral data, and temporal information. The complete Transformer architecture is as follows: Figure 4As shown, it includes the encoder and decoder structure.
[0209] (1) Feature encoding and mapping:
[0210] The spectral feature encoder employs a three-layer fully connected network with 256, 128, and 64 neurons in the hidden layers, all using ReLU activation and Dropout (p=0.3) regularization. This maps the 121-dimensional spectral features to a 64-dimensional unified representation space.
[0211] The forward propagation calculation for a fully connected layer is as follows:
[0212] h (l) =Dropout(ReLU(W (l) h (l-1) +b (l) ))
[0213] Where h (l) For the output of layer l, W (l) and b (l) These are the weight matrix and the bias vector, respectively.
[0214] The mathematical representation of the Dropout operation is:
[0215]
[0216] Where p = 0.3 is the probability of discarding.
[0217] Image features are processed through a global average pooling layer of the YOLOv8 network to output a 512-dimensional feature vector, which is then linearly projected to a 64-dimensional vector. Hierarchical features are represented using one-hot encoding and mapped to a 64-dimensional vector through an embedding layer.
[0218] The formula for calculating global average pooling is:
[0219]
[0220] Where F is the feature map, and H and W are the height and width of the feature map, respectively.
[0221] (2) Implementation of multi-head self-attention mechanism:
[0222] An 8-head attention mechanism is employed, with each attention head having a dimension of 64 / 8 = 8. The query (Q), key (K), and value (V) matrices are obtained through a learnable linear transformation, with a scaling factor of √8. Attention weights are calculated using Softmax normalization to ensure that the sum of the weights is 1.
[0223] The scaled dot product of single-head attention is calculated as follows:
[0224]
[0225] in d k =8 represents the attention head dimension.
[0226] The formula for parallel computation of multi-head attention is:
[0227] MultiHead(Q,K,V)=Concat(head1,…,head h W O
[0228] Where h = 8 is the number of attention heads, W O This is for outputting the projection matrix.
[0229] The Transformer encoder has a stacking depth of 6 layers, with each layer containing two sub-layers: a multi-head self-attention layer and a feedforward neural network. The hidden layer dimension of the feedforward network is expanded to 256, using the GELU activation function. Layer normalization is applied to the input of each sub-layer, and residual connections ensure stable gradient flow.
[0230] The calculation formula for a feedforward neural network is:
[0231] FFN(x)=GELU(xW1+b1)W2+b2
[0232] Where W1∈R 64×256 W2∈R 256×64 This is the weight matrix.
[0233] The mathematical expression for the GELU activation function is:
[0234]
[0235] Where Φ(x) is the cumulative distribution function of the standard normal distribution, and erf is the error function.
[0236] The formula for calculating layer normalization is:
[0237]
[0238] in The mean, Let be the standard deviation, and γ and β be learnable parameters.
[0239] (3) Position encoding and timing processing:
[0240] Sine and cosine position encoding is used to process time-series information, supporting a maximum sequence length of 66 (corresponding to a 66-hour withering cycle). The dimension parameter d_model in the position encoding formula is set to 64, consistent with the feature dimension.
[0241] like Figure 9As shown in the figure, the validation results of the YOLOv8 tea withering classification model demonstrate the system's recognition performance for tea leaves at different withering stages. The figure clearly shows a comparison of the visual features of the six withering stages (levels 1-6): Levels 1-2 are fresh or lightly withered tea leaves, exhibiting a bright green color and flat, plump leaves; Levels 3-4 are moderately withered tea leaves, gradually turning yellowish-green in color, with slightly curled leaves; Levels 5-6 are heavily withered tea leaves, ranging from dark green to brownish-black, with noticeably curled and wrinkled leaves. These validation results demonstrate the effectiveness of the multimodal feature extraction and classification algorithm, proving that the model can accurately identify the state of tea leaves at different withering stages.
[0242] Step 4: Construction of the Time Series Modeling Module Based on TimesNet (e.g.) Figure 5 (As shown in the architecture of the time series prediction model)
[0243] A multi-branch time series prediction network was constructed to achieve accurate prediction of withering completion time, multi-step sequence prediction, and withering rate.
[0244] (1) TimesNet infrastructure design:
[0245] The time series forecasting module uses the TimesNet architecture as its backbone network, which can effectively capture multi-periodic variation patterns in time series. The input sequence length is set to 12 (corresponding to the data of the previous 12 hours), and the forecast sequence length is 4 (for the forecast of the next 4 hours).
[0246] The time-to-image conversion formula of TimesNet is:
[0247] X 2D =Reshape(X) 1D ,(p,T / p))
[0248] Where X 1D ∈R T×d For a one-dimensional time series input, p is the period length, T=12 is the sequence length, and d is the feature dimension.
[0249] The mathematical expression for 2D convolution processing is:
[0250] Y = σ(Conv2D(X) 2D )+b)
[0251] Where Conv2D is a two-dimensional convolution operation, σ is the activation function, and b is the bias term.
[0252] The network consists of four TimesNet blocks, each employing 2D convolutions to handle the periodic transformations of temporal features. The convolutional kernel size is set to 3×3, and the number of channels increases progressively from 64 to 256. Each TimesNet block is followed by adaptive pooling and residual connections to ensure full utilization of feature information.
[0253] (2) Multi-task prediction head design:
[0254] Completion time prediction branch: Employs a fully connected network with two hidden layers (128-64 neurons), outputting a single numerical value representing the estimated completion time. The activation function uses Softplus to ensure a positive output, with the time range limited to 0-66 hours.
[0255] The mathematical expression for the Softplus activation function is:
[0256] Softplus(x) = ln(1+e) x )
[0257] The formula for calculating the time prediction head is:
[0258] T pred =Softplus(FC 64→1 (ReLU(FC 128→64 (ReLU(FC 256→128 (F fused ))))))×66
[0259] Where F fused To fuse features, FC is a fully connected layer, and finally multiplied by 66 for time scaling.
[0260] Multi-step temporal prediction branch: Based on an LSTM network architecture, it contains 3 layers of LSTM units with a hidden state dimension of 128. It outputs a 6-dimensional vector, corresponding to the time predictions for reaching withering levels 1-6. A sequence-to-sequence training strategy is employed, supporting variable-length outputs.
[0261] The calculation formula for an LSTM cell is:
[0262] f t =σ(W f ·[h t-1 ,x t ]+b f )
[0263] i t =σ(W i ·[h t-1 ,x t ]+b i )
[0264]
[0265]
[0266] o t =σ(W o ·[ht-1 ,x t ]+b o )
[0267] h t =o t *tanh(C t )
[0268] Where f t i t o t These are the forget gate, input gate, and output gate, respectively. (C) t In cellular state, h t It is in a hidden state.
[0269] Wilting rate prediction branch: A classification network design is used to output three types of wilting rates (slow, medium, and fast). The network contains a combination of convolutional and fully connected layers, and finally outputs the probability distribution through Softmax.
[0270] (3) Optimization of multi-task loss function:
[0271] Design a comprehensive loss function to balance multiple prediction tasks:
[0272] L total =λ1L time +λ2L sequence +λ3L rate +λ4L balance
[0273] The mean squared error loss is used for time prediction, the sequence loss is used for sequence prediction, and the cross-entropy loss is used for rate prediction. The weight coefficients are determined through grid search: λ1 = 0.4, λ2 = 0.3, λ3 = 0.2, λ4 = 0.1.
[0274] like Figure 7 As shown, the verification results of the system's time prediction accuracy indicate that the predicted completion time and the actual completion time exhibit a highly linear correlation, R0 2 The value reached 0.982, indicating that the model can explain 98.2% of the prediction variance. The mean absolute error (MAE) was only 1.2 hours, with an error rate of less than 2% over a complete 66-hour withering cycle, fully validating the high-precision prediction capability of the TimesNet-based time-series modeling module. The data points in the scatter plot are closely distributed around the perfect prediction line (red dashed line), demonstrating that the system maintains stable prediction accuracy at different withering stages.
[0275] Step 5: Construction and implementation of transfer learning modules (e.g.) Figure 6 (As shown in the transfer learning architecture)
[0276] We design a lightweight fusion strategy transfer framework to enable pre-trained models to adapt quickly to different environmental conditions, thus solving the problem of performance degradation of traditional deep learning models in new scenarios.
[0277] (1) Design of the fusion strategy migration submodule:
[0278] Four different fusion strategy networks were constructed to form a fusion strategy pool: Random Forest Fusion (RF) uses 100 decision trees with a maximum depth of 20 layers and reduces the risk of overfitting through Bagging ensemble; Weighted Fusion (WF) dynamically adjusts the importance of different modalities through a learnable weight matrix; Multilayer Perceptron Fusion (MLP) uses a three-layer fully connected network (12-dimensional input → 64-dimensional hidden → 32-dimensional hidden → 6-dimensional output) and uses the ReLU activation function and batch normalization; Attention Fusion (ATT) calculates the correlation weights between modalities based on the self-attention mechanism.
[0279] The knowledge distillation process employs a soft-label strategy. The teacher model is a complex attention fusion network, while the student model is a lightweight MLP network. The distillation loss function combines cross-entropy loss and KL divergence loss, with a temperature parameter set to 4 to balance the contributions of hard and soft labels.
[0280] (2) Implementation of the environment domain adaptive module:
[0281] Parameter standardization: A standardized mapping function for environmental parameters was established, linearly mapping the temperature range of 20-30℃ to the [0,1] interval, and the humidity range of 40-85% was similarly standardized. Time parameters were normalized by dividing by 66 to ensure the consistency of data distribution across different batches.
[0282] Inter-domain variance minimization: The maximum mean difference (MMD) loss function is used to quantify the difference in feature distributions between the source and target domains, and the inter-domain variance is minimized through an adversarial training strategy. The MMD kernel function uses a Gaussian kernel, and the bandwidth parameter is adaptively determined using a median heuristic.
[0283] Online adaptation mechanism: An incremental learning algorithm is designed to trigger online updates of model parameters when a significant change in environmental conditions is detected. Elastic weight consolidation (EWC) technology is employed to prevent catastrophic forgetting, retaining existing knowledge while learning new environmental features.
[0284] (3) Evaluation and optimization of migration effects:
[0285] Establish a transfer learning performance evaluation system, quantifying the transfer effect through three dimensions: source domain performance retention rate, target domain adaptation speed, and cross-domain generalization ability. Set a performance threshold of 85%; automatically trigger model retraining when the performance falls below the threshold.
[0286] like Figure 8As shown, the feature distribution comparison chart illustrates the distribution comparison between the source and target domains across 10 key feature dimensions. Visual analysis reveals significant differences in the distribution of the source (red) and target (blue) domains across multiple feature dimensions before transfer learning. Particularly in dimensions 1, 4, 5, 7, 8, 9, and 10, the target domain's distribution is shifted to the right and is more widespread than the source domain, reflecting inter-domain differences under different environmental conditions. After domain adaptive transfer learning, the feature distributions of the two domains converge, laying a solid foundation for the model's effective generalization in new environments.
[0287] like Figure 10 As shown in the figure, the comparison of the training convergence curves of the multimodal fusion model demonstrates the accuracy trends of three fusion strategies—random forest, MLP, and weighted fusion—during the training and validation processes. The figure shows that: (1) the training accuracy of all models reaches over 95%, and the validation accuracy is between 80% and 92%, proving the effectiveness of the multimodal fusion method; (2) the weighted fusion strategy exhibits good convergence stability during both training and validation, with relatively small fluctuations in validation accuracy; (3) the random forest method is the most stable on the validation set, maintaining a validation accuracy between 85% and 92%, demonstrating strong generalization ability; (4) the MLP and weighted fusion methods show slight overfitting in the later stages of training, but their overall performance is still superior to the single-modal method. This comparison demonstrates the superiority and convergence stability of the proposed multimodal fusion method.
[0288] Step 6: Construction of Intelligent Decision Control Module
[0289] A hierarchical intelligent control system is constructed, integrating multivariable coupled PID control, decision reasoning engine and safety protection mechanism to achieve precise and coordinated control of environmental parameters such as temperature and humidity.
[0290] (1) Design of a multivariable coupled PID controller:
[0291] Incremental PID algorithm implementation: An incremental PID control algorithm is used to avoid integral saturation. The control period is set to 30 seconds to meet the time scale requirements of the withering process. Initial PID parameter settings: proportional coefficient Kp = 1.2, integral coefficient Ki = 0.3, derivative coefficient Kd = 0.1.
[0292] Adaptive parameter adjustment mechanism: Dynamically adjust PID parameters based on the magnitude and rate of change of system error. When |error|>2.0, increase the proportional coefficient to improve response speed; when the error rate of change>0.5, increase the derivative coefficient to suppress overshoot; when steady-state error persists, appropriately increase the integral coefficient to eliminate steady-state error.
[0293] Temperature and humidity coupled control strategy: Considering the mutual influence between temperature and humidity, a coupled compensation algorithm is designed. The temperature control output is affected by the humidity control output at the previous moment, with a coupling coefficient CTH = 0.15; the humidity control output is affected by the temperature control output at the previous moment, with a coupling coefficient CHT = 0.12. The coupling coefficients are determined through system identification experiments.
[0294] (2) Intelligent Decision Reasoning Engine:
[0295] Multimodal fusion decision-making: The outputs of the YOLOv8 image classifier, random forest spectral classifier, and time-series prediction model are integrated and weighted to generate a comprehensive decision. The weight allocation is dynamically calculated using an attention mechanism and is adaptively adjusted according to the current withering stage and environmental conditions.
[0296] Control strategy generation algorithm: A standard 66-hour wilting control table is established, containing target temperature and humidity setpoints for different time periods. Based on the current wilting level and predicted completion time, a real-time control target is generated using an interpolation algorithm. Considering the influence of environmental factors, a correction coefficient is introduced to adjust the control parameters.
[0297] Dynamic optimization and adjustment mechanism: Real-time monitoring of the deviation between the withering progress and the expected target; when the deviation exceeds a threshold, control parameter adjustments are triggered. Temperature deviation coefficient KT = 0.8, humidity deviation coefficient KH = 1.2, and the optimization algorithm uses particle swarm optimization to find the optimal combination of control parameters.
[0298] (3) System security protection and monitoring:
[0299] Multiple safety protection mechanisms: Temperature control range is set at 18-32℃, and humidity control range is set at 35-90%. If the temperature exceeds the safe range, the corresponding actuator will immediately stop working. Multiple safety measures, including over-temperature protection, over-humidity protection, and power failure protection, are configured to ensure safe system operation.
[0300] Actuator response and feedback: Heater power 1500W, response time <3 seconds; Humidifier capacity 2L / h, response time <5 seconds; Exhaust fan airflow 300m³ / h. 3 / h, start / stop time <2 seconds. All actuators are equipped with position feedback sensors to achieve closed-loop control.
[0301] Data Recording and Traceability System: A complete data recording system is established, recording system status parameters every 10 seconds, including environmental parameters, control outputs, and wilting level predictions. Data storage uses a time-series database, supporting high-concurrency writes and complex query analysis.
[0302] Example 2: System Integration Testing and Performance Verification
[0303] This embodiment deploys the complete system in an actual tea factory environment and conducts a continuous 3-month test. The test dataset contains 180 complete withering batches under different seasons (spring, summer, and autumn), different varieties (Silver Needle, White Peony), and different environmental conditions.
[0304] System integration and deployment: Utilizing an edge computing architecture, the core processing unit employs NVIDIA Jetson Xavier NX, with a GPU computing power of 21 TOPS and 32GB of memory. The data acquisition system automatically collects multimodal data every 30 minutes, with real-time processing latency controlled within 5 seconds. The control system's average response time is 2.3 seconds, meeting real-time control requirements.
[0305] Performance test results: The accuracy rate of wilting level classification reached 93.33%, which is 15% higher than that of traditional manual judgment; the average error of completion time prediction was 3.2 hours, and the prediction accuracy was 85.6%; the temperature control accuracy was ±0.5℃, and the humidity control accuracy was ±2.0%; the system's continuous operation stability was 99.2%, and no major failures occurred.
[0306] Economic Benefit Analysis: The system reduces manual monitoring costs by 60%, improves withering quality consistency by 35%, shortens the withering cycle by an average of 4 hours, and raises the tea quality rating by one grade. The system's investment payback period is approximately 18 months, demonstrating significant economic benefits and application value.
[0307] The above description is merely a preferred embodiment of a multimodal intelligent control system for tea withering. The scope of protection for this system is not limited to the above embodiments; all technical solutions falling within this conceptual framework are within the scope of protection of this invention. It should be noted that for those skilled in the art, any improvements and variations made without departing from the principles of this invention should also be considered within the scope of protection of this invention. This invention can be widely applied to tea processing, intelligent manufacturing of agricultural products, automation of the food industry, and other related fields, possessing significant practical value and promising prospects for promotion.
Claims
1. A smart control system for tea withering based on multimodal feature fusion and temporal prediction, characterized in that, Includes the following steps: Step 1: Acquire RGB images, hyperspectral data, and environmental parameters during the tea withering process; Step 2: Construct a comprehensive system that includes a multimodal feature extraction module, a temporal modeling module, a transfer learning module, and an intelligent decision control module; Step 3: Construct a cross-modal fusion framework based on the Transformer attention mechanism in the multimodal feature extraction module, and achieve deep integration of RGB images, hyperspectral data and temporal information through adaptive weight learning; Step 4: In the temporal prediction stage, an adaptive attention fusion network is introduced to achieve the optimal combination of image, spectral, and hierarchical information through a dynamic modality weight allocation mechanism. Step 5: Employ fusion strategy transfer learning techniques to achieve cross-environment adaptation and improve model generalization ability; Step 6: Achieve coordinated optimization of temperature and humidity parameters through multivariable coupling control theory and establish an intelligent decision engine.
2. The system according to claim 1, characterized in that, The specific steps for obtaining multimodal data in step 1 are as follows: This system utilizes existing technologies to acquire RGB images and hyperspectral data during the tea withering process, while simultaneously recording temporal information and environmental parameters. RGB images capture changes in the tea's appearance, including color, shape, and texture. Hyperspectral data, containing reflectance information in the 397.66-1003.81 nm wavelength band, is used to extract spectral characteristics of internal tissue changes in the tea leaves. Environmental parameter records include physical parameters such as temperature and humidity, used to establish a correlation model between the withering process and environmental conditions. The system monitors the complete withering cycle (0-66 hours), constructing a comprehensive time-series dataset. Finally, through comprehensive processing of multimodal data, it achieves accurate judgment and prediction of the withering state.
3. The system according to claim 1, characterized in that, The multimodal feature extraction module in step 2 specifically includes: A comprehensive feature extraction system based on five categories—color, morphology, spectral, temporal, and texture—was constructed. The system was designed to comprehensively capture the multi-dimensional changes in tea leaves during the withering process, providing a reliable data foundation for subsequent state classification and prediction. Color feature extraction: Statistical features are extracted from the RGB, HSV, and LAB color spaces to quantify the color changes of tea leaves during withering. in Represents the average values of the RGB three channels of the image, where N represents the total number of pixels and R represents the average value of the RGB three channels. i G i B i This represents the RGB value of the i-th pixel. These features can quantitatively describe the process by which tea leaves gradually change from bright green to yellowish-green and then reddish-brown. Formula for calculating the proportion of reddened area: Where R area The red-color area ratio represents the proportion of the red-color area, where W and H are the width and height of the image, respectively. 1[RedPixel(i,j)] is an indicator function, with a value of 1 when pixel (i,j) is determined to be a red-color area, and 0 otherwise. The red-color area ratio is an important visual indicator of the degree of withering, gradually increasing as withering progresses. Morphological feature extraction: After obtaining the tea leaf contour through image segmentation, geometric features and curl features are extracted to quantify the morphological changes of the tea leaves. Where A is the area of the tea leaf, P is the perimeter, and L is the length of the tea leaf. major and L minor , where a and b are the lengths of the principal axis and minor axis, respectively, and the major axis and minor axis of the ellipse fitting. These features can quantify the morphological change process of tea leaves from flat to curled, and are physical indicators of withering depth. Spectral feature extraction: Vegetation index and moisture index are extracted from hyperspectral data to reflect changes in the internal components of tea leaves. Where ρ NIR and ρ RED ρ represents the reflectance in the near-infrared and red light bands, respectively. 857 and ρ 1241 These represent the reflectance at wavelengths of 857 nm and 1241 nm, respectively. NDVI (Normalized Difference Vegetation Index) reflects changes in chlorophyll content, while NDWI (Normalized Difference Moisture Index) reflects changes in the moisture content of tea leaves. These indices can penetrate the surface and directly reflect the internal composition. Temporal Feature Extraction: Quantifying the temporal characteristics of the withering process through rate of change analysis and trend fitting. Linear trend: f(t) = at + b Quadratic trend: f(t) = at 2 +bt+c Where t represents the withering time, R(t), G(t), and B(t) represent the RGB mean values at time t, and a, b, and c are fitting parameters. The color change rate reflects the withering speed, and the linear and quadratic trends describe the overall pattern of the withering process, which helps predict subsequent development. Texture feature extraction: Texture features are extracted using the Gray-Level Co-occurrence Matrix (GLCM). Where P(i,j) represents the probability that a pair of pixels with gray values i and j co-occur at a specific distance and direction. Contrast reflects the roughness of the tea surface, and entropy describes the complexity of the texture; both increase as withering deepens.
4. The system according to claim 1, characterized in that, The timing modeling module in step 2 specifically includes: The time-series prediction system comprises three core components: a single-point prediction model, a sequential time prediction model, and a trajectory prediction model, enabling multi-level prediction of the dynamic changes in the withering process. This module, through in-depth mining of time-series data, can predict the withering completion time and future trajectory, providing a basis for intelligent control decisions. The single-point prediction model: Accurate prediction of withering completion time is achieved based on a multi-modal feature fusion network. F spectral =ReLU(Linear3(Dropout(ReLU(Linear2(Dropout(ReLU(Linear1(S)))))))) where S is the input spectral feature (121 dimensions), Linear... i This represents the i-th linear layer, where ReLU is the activation function and Dropout is a random deactivation layer (to prevent overfitting). This encoder converts the raw spectral data into a 256-dimensional feature vector. F fused =CrossModalFusion([a spectral ·F spectral ;a image ·F image ;a grade ·F grade ]) Where α spectral α image α grade The attention weights for spectral features, image features, and hierarchical features are respectively, F. spectral F image F grade These are the corresponding feature vectors. Through an adaptive attention mechanism, the model can dynamically adjust the importance of each modality according to different scenarios. Time series prediction model: Based on the Transformer architecture, it processes time series data and uses positional encoding to provide sequence information. Where pos represents the sequence position, i represents the dimension index, and d... model =384 represents the model dimension. Sine positional encoding enables the model to distinguish features at different time points, maintaining the integrity of temporal information.
5. The system according to claim 1, characterized in that, The cross-modal fusion framework of the Transformer attention mechanism in step 3 is as follows: This module achieves deep fusion of features from different modalities through an adaptive attention mechanism, fully leveraging the complementarity of RGB images, hyperspectral data, and temporal information to generate a comprehensive feature representation with strong expressive power. Feature projection mapping process: P image =W image F image P spectral =W spectral F spectral P temporal =W temporal F temporal Among them W image W spectral W temporal These are projection matrices for image, spectral, and temporal features, respectively, which map features of different dimensions to a unified feature space. Multi-head self-attention calculation: Where Q, K, and V are the query, key, and value matrices, respectively, and d k This is the scaling factor. By calculating the similarity between the query and the key, different weights are assigned to different elements in the value matrix, thereby enabling the learning of correlations between features. MultiHead(Q,K,V)=Concat[head1,…,head h ]W O head i This represents the output of the i-th attention head, where h = 8 is the number of attention heads, and W... O This is the output projection matrix. Multi-head attention allows the model to simultaneously focus on feature relationships in different subspaces, enriching feature representations. Residual connectivity and layer normalization: F cross =LayerNorm(F+MultiHead(F,F,F)) F final =LayerNorm(F cross +FFN(F cross )) LayerNorm is a layer normalization algorithm, and FFN is a feedforward neural network. The combination of residual connections and layer normalization can effectively alleviate the gradient problem in deep network training and improve network stability.
6. The system according to claim 1, characterized in that, The adaptive attention fusion network in step 4: This network is specifically designed for time-series forecasting tasks. Through a dynamic modal weight allocation mechanism, it achieves the optimal combination of image, spectral, and hierarchical information. The network includes a time prediction head, a multi-step time prediction module, and a trajectory prediction module, enabling it to comprehensively predict the future development of the withering process. Predicted completion time of withering: T completion =Softplus(TimePredictor(F fused ))×TimeScale Where TimePredictor is the time predictor, Softplus is the activation function to ensure that the output is positive, and TimeScale=60.0 is the time scaling factor to map the output to a reasonable hour range. Multi-step time prediction: T multistep =ReLU(MultiStepPredictor(F fused )) The MultiStepPredictor is a multi-step time predictor that outputs a 6-dimensional vector, corresponding to the time to reach levels 1-6. Withering rate prediction: R withering =Sigmoid(RatePredictor(F fused )) RatePredictor is the rate predictor, and Sigmoid compresses the output to the range of 0-1, representing the relative rate of withering. Modal balance loss function: L balance <0.4·L variance +0.3·L entropy +0.2·L min_penalty +0.1·L max_penalty Where L variance For variance loss, L entropy For entropy loss, L min_penalty and L max_penalty These are the minimum and maximum penalty terms, respectively. This loss function prevents any one mode from becoming overly dominant and promotes a balanced integration of modes.
7. The system according to claim 1, characterized in that, The specific implementation of the fusion strategy transfer learning technique in step 5 is as follows: This technology enables cross-domain transfer of knowledge from pre-trained models through a lightweight fusion network, maintaining the feature extraction capabilities of the original model and making adaptive adjustments only to the fusion layer, thereby significantly improving transfer efficiency and generalization ability. Four fusion strategies are implemented: Where P img and P spec These represent the predicted probabilities of the image and spectral models (6 dimensions each), respectively. || denotes the vector concatenation operation, w is the learnable weight vector, MLP stands for Multilayer Perceptron, and Attention is the attention mechanism. Random Forest fusion has strong anti-overfitting ability, weighted fusion has high computational efficiency, MLP fusion has strong expressive power, and attention fusion has strong adaptive ability. Environmental parameter normalization: Where T is temperature, H is humidity, and t is time (hours), T min =20℃, T max =30℃, H min =40%, H max =85%. Parameter normalization ensures the consistency of data distribution under different environmental conditions, providing a unified input feature for the model.
8. The system according to claim 1, characterized in that, The multivariable coupling control theory in step 6 specifically refers to: The temperature and humidity controller adopts a four-layer hierarchical architecture design: physical hardware layer, hardware abstraction layer, control algorithm layer, and safety monitoring layer. Through coupling compensation, it achieves coordinated optimization of temperature and humidity parameters, effectively solving the mutual interference problem in traditional single-variable control. Incremental PID control algorithm: u(k)=u(k-1)+K p [e(k)-e(k-1)]+K i e(k)+K d [e(k)-2e(k-1)+e(k-2)] Where u(k) is the control output at step k, e(k) is the control error, and K p K i K d These are the proportional, integral, and derivative coefficients, respectively. Incremental PID avoids integral saturation problems and has better stability and anti-interference capabilities. Adaptive parameter adjustment: K p (k)=K p0 +ΔK p ·f(Error,ErrorRate) Where K p0 The basic proportionality coefficient, ΔK p The adjustment range is given by f, which is an adaptive function based on the error and the rate of change of the error. This mechanism can dynamically adjust the PID parameters according to the system response characteristics, thereby improving control accuracy. Temperature and humidity coupled control: u temp (k)=PID temp (e temp (k))+C TH ·u humidity (k-1) u humidity (k)=PID humidity (e humidity (k))+C HT ·u temp (k-1) Where C TH and C HT Here, represents the coupling compensation coefficients, and represents the effects of humidity on temperature and temperature on humidity, respectively. Through coupling compensation, the system can coordinate the mutual influence between temperature and humidity, improving control stability.
9. The system according to claim 1, characterized in that, The intelligent decision engine is specifically implemented as follows: The intelligent wilting decision engine adopts a modular design, integrating three classification models: YOLOv8 image classifier, random forest spectral classifier, and Transformer multimodal fusion model. It achieves high-precision wilting state judgment and optimization control decision generation through multi-source information fusion. Multimodal feature fusion classification: P fusion =Softmax(αP image +βP spectral +γP temporal ) Where P image P spectral P temporal These represent the predicted probabilities for image, spectral, and temporal models, respectively, with α, β, and γ being the corresponding attention weights, obtained through learning. a=Attention image (Context) β=Attention spectral (Context) c = Attention temporal (Context) Context refers to the current decision context, which includes the withering stage, environmental conditions, and historical information. Temperature and humidity control strategy generation: T target (t)=Interpolate(ControlTable,t) H target (t)=PhaseAdaptive(Grade,t,Environment) Where ControlTable is the standard 66-hour withering control table, t is the withering time, Grade is the current withering level, and Environment is the environmental condition. Based on the standard process and the current state, the system generates optimal temperature and humidity control parameters. Dynamic adjustment of control parameters: ΔT=K T ·(PredictedGrade-ExpectedGrade) ΔH=K H ·(WitheringRate-OptimalRate) Where K T and K H Here, represents the adaptive gain coefficient, PredictedGrade is the predicted grade, ExpectedGrade is the expected grade, WitheringRate is the withering rate, and OptimalRate is the optimal rate. This mechanism can automatically adjust temperature and humidity parameters according to the progress of withering, ensuring that the withering process proceeds as expected. Wilting progress assessment: CurrentGrade represents the current wilting level (1-6), and MaxGrade=6 is the highest level. Progress assessment provides the decision-making system with a quantitative indicator of wilting completion, facilitating control and adjustment.
10. A smart control system for tea withering based on multimodal feature fusion and temporal prediction, characterized in that: This system adopts a multi-stage closed-loop control architecture, comprising four core components: a multimodal feature extraction module, a temporal modeling module, a transfer learning module, and an intelligent decision control module. It achieves deep fusion of RGB images, hyperspectral data, and temporal information through an innovative Transformer attention mechanism, dynamically allocates modal weights using an adaptive attention mechanism, and combines fusion strategy transfer learning technology to achieve cross-environmental adaptation of the model. Based on multivariate coupled control theory, it achieves precise coordinated control of temperature and humidity. By identifying the withering state of tea leaves, predicting the completion time and future development trajectory, the system generates precise temperature and humidity control strategies, effectively improving the quality and consistency of tea withering. On 180 validation samples, the system achieves a classification accuracy of 93.33%, control precision of ±0.5℃ and ±2.0%, and a system response time of 2.5 seconds, reducing the withering judgment time from 30 minutes to 3 seconds, demonstrating significant technical effectiveness and industrial application value.
Citation Information
Cited By
Multi-modal fusion evaluation method and system for grading green tea
CN121545141A
Medical data processing method and product based on multi-stage transfer learning and multi-modal data collaborative fusion
CN122265749A