Tobacco leaf curing process control method and device based on reinforcement learning
By using a reinforcement learning-based method, environmental parameters of the curing barn and physicochemical properties of tobacco leaves are obtained, a multi-dimensional state vector is constructed, and dynamic control parameters are generated. This solves the problem of poor quality in tobacco curing caused by traditional PID control methods, and achieves efficient and adaptive tobacco curing.
Patent Information
- Application Number
- CN202511361674.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-05-12
- Filing Date
- 2025-09-23
- Publication Date
- 2025-10-28
AI Technical Summary
Traditional PID control methods cannot adaptively adjust during the tobacco curing process, resulting in poor quality of cured tobacco leaves and insufficient dynamic adaptability.
A reinforcement learning-based approach is adopted to construct a multi-dimensional state vector by acquiring environmental parameters and physicochemical indicators of tobacco leaves in the curing chamber. A pre-trained hierarchical reinforcement learning model is then used to generate dynamic control parameters to adjust the heating power, ventilation rate, and humidity compensation of the curing equipment.
It improves the adaptability of the tobacco curing control process, enhances the quality of tobacco curing, reduces reliance on manual experience, lowers the curing failure rate, and optimizes the chemical composition and taste of tobacco.
Smart Images

Figure CN120836784A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of tobacco processing, and in particular to a method and apparatus for controlling the tobacco curing process based on reinforcement learning. Background Technology
[0002] In related technologies, when using the proportional-integral-derivative (PID) control algorithm to cure tobacco leaves, the PID parameters rely on manual adjustment, resulting in weak dynamic adaptability. It cannot adaptively adjust to different environmental parameters in the curing room and tobacco leaves with different physicochemical properties, thus failing to produce high-quality tobacco leaves. Summary of the Invention
[0003] This disclosure provides a method and apparatus for controlling the tobacco curing process based on reinforcement learning, in order to solve the problem of insufficient adaptability of traditional PID control methods for curing tobacco leaves.
[0004] The first aspect of this disclosure proposes a method for controlling the tobacco curing process based on reinforcement learning, the method comprising: The environmental parameters and physicochemical indicators of tobacco leaves inside the curing room were obtained. The environmental parameters included temperature field distribution data, humidity data, and gas concentration data. The physicochemical indicators of tobacco leaves included moisture content data, morphological data, and physiological activity data. Based on environmental parameters and physicochemical indicators of tobacco leaves, a multi-dimensional state vector is constructed. Multi-dimensional state vectors are processed through a pre-trained hierarchical reinforcement learning model to generate dynamic control parameters; Adjust the heating power, ventilation rate, and humidity compensation of the baking equipment according to dynamic control parameters.
[0005] In one embodiment of this disclosure, a pre-trained hierarchical reinforcement learning model is used to process multi-dimensional state vectors and generate dynamic control parameters, including: The surface spectral data of tobacco leaves are matched with the standard quality spectrum to calculate the current quality deviation coefficient. The priority of each optimization objective in the reward function is dynamically adjusted based on the time decay factor and the stage objective weight. The long-term benefits of state-action pairs are evaluated through a dual value network, and the optimal set of control instructions is output as dynamic control parameters.
[0006] In one embodiment of this disclosure, obtaining environmental parameter data and physicochemical indicators of tobacco leaves within the curing barn includes: Collect temperature field distribution data, humidity data, and gas concentration data inside the baking chamber; Multispectral imaging data of surface color parameters of tobacco leaves were collected as morphological data. The offset of the microwave resonant frequency for real-time measurement of the moisture content of tobacco leaves is used as the moisture content data. Physiological activity data of tobacco leaves were analyzed by online near-infrared spectroscopy, including the characteristic absorption intensity peaks of chlorophyll, carotenoids and starch in tobacco leaves.
[0007] In one embodiment of this disclosure, a multi-dimensional state vector is constructed based on environmental parameters and physicochemical indicators of tobacco leaves, including: The baking chamber is divided into at least two three-dimensional grid units, and the temperature gradient vector of each three-dimensional grid unit is calculated. The moisture content difference coefficient between the internal and outer regions of the space formed by tobacco leaf accumulation; A multi-dimensional state vector is constructed based on environmental parameters, temperature gradient vector, and water content difference coefficient.
[0008] In one embodiment of this disclosure, before processing the multi-dimensional state vector through a pre-trained hierarchical reinforcement learning model to generate dynamic control parameters, the method further includes: Based on the LSTM-DDPG architecture, a hierarchical reinforcement learning upper-layer policy network is pre-trained in the digital twin system. The historical state is weighted through a time attention mechanism to output the stage temperature and humidity target curve. The Q value is evaluated through the Critic network and the instructions are optimized through the Actor network. The lower layer training is started after the error reaches the target. The lower-level execution network of the hierarchical reinforcement learning receives target instructions through a distributed PPO cluster. Each node collects the environmental state, the GAE calculates the advantage function, generates fan / heating plate control instructions, aggregates gradients to update the policy, and constrains the action safety boundary to ensure the convergence of execution variance. A shared experience pool stores cross-layer data, priority sampling coordinates learning focus, a gradient alignment strategy is adopted using target consistency loss, and the time sequence is synchronized by hierarchical update cycle. The complexity is gradually increased through learning, and training is terminated when the joint convergence index is reached.
[0009] In one embodiment of this disclosure, it further includes: Constructing a three-dimensional computational fluid dynamics model in a virtual training environment; Simulation strategy parameters are mapped to physical actuators through transfer learning; A safe exploration mechanism is used to limit the feasible domain boundary of the action space.
[0010] A second aspect of this disclosure provides a reinforcement learning-based control device for tobacco curing processes, the device comprising: The acquisition module is used to acquire environmental parameters and physicochemical indicators of tobacco leaves inside the curing room. The environmental parameters include temperature field distribution data, humidity data, and gas concentration data. The physicochemical indicators of tobacco leaves include moisture content data, morphological data, and physiological activity data. The module is used to construct multi-dimensional state vectors based on environmental parameters and physicochemical indicators of tobacco leaves; The generation module is used to process multi-dimensional state vectors through a pre-trained hierarchical reinforcement learning model to generate dynamic control parameters. The control module is used to adjust the heating power, ventilation rate, and humidity compensation of the baking equipment according to dynamic control parameters.
[0011] A third aspect of this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform any of the methods described in the first aspect of this disclosure.
[0012] A fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing computer instructions, characterized in that the computer instructions are used to cause a computer to perform the method described in the first aspect of this disclosure.
[0013] A fifth aspect of this disclosure provides a computer program product characterized by comprising a computer program that, when executed by a processor, implements the method of any one of the first aspects of this disclosure.
[0014] A sixth aspect of this disclosure provides a chip including at least one processor and a communication interface; the communication interface is used to receive signals input to the chip or signals output from the chip, and the processor communicates with the communication interface and implements the method of any one of the first aspects of this disclosure through logic circuits or executing code instructions.
[0015] In summary, the reinforcement learning-based tobacco curing process control method proposed in this disclosure achieves the following beneficial effects: It acquires environmental parameters and physicochemical indicators of tobacco leaves within the curing chamber. Environmental parameters include temperature field distribution data, humidity data, and gas concentration data; physicochemical indicators include moisture content data, morphological data, and physiological activity data, providing a raw data source for the control algorithm of the tobacco curing process. Based on the environmental parameters and physicochemical indicators, a multi-dimensional state vector is constructed, realizing the dynamic coupling representation of environmental parameters and physicochemical indicators, providing a high-precision state perception foundation for the curing process. A pre-trained hierarchical reinforcement learning model processes the multi-dimensional state vector to generate dynamic control parameters, providing control parameters for the curing equipment. Adjusting the heating power, ventilation rate, and humidity compensation of the curing equipment according to the dynamic control parameters improves the adaptability of the tobacco curing control process and enhances the quality of the cured tobacco. It should be understood that the above general description and the following detailed description are merely exemplary and explanatory and do not limit this disclosure. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0017] Figure 1 This is a flowchart of a tobacco curing process control method based on reinforcement learning, according to an embodiment of this disclosure. Figure 2 This is a flowchart illustrating how a pre-trained hierarchical reinforcement learning model processes multi-dimensional state vectors to generate dynamic control parameters, according to an embodiment of this disclosure. Figure 3 This is a flowchart illustrating an embodiment of the present disclosure for acquiring environmental parameter data and physicochemical indicators of tobacco leaves within a curing barn; Figure 4 This is a flowchart illustrating how to construct a multi-dimensional state vector based on environmental parameters and physicochemical indicators of tobacco leaves, according to an embodiment of this disclosure. Figure 5 This is a flowchart illustrating a training hierarchical reinforcement learning model according to an embodiment of the present disclosure; Figure 6 This is a schematic diagram of a tobacco curing process control device based on reinforcement learning, according to an embodiment of this disclosure. Figure 7 This is a block diagram illustrating an electronic device for implementing the reinforcement learning-based tobacco curing process control method of this disclosure, according to an exemplary embodiment. Figure 8 This is a schematic diagram of the chip structure according to an embodiment of the present disclosure. Detailed Implementation
[0018] Embodiments of this disclosure are described in detail below, with examples of embodiments shown in the accompanying drawings, wherein the same or similar reference numerals identify the same or similar originals or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.
[0019] First, let's briefly introduce the relevant terms used in this disclosure: Hierarchical Reinforcement Learning (HRL) is a reinforcement learning method that decomposes complex tasks into multi-level sub-tasks. High-level policies select sub-goals or abstract behaviors, while low-level policies execute specific actions, thus addressing problems such as sparse rewards and complex action spaces.
[0020] Long Short Term Memory - Deep Deterministic Policy Gradient (LSTM-DDPG) is a hybrid model combining LSTM networks and the DDPG algorithm, used to handle reinforcement learning tasks with strong temporal dependencies and continuous action spaces. DDPG is an algorithm that combines deterministic policy gradients and deep Q-networks, belonging to the Actor-Critic framework.
[0021] The Actor-Critic framework is a classic architecture in reinforcement learning that combines policy optimization and value evaluation mechanisms, aiming to achieve more efficient learning through their synergy. The Actor is responsible for generating specific actions based on the current state and interacting with the environment. The Critic is responsible for evaluating the value of the state or state-action pair, guiding the Actor's policy optimization.
[0022] Generalized Advantage Estimation (GAE) is a method in reinforcement learning that estimates the advantage function by incorporating multi-step temporal difference errors (TD Errors). The core idea of GAE is to find a balance between bias and variance, thereby improving the accuracy and stability of the advantage function estimation.
[0023] The Deep Deterministic Policy Gradient (DDPG) algorithm, based on the Actor-Critic framework, is suitable for continuous action spaces and outputs actions through deterministic policies.
[0024] Proximal Policy Optimization (PPO) is a reinforcement learning algorithm based on policy gradients that ensures training stability by limiting the magnitude of policy updates (such as KL divergence constraints).
[0025] Generalized Advantage Estimation (GAE) is a method for estimating the advantage function in policy gradient algorithms, balancing bias and variance to optimize policy updates.
[0026] The method proposed in this disclosure is applied to the control task of tobacco curing process based on reinforcement learning. Its applications are diverse. In the tobacco industry, it can be used for tobacco processing and quality improvement. In tobacco curing plants, it replaces traditional PID control, dynamically adjusting temperature and humidity based on the real-time state of the tobacco leaves (moisture content, enzyme activity) to solve quality fluctuations caused by different batches of tobacco leaves and climatic conditions. It can also be used for refined grading and curing, developing differentiated control strategies for different varieties of tobacco leaves (such as flue-cured and air-cured tobacco) to increase the proportion of high-value tobacco leaves. It reduces reliance on human experience, decreasing curing failure rates (such as green tobacco or charred tobacco); it optimizes the chemical composition of tobacco leaves (such as starch degradation and nicotine retention), improving taste and aroma. It can also be used in agricultural product processing, such as drying tea, coffee, and traditional Chinese medicine. Furthermore, it can be used for tea fixation and drying, dynamically adjusting hot air temperature based on tea leaf morphology (curl, color) to prevent excessive oxidation or charring. Finally, it can be used for coffee bean roasting, combining bean expansion rate and gas concentration (CO2 release) to precisely control the roasting curve and enhance flavor profiles. It can also be used for drying traditional Chinese medicinal herbs, monitoring their moisture content and active ingredients (such as polysaccharides and volatile oils), and preventing high temperatures from damaging their efficacy. It solves the quality instability problem caused by the "one-size-fits-all" temperature control in traditional processes; extends the shelf life of agricultural products, and increases added value. In the food industry, it controls complex processing environments. In meat smoking, it adjusts the smoking time based on meat moisture and smoke concentration to balance flavor and food safety. In fermented food production (such as soy sauce and yogurt), it monitors microbial activity and metabolites (pH value, amino acid content) in real time, dynamically adjusting fermentation environment parameters. In vegetable dehydration, it optimizes the drying rate based on cell structure changes (morphological data) to retain nutrients. It improves product consistency and safety; reduces energy waste (e.g., by avoiding overheating). In industrial manufacturing, in material processing and energy optimization, it can be used for wood drying, adjusting the drying curve based on wood density and moisture content distribution to prevent cracking and deformation. It can also be used in ceramic firing, optimizing the firing cycle by combining the shrinkage rate of the green body with the temperature field distribution inside the kiln. It can also be used for chemical reaction control, monitoring reactant concentrations and exothermic rates, dynamically adjusting cooling systems to prevent explosive polymerization or incomplete reactions. This improves yield, reduces waste costs, and optimizes energy consumption (e.g., reducing ineffective ventilation or heating time). In smart agriculture and environmental protection, in intelligent greenhouse control scenarios, it integrates crop growth data (leaf morphology, transpiration rate) with environmental parameters to achieve precise temperature and humidity control. In environmentally friendly drying equipment, during biomass fuel drying, it optimizes combustion efficiency based on material characteristics (e.g., straw moisture content) to reduce carbon emissions. This promotes the green transformation of agriculture, aligning with the "dual carbon" goal; and improves resource utilization (e.g., water and electricity conservation). In scientific research and education, it can be used for complex system modeling, serving as a typical example of multivariable coupled control, and for control algorithm research (e.g., the fusion of HRL and LSTM).It can also be used as a teaching and experimental platform to build a small baking simulation device, demonstrate the practical application of reinforcement learning in industrial control, accelerate the verification of new technologies, and cultivate interdisciplinary (agriculture + artificial intelligence) talents.
[0027] This technical solution can be widely applied in fields such as agricultural processing, food industry, materials manufacturing, and smart agriculture, and is particularly suitable for scenarios requiring high-precision environmental control and relying on experience-based parameter tuning. Its core value lies in overcoming the limitations of traditional control methods in complex systems through data-driven and adaptive decision-making, thereby promoting the intelligent and green upgrading of industries. The application scenarios are not limited in the embodiments disclosed herein.
[0028] The following section provides a detailed description of the reinforcement learning-based tobacco curing process control method provided in this disclosure, with reference to the accompanying drawings.
[0029] Figure 1 This is a flowchart illustrating a reinforcement learning-based method for controlling the tobacco curing process, according to an embodiment of this disclosure. Figure 1 The embodiment shown illustrates that the reinforcement learning-based tobacco curing process control method includes: Step 101: Obtain environmental parameters and physicochemical indicators of tobacco leaves inside the curing room. Environmental parameters include temperature field distribution data, humidity data, and gas concentration data. Physicochemical indicators of tobacco leaves include moisture content data, morphological data, and physiological activity data.
[0030] In this embodiment, the environmental parameters within the curing chamber include temperature field distribution data, humidity data, and gas concentration data. Temperature field distribution data refers to temperature data measured at different locations within the curing chamber using distributed sensors (such as a thermocouple array), reflecting the spatial uniformity of temperature. For example, the temperature difference between the top and bottom of the curing chamber may affect the dehydration rate of the tobacco leaves. Humidity data refers to the air humidity values within the curing chamber obtained through wet-bulb and dry-bulb sensors or capacitive humidity sensors, directly affecting the rate of moisture evaporation from the tobacco leaves. Gas concentration data includes the concentrations of gases such as oxygen (O2), carbon dioxide (CO2), and carbon monoxide (CO), reflecting combustion efficiency and the biochemical reaction state of the tobacco leaves (such as enzyme activity).
[0031] Optionally, distributed sensor networks (such as ZigBee or LoRa wireless transmission) are used to collect environmental parameters in real time, and online detection devices (such as NIRS probes and cameras) are used to obtain the physicochemical indicators of tobacco leaves. Heterogeneous data (temperature, images, spectra, etc.) are aligned and normalized through timestamps to form a unified time-series dataset.
[0032] Physicochemical indicators of tobacco leaves include data on moisture content, morphology, and physiological activity, reflecting the physiological characteristics of tobacco leaves. Moisture content, measured online using near-infrared spectroscopy or dielectric sensors, is a key indicator for determining the curing stage (e.g., yellowing and color-fixing). Morphological data, acquired through machine vision systems, refers to the morphological characteristics of tobacco leaves (e.g., leaf shrinkage rate and color changes), used to assess the physical state of the leaves (e.g., whether they have reached the "wilt" standard). Physiological activity data, measured using biosensors (e.g., electrochemical enzyme sensors), refers to the intensity of internal biochemical reactions in tobacco leaves, such as polyphenol oxidase activity, reflecting the metabolic state of the leaves. Environmental parameters within the curing chamber are collected using temperature sensors, humidity sensors, and gas concentration detectors; moisture content and morphological data are collected using image or electrical sensors; and physiological activity data are collected using biosensors, enabling real-time and comprehensive monitoring of the curing status.
[0033] Step 102: Construct a multi-dimensional state vector based on environmental parameters and physicochemical indicators of tobacco leaves.
[0034] In this embodiment, the multi-dimensional state vector is formed by converting the heterogeneous data collected in step 101 into a standardized numerical vector that can be processed by the reinforcement learning model, with each dimension corresponding to a key state feature.
[0035] Optionally, features from heterogeneous data can be extracted through feature engineering, and temperature field data can be compressed using Principal Component Analysis (PCA) or an autoencoder (e.g., compressing 100 temperature measurement points into 3 principal components) for dimensionality reduction. Data with different dimensions (e.g., temperature in °C, humidity in %) can be mapped to the [0,1] interval for normalization to avoid model training bias. The extracted features can be temporally aligned, for example, by interpolating or averaging asynchronously acquired data (e.g., temperature every 5 seconds, image every 30 seconds) to ensure temporal consistency of the state vector. This improves the quality of model input and reduces computational complexity.
[0036] Step 103: Process the multi-dimensional state vector through a pre-trained hierarchical reinforcement learning model to generate dynamic control parameters.
[0037] Dynamic control parameters refer to the execution commands that are adjusted in real time, including heating power (kW), ventilation rate (m³ / h), and humidity compensation (g / m³). Using a multi-dimensional state vector as input, a pre-trained hierarchical reinforcement learning model is used for analysis to obtain the dynamic control parameters for adjusting the baking equipment in the baking chamber. This achieves complex task decomposition and dynamic optimization.
[0038] Step 104: Adjust the heating power, ventilation rate, and humidity compensation of the baking equipment according to the dynamic control parameters.
[0039] In this embodiment, heating power is adjusted by using a PID controller or pulse width modulation (PWM) technology to regulate the output power of the heating element or burner. Ventilation rate control is achieved by adjusting the fan speed via a frequency converter, thereby controlling the air circulation speed within the curing chamber (affecting heat distribution and humidity diffusion). Humidity compensation is adjusted by using atomizing nozzles or a steam generator to increase or decrease the water vapor content within the curing chamber, balancing the tobacco dehydration rate with the risk of epidermal hardening. Based on dynamic control parameters, the heating power, ventilation rate, and humidity compensation of the curing equipment are adjusted using the methods described above.
[0040] Optionally, control parameters (such as 80% heating power) can be converted into equipment commands (such as 60% PWM duty cycle), and redundant actuators (such as dual fans) can be used to improve system reliability. After execution, environmental parameters are re-acquired (step 101), forming a "perception-decision-execution-feedback" closed loop. If the actual state deviates from expectations (such as the temperature not reaching the target), model re-inference is triggered (step 103). This ensures that control commands are accurately implemented and enhances the system's anti-interference capability.
[0041] In summary, the reinforcement learning-based tobacco curing process control method proposed in this disclosure acquires environmental parameters and physicochemical indicators of tobacco leaves within the curing chamber. The environmental parameters include temperature field distribution data, humidity data, and gas concentration data, while the physicochemical indicators include moisture content, morphological data, and physiological activity data, providing the original data source for the control algorithm of the tobacco curing process. Based on the environmental parameters and physicochemical indicators, a multi-dimensional state vector is constructed, realizing the dynamic coupling representation of these parameters and indicators, providing a high-precision state perception foundation for the curing process. A pre-trained hierarchical reinforcement learning model processes the multi-dimensional state vector to generate dynamic control parameters, providing control parameters for the curing equipment. Adjusting the heating power, ventilation rate, and humidity compensation of the curing equipment according to the dynamic control parameters improves the adaptability of the tobacco curing control process and enhances the quality of the cured tobacco.
[0042] Figure 2 This is a flowchart illustrating how a pre-trained hierarchical reinforcement learning model processes multi-dimensional state vectors to generate dynamic control parameters, according to an embodiment of this disclosure. Figure 2 Yes Figure 1 Further explanation of step 103, based on Figure 2 The illustrated embodiment includes the following steps: Step 201: Perform feature matching between the surface spectral data of tobacco leaves and the standard quality spectrum to calculate the current quality deviation coefficient.
[0043] In this embodiment, the surface spectral data of tobacco leaves refers to the reflectance spectral curves of tobacco leaf surfaces (typically in the wavelength range of 900-1700 nm) obtained through near-infrared spectroscopy (NIRS) or hyperspectral imaging technology, reflecting the distribution and content of chemical components (such as starch, total sugar, and nicotine) in the tobacco leaves. For example, starch has a characteristic absorption peak near 1200 nm. The standard quality spectrum refers to a reference spectral library established based on historical high-quality tobacco leaf samples, containing ideal spectral characteristics for different curing stages (such as yellowing stage and color-fixing stage). For example, the standard spectrum for the color-fixing stage may require a total sugar content ≥18% and nicotine ≤2.5%. The quality deviation coefficient refers to a numerical index that quantifies the difference between the current tobacco leaf spectrum and the standard spectrum. The surface of the tobacco leaves is scanned online using a hyperspectral camera to generate a spectral image with a spatial resolution of 1 mm². The original spectrum is denoised and baseline corrected to eliminate ambient light interference. Then, feature matching is performed, for example, through dynamic time warping (DTW), to align the time series of the current spectrum with the standard spectrum (for different curing stages) and resolve phase shift issues. Alternatively, principal component analysis (PCA) can be used to extract principal component features of the spectrum (e.g., the first three principal components accounting for 90% of the variance), reducing computational complexity. Finally, the deviation coefficient is calculated, and the corresponding template from the standard spectral library is dynamically selected based on the tobacco leaf grade (e.g., upper-middle leaves, lower leaves). The differences in each band are weighted and the comprehensive deviation coefficient is output (δ∈[0,1], 0 indicating a perfect match). Based on this method, precise quality monitoring can be achieved, detecting anomalies in tobacco leaf chemical composition in real time through spectral feature matching, such as insufficient starch degradation. Quantitative feedback can also be achieved, with the deviation coefficient providing a clear optimization objective for reinforcement learning, such as reducing the δ value.
[0044] Step 202: Based on the time decay factor and the stage target weight, dynamically adjust the priority of each optimization objective in the reward function.
[0045] The time decay factor α(t) refers to a weighting coefficient that decreases with time, such as... This is used to reduce the impact of early actions on long-term rewards and prevent the strategy from focusing excessively on short-term gains. For example, excessively rapid temperature increases in the early stages of baking may lead to tobacco leaf scorching later, requiring a decay factor to suppress the weight of short-term temperature increase rewards. Stage objective weights refer to the dynamic adjustment of the priority of each optimization term in the reward function based on the core objectives of different baking stages (such as yellowing, color-fixing, and drying stages). For example, in the yellowing stage, "starch degradation rate" is given a higher weight. In the color-fixing stage, the emphasis is on "color uniformity" and "moisture gradient control." The reward function is a weighted function that integrates multiple objectives and can be defined as follows: ,in Rewards are given for sub-targets (such as temperature stability and humidity deviation). The stages are weighted accordingly. The curing stages are automatically divided based on tobacco leaf moisture content and color changes (L*a*b*color space). For example, when the moisture content drops from 50% to 40%, the transition from "yellowing stage" to "color setting stage" is triggered. A fuzzy logic controller is used to adjust the process based on real-time status. For example, if local overheating is detected (temperature field variance > 2℃), the weight of "temperature uniformity" is increased. The time decay factor α(t) is negatively correlated with the remaining baking time (e.g., the shorter the remaining time, the faster the decay). Sub-objective reward normalization: [The text abruptly ends here, likely due to an incomplete sentence or missing information.] Mapping to the [-1,1] interval (e.g., deducting 0.1 points for every 1℃ deviation in temperature). R(t) is updated in real-time and input into the reinforcement learning model to guide policy optimization. This step avoids over-optimization of a single objective (e.g., excessive pursuit of energy saving leading to insufficient baking). It improves stage adaptability by using dynamic weights to ensure that the core needs of different baking stages are prioritized.
[0046] Step 203: Evaluate the long-term benefits of the state-action pair through a dual value network, and output the optimal control instruction set as dynamic control parameters.
[0047] In this embodiment, the dual-value network comprises an architecture with two independent Critic networks (Main Critic and Target Critic) to reduce the Q-value overestimation problem. Its update rule is as follows: The Main Critic is used to evaluate the value of the current action, while the Target Critic provides a stable updated target.
[0048] The long-term returns of state-action pairs are evaluated using a dual value network, taking into account a discount factor. The cumulative reward is calculated using the following formula: ,in, The closer it is to 1, the more emphasis is placed on long-term returns.
[0049] The optimal control instruction set refers to a set of continuous control parameters (such as a combination of heating power, ventilation rate, and humidity compensation) generated by the Actor network, and the instruction set with the highest long-term benefit is selected through dual Critic evaluation.
[0050] Optionally, the network architecture is designed as a Main Critic, Target Critic, and Actor network, where the Main Critic takes input state s and action a, and outputs the Q-value Q(s, a); the Target Critic has the same structure as the Main Critic, and its parameters are updated via... Synchronization is performed to improve training stability. Long-term performance optimization utilizes a priority experience replay mechanism based on TD error. Prioritizing samples accelerates learning from key samples. Noise is introduced into the action space to encourage the exploration of non-greedy actions. The Actor network outputs multi-dimensional continuous action vectors, such as heating power and ventilation rate. All candidate actions are evaluated using a dual Critic algorithm, and the selected action is chosen. The highest instruction set. Reduced overestimation and a dual-network design lower Q-value bias, improving policy reliability through a discount factor. The priority experience playback mechanism avoids local optima and provides control parameters for the control of the curing equipment, which helps to ensure the final quality of tobacco leaves.
[0051] Figure 3 This is a flowchart illustrating an embodiment of the present disclosure for acquiring environmental parameter data and physicochemical indicators of tobacco leaves within a curing oven. Figure 3 Yes Figure 1 Further explanation of step 101, based on Figure 3 The illustrated embodiment includes the following steps: Step 301: Collect temperature field distribution data, humidity data, and gas concentration data inside the baking chamber.
[0052] Distributed sensors (such as thermocouple arrays, capacitive hygrometers, and electrochemical gas sensors) are used to monitor the environmental parameters of the curing oven in real time to ensure temperature and humidity uniformity and combustion safety.
[0053] Step 302: Collect multispectral imaging data of the surface color parameters of tobacco leaves as morphological data.
[0054] The reflectance spectrum of tobacco leaves is captured by a multispectral camera (with filters for specific wavelengths), and the color (such as RGB and Lab values) and texture are quantified to assess the physical state of the tobacco leaves (such as the degree of browning).
[0055] Step 303: Measure the microwave resonant frequency offset of the tobacco leaves in real time as the moisture content data.
[0056] Based on the principle of microwave resonant cavity, changes in tobacco leaf moisture alter the dielectric constant, causing a shift in the resonant frequency. The moisture content is then retrieved using the frequency difference (accuracy ±0.5%).
[0057] Step 304: Analyze the physiological activity data of tobacco leaves by online near-infrared spectroscopy. The physiological activity data includes the characteristic absorption intensity peaks of chlorophyll, carotenoids and starch in tobacco leaves.
[0058] Near-infrared light (wavelength 900-1700 nm) penetrates tobacco leaves to detect the intensity of characteristic absorption peaks of chlorophyll (670 nm), starch (1200 nm), carotenoids, etc., and to assess metabolic activity in real time.
[0059] In this embodiment, by collecting multidimensional data on the environment, morphology, moisture, and biochemical data, data on the curing barn and the tobacco leaves themselves are provided for a comprehensive evaluation of the tobacco leaves, and the original data source is provided for the control algorithm of the tobacco curing process.
[0060] Figure 4 This is a flowchart illustrating how a multi-dimensional state vector is constructed based on environmental parameters and physicochemical indicators of tobacco leaves, according to an embodiment of this disclosure. Figure 4 Yes Figure 1 The specific explanation of step 102 is based on Figure 4 The illustrated embodiment includes the following steps: Step 401: Divide the baking chamber into at least two three-dimensional grid units and calculate the temperature gradient vector of each three-dimensional grid unit.
[0061] The three-dimensional space of the baking chamber is divided into several equal-volume cubic grids (e.g., 10 cm × 10 cm × 10 cm). Temperature sensors (e.g., thermocouples) are deployed in each grid to collect local temperature data in real time. For each grid, the temperature difference between it and its adjacent grids is calculated, generating a three-dimensional vector. This three-dimensional vector can characterize the direction and rate of heat transfer. Through spatial discretization and gradient analysis, abnormal temperature regions can be accurately located, and heat distribution can be optimized.
[0062] Step 402: Integrate the moisture content difference coefficient between the internal and outer regions of the space formed by tobacco leaf accumulation.
[0063] For the internal region formed by the tobacco leaf pile, the core moisture content (Wcore) was measured by inserting a microwave resonant probe into the center of the pile. For the outer edge region, an infrared moisture meter was used to scan the surface of the pile to obtain the edge moisture content (Wedge). The moisture content difference coefficient was then calculated. This method quantifies the uneven distribution of moisture. It detects the synchronicity of tobacco leaf dehydration, preventing surface hardening ("crusting") or internal mold growth caused by excessive internal and external moisture gradients. Layered moisture monitoring and differential quantification prevent quality defects caused by uneven tobacco leaf dehydration.
[0064] Step 403: Construct a multi-dimensional state vector based on environmental parameters, temperature gradient vector, and water content difference coefficient.
[0065] Optionally, temperature, humidity, and gas concentration are normalized, and the average magnitude of the gradient vectors of each grid is used to characterize the overall heat flux intensity. The moisture content difference coefficient is weighted and merged with morphological data to generate a comprehensive difference index. Then, environmental parameters, temperature gradient vectors, and moisture content difference coefficients are combined into a matrix to form a multi-dimensional state vector. Multi-source data fusion and feature compression improve model convergence speed and control accuracy. This achieves dynamic coupling characterization of environmental parameters and tobacco leaf physicochemical indicators, providing a high-precision state perception foundation for the curing process.
[0066] Figure 5 This is a flowchart illustrating a training hierarchical reinforcement learning model according to an embodiment of the present disclosure. Figure 5 Yes Figure 1 The specific explanation prior to step 103 is based on Figure 5 The illustrated embodiment includes the following steps: Step 501: Based on the LSTM-DDPG architecture, a hierarchical reinforcement learning upper-layer policy network is pre-trained in the digital twin system. The historical state is weighted through a time attention mechanism to output the stage temperature and humidity target curve. The Q value is evaluated through the Critic network and the instructions are optimized through the Actor network. After the error reaches the target, the lower-layer training is started.
[0067] In this embodiment, a Long Short-Term Memory (LSTM) network and a Deep Deterministic Policy Gradient (DDPG) are combined to process time-series data and output continuous actions. The baking process is simulated in a virtual oven model to avoid damage to real equipment. A time attention mechanism is used to automatically weight key time points of historical states (such as periods of drastic temperature changes) to optimize the generation of the target curve. Layered training is used to trigger the next layer of the network training; once the Q-value error of the Critic evaluation (e.g., MAE < 0.5) meets the target, training of the next layer is initiated.
[0068] Step 502: The lower-level execution network of the hierarchical reinforcement learning receives the target instruction through a distributed PPO cluster. Each node collects the environmental state, the GAE calculates the advantage function, generates fan / heating plate control instructions, aggregates gradients to update the strategy, and constrains the action safety boundary to ensure the convergence of execution variance.
[0069] In this embodiment, a distributed PPO cluster is used to collect environmental data in parallel across multiple nodes (e.g., nodes of 10 image processing units), accelerating experience collection. The advantage function is calculated using GAE. Among these measures, safety boundary constraints prevent equipment overload by limiting the range of actions. After each node calculates the policy gradient, the main network parameters are updated synchronously using the AllReduce algorithm. Distributed parallel computing improves data throughput. The PPO's Clip mechanism limits policy mutations, thereby improving variance convergence speed.
[0070] Step 503: Share the experience pool to store cross-layer data, prioritize sampling to coordinate learning focus, adopt the target consistency loss to align gradients, update the periodic synchronization time sequence in layers, gradually increase complexity through learning, and guide the training to terminate when the joint convergence index is reached.
[0071] In this embodiment, an experience pool is shared by storing interaction data (state, action, reward) between upper and lower layers and supporting cross-layer transfer learning. Sampling probabilities are allocated according to the absolute value of the TD error, prioritizing the replay of high-value experiences. The loss function is defined to align the policy gradients between upper and lower layers: .
[0072] Optionally, a simple training scenario (such as constant temperature) can be initialized, and perturbations can be gradually introduced until the joint metric (such as quality deviation) meets the target.
[0073] The technical solution in this embodiment provides reinforcement learning model support for the tobacco curing control algorithm, reduces reliance on manual parameter tuning, and improves the yield of high-quality tobacco leaves by reducing energy consumption through safety constraints and precise control.
[0074] In one embodiment of this disclosure, the reinforcement learning-based tobacco curing process control method further includes: A three-dimensional computational fluid dynamics (CFD) model is constructed in a virtual training environment. Simulation policy parameters are mapped to physical actuators through transfer learning. A safe exploration mechanism is employed to constrain the feasible domain boundary of the action space. Optionally, a high-fidelity virtual environment is established by simulating airflow, temperature, and humidity distribution within the curing barn using CFD, providing a low-cost training scenario for reinforcement learning and accurately simulating heat transfer and tobacco dehydration dynamics. The policy network parameters trained in the simulation environment are transferred to the physical equipment through domain adaptation (such as adversarial training) and fine-tuning techniques to compensate for real-world differences such as sensor noise and actuator latency, thereby improving policy generalization. Constraining the action space (e.g., heating power ≤ 90%), combined with real-time risk detection and action clipping, prevents dangerous operations such as overheating and overload, ensuring that the training process complies with process safety standards. Through the technical solution of this embodiment, simulation using CFD can reduce trial-and-error costs, accelerate deployment through transfer learning, and ensure the safety of equipment and tobacco leaves through safety mechanisms. These three aspects work together to achieve efficient and reliable curing control.
[0075] This disclosure proposes a reinforcement learning-based control method for tobacco curing. It acquires environmental parameters and physicochemical indicators of tobacco leaves within the curing chamber. Environmental parameters include temperature field distribution data, humidity data, and gas concentration data. Physicochemical indicators include moisture content, morphological data, and physiological activity data, providing the original data source for the control algorithm of the tobacco curing process. Based on the environmental parameters and physicochemical indicators, a multi-dimensional state vector is constructed, realizing the dynamic coupling representation of these parameters and indicators, providing a high-precision state perception foundation for the curing process. A pre-trained hierarchical reinforcement learning model processes the multi-dimensional state vector to generate dynamic control parameters, providing control parameters for the curing equipment. The heating power, ventilation rate, and humidity compensation of the curing equipment are adjusted according to the dynamic control parameters, improving the adaptability of the tobacco curing control process and enhancing the quality of the cured tobacco.
[0076] Corresponding to the methods provided in the above embodiments, this disclosure also provides a tobacco curing process control device based on reinforcement learning. Since the device provided in this disclosure corresponds to the methods provided in the above embodiments, the implementation of the methods is also applicable to the device provided in this embodiment, and will not be described in detail in this embodiment.
[0077] Figure 6 This is a schematic diagram of a reinforcement learning-based tobacco curing process control device 600 according to an embodiment of this disclosure. Figure 6 As shown, the reinforcement learning-based tobacco curing process control device includes: The acquisition module 610 is used to acquire environmental parameters and physicochemical indicators of tobacco leaves inside the curing room. The environmental parameters include temperature field distribution data, humidity data, and gas concentration data. The physicochemical indicators of tobacco leaves include moisture content data, morphological data, and physiological activity data of tobacco leaves. Module 620 is used to construct a multi-dimensional state vector based on environmental parameters and physicochemical indicators of tobacco leaves; The generation module 630 is used to process multi-dimensional state vectors through a pre-trained hierarchical reinforcement learning model to generate dynamic control parameters. The control module 640 is used to adjust the heating power, ventilation rate and humidity compensation of the baking equipment according to dynamic control parameters.
[0078] In some embodiments, the generation module 630 is used for: The surface spectral data of tobacco leaves are matched with the standard quality spectrum to calculate the current quality deviation coefficient. The priority of each optimization objective in the reward function is dynamically adjusted based on the time decay factor and the stage objective weight. The long-term benefits of state-action pairs are evaluated through a dual value network, and the optimal set of control instructions is output as dynamic control parameters.
[0079] In some embodiments, the acquisition module 610 is used for: Collect temperature field distribution data, humidity data, and gas concentration data inside the baking chamber; Multispectral imaging data of surface color parameters of tobacco leaves were collected as morphological data. The offset of the microwave resonant frequency for real-time measurement of the moisture content of tobacco leaves is used as the moisture content data. Physiological activity data of tobacco leaves were analyzed by online near-infrared spectroscopy, including the characteristic absorption intensity peaks of chlorophyll, carotenoids and starch in tobacco leaves.
[0080] In some embodiments, the construction module 620 is used for: The baking chamber is divided into at least two three-dimensional grid units, and the temperature gradient vector of each three-dimensional grid unit is calculated. The moisture content difference coefficient between the internal and outer regions of the space formed by tobacco leaf accumulation; A multi-dimensional state vector is constructed based on environmental parameters, temperature gradient vector, and water content difference coefficient.
[0081] In some embodiments, before the generation module 630 processes the multi-dimensional state vector through a pre-trained hierarchical reinforcement learning model to generate dynamic control parameters, it is further configured to: Based on the LSTM-DDPG architecture, a hierarchical reinforcement learning upper-layer policy network is pre-trained in the digital twin system. The historical state is weighted through a time attention mechanism to output the stage temperature and humidity target curve. The Q value is evaluated through the Critic network and the instructions are optimized through the Actor network. The lower layer training is started after the error reaches the target. The lower-level execution network of the hierarchical reinforcement learning receives target instructions through a distributed PPO cluster. Each node collects the environmental state, the GAE calculates the advantage function, generates fan / heating plate control instructions, aggregates gradients to update the policy, and constrains the action safety boundary to ensure the convergence of execution variance. A shared experience pool stores cross-layer data, priority sampling coordinates learning focus, a gradient alignment strategy is adopted using target consistency loss, and the time sequence is synchronized by hierarchical update cycle. The complexity is gradually increased through learning, and training is terminated when the joint convergence index is reached.
[0082] In some embodiments, the device 6 is further configured to: Constructing a three-dimensional computational fluid dynamics model in a virtual training environment; Simulation strategy parameters are mapped to physical actuators through transfer learning; A safe exploration mechanism is used to limit the feasible domain boundary of the action space.
[0083] In summary, a reinforcement learning-based tobacco curing process control device acquires environmental parameters and physicochemical indicators of tobacco leaves within the curing chamber. Environmental parameters include temperature field distribution data, humidity data, and gas concentration data, while physicochemical indicators include moisture content, morphological data, and physiological activity data. Based on these environmental parameters and physicochemical indicators, a multi-dimensional state vector is constructed. A pre-trained hierarchical reinforcement learning model processes this multi-dimensional state vector to generate dynamic control parameters. These dynamic control parameters are then used to adjust the heating power, ventilation rate, and humidity compensation of the curing equipment.
[0084] This device solves the problem of insufficient adaptability in the traditional PID control method for tobacco curing, improves the adaptability of the tobacco curing control process, and enhances the quality of the cured tobacco.
[0085] The methods and apparatus provided in the embodiments of this disclosure have been described above. To implement the functions of the methods provided in the embodiments of this disclosure, the electronic device may include a hardware structure and software modules, and may implement the above functions in the form of a hardware structure, software modules, or a hardware structure plus software modules. One of the above functions may be executed in the form of a hardware structure, software modules, or a hardware structure plus software modules.
[0086] Figure 7 This is a block diagram illustrating an electronic device 700 for implementing the above-described reinforcement learning-based tobacco curing process control method, according to an exemplary embodiment.
[0087] For example, electronic device 700 can be a mobile phone, computer, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0088] Reference Figure 7 The electronic device 700 may include one or more of the following components: a processing component 702, a memory 704, a power supply component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.
[0089] Processing component 702 typically controls the overall operation of electronic device 700, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 702 may include one or more modules to facilitate interaction between processing component 702 and other components. For example, processing component 702 may include a multimedia module to facilitate interaction between multimedia component 708 and processing component 702.
[0090] Memory 704 is configured to store various types of data to support the operation of electronic device 700. Examples of this data include instructions for any application or method operating on electronic device 700, contact data, phonebook data, messages, pictures, videos, etc. Memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0091] Power supply component 706 provides power to various components of electronic device 700. Power supply component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 700.
[0092] Multimedia component 708 includes a screen that provides an output interface between electronic device 700 and user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 708 includes a front-facing camera and / or a rear-facing camera. When electronic device 700 is in an operating mode, such as a shooting mode or video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0093] Audio component 710 is configured to output and / or input audio signals. For example, audio component 710 includes a microphone (MIC) configured to receive external audio signals when electronic device 700 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 704 or transmitted via communication component 716. In some embodiments, audio component 710 also includes a speaker for outputting audio signals.
[0094] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0095] Sensor assembly 714 includes one or more sensors for providing state assessments of various aspects of electronic device 700. For example, sensor assembly 714 may detect the on / off state of electronic device 700, the relative positioning of components such as the display and keypad of electronic device 700, changes in position of electronic device 700 or a component of electronic device 700, the presence or absence of user contact with electronic device 700, orientation or acceleration / deceleration of electronic device 700, and temperature changes of electronic device 700. Sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 714 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 714 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0096] Communication component 716 is configured to facilitate wired or wireless communication between electronic device 700 and other devices. Electronic device 700 can access wireless networks based on communication standards, such as WiFi, 2G or 3G, 4G LTE, 5G NR (NewRadio), or combinations thereof. In one exemplary embodiment, communication component 716 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 716 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0097] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0098] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, which can be executed by a processor 720 of an electronic device 700 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0099] Embodiments of this disclosure also propose a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the reinforcement learning-based tobacco curing process control method described in the above embodiments of this disclosure.
[0100] Embodiments of this disclosure also propose a computer program product, including a computer program that is executed by a processor to perform the reinforcement learning-based tobacco curing process control method described in the above embodiments of this disclosure.
[0101] Figure 8 This is a schematic diagram of the structure of a chip 800 for implementing the above-described reinforcement learning-based tobacco curing process control method, according to an exemplary embodiment.
[0102] Reference Figure 8 The chip 800 includes at least one communication interface 801 and a processor 802; the communication interface 801 is used to receive signals input to the chip 800 or signals output from the chip 800, and the processor 802 communicates with the communication interface 801 and implements the reinforcement learning-based tobacco curing process control method described in the above embodiments through logic circuits or execution code instructions.
[0103] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0104] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0105] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this disclosure pertain.
[0106] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processing module, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (control method), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic device, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0107] It should be understood that various parts of the embodiments of this disclosure can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0108] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0109] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into a single processing module, or each unit can exist physically separately, or two or more units can be integrated into a single module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The aforementioned storage medium can be a read-only memory, a hard disk, or an optical disk, etc.
[0110] Although embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A method for controlling the tobacco curing process based on reinforcement learning, characterized in that, The method includes: The environmental parameters and physicochemical indicators of tobacco leaves inside the curing room are obtained. The environmental parameters include temperature field distribution data, humidity data, and gas concentration data. The physicochemical indicators of tobacco leaves include moisture content data, morphological data, and physiological activity data of the tobacco leaves. Based on the environmental parameters and the physicochemical indicators of the tobacco leaves, a multi-dimensional state vector is constructed. The multi-dimensional state vector is processed by a pre-trained hierarchical reinforcement learning model to generate dynamic control parameters. The heating power, ventilation rate, and humidity compensation of the baking equipment are adjusted according to the dynamic control parameters.
2. The method according to claim 1, characterized in that, The process of processing the multi-dimensional state vector using a pre-trained hierarchical reinforcement learning model to generate dynamic control parameters includes: The surface spectral data of tobacco leaves are matched with the standard quality spectrum to calculate the current quality deviation coefficient. The priority of each optimization objective in the reward function is dynamically adjusted based on the time decay factor and the stage objective weight. The long-term benefits of state-action pairs are evaluated using a dual value network, and the optimal set of control instructions is output as the dynamic control parameters.
3. The method according to claim 1, characterized in that, The acquisition of environmental parameter data and physicochemical indicators of tobacco leaves inside the curing barn includes: Collect the temperature field distribution data, humidity data, and gas concentration data within the baking chamber; Multispectral imaging data of the surface color parameters of the tobacco leaves were collected as the morphological data. The offset of the microwave resonant frequency for measuring the moisture content of the tobacco leaves in real time is used as the moisture content data. The physiological activity data of the tobacco leaves were analyzed by online near-infrared spectroscopy, and the physiological activity data included the characteristic absorption intensity peaks of chlorophyll, carotenoids and starch of the tobacco leaves.
4. The method according to claim 1, characterized in that, The construction of a multi-dimensional state vector based on the environmental parameters and the physicochemical indicators of the tobacco leaves includes: The baking chamber is divided into at least two three-dimensional grid units, and the temperature gradient vector of each three-dimensional grid unit is calculated; The moisture content difference coefficient between the internal and outer regions of the space formed by tobacco leaf accumulation; The multi-dimensional state vector is constructed based on the environmental parameters, the temperature gradient vector, and the water content difference coefficient.
5. The method according to claim 1, characterized in that, Before processing the multi-dimensional state vector through the pre-trained hierarchical reinforcement learning model to generate dynamic control parameters, the method further includes: Based on the LSTM-DDPG architecture, the upper-layer policy network of the hierarchical reinforcement learning is pre-trained in the digital twin system. The historical state is weighted through the time attention mechanism, and the stage temperature and humidity target curve is output. The Q value is evaluated through the Critic network, and the instructions are optimized through the Actor network. The lower-layer training is started after the error reaches the target. The lower-level execution network of the hierarchical reinforcement learning receives target instructions through a distributed PPO cluster. Each node collects environmental status, GAE calculates the advantage function, generates fan / heating plate control instructions, aggregates gradients to update the strategy, and constrains the action safety boundary to ensure the convergence of execution variance. A shared experience pool stores cross-layer data, priority sampling coordinates learning focus, a gradient alignment strategy is adopted using target consistency loss, and the time sequence is synchronized by hierarchical update cycle. The complexity is gradually increased through learning, and training is terminated when the joint convergence index is reached.
6. The method according to claim 1, characterized in that, Also includes: Constructing a three-dimensional computational fluid dynamics model in a virtual training environment; Simulation strategy parameters are mapped to physical actuators through transfer learning; A safe exploration mechanism is used to limit the feasible domain boundary of the action space.
7. A tobacco leaf curing process control device based on reinforcement learning, characterized in that, The device includes: The acquisition module is used to acquire environmental parameters and physicochemical indicators of tobacco leaves inside the curing room. The environmental parameters include temperature field distribution data, humidity data, and gas concentration data. The physicochemical indicators of tobacco leaves include moisture content data, morphological data, and physiological activity data of the tobacco leaves. A construction module is used to construct a multi-dimensional state vector based on the environmental parameters and the physicochemical indicators of the tobacco leaves; The generation module is used to process the multi-dimensional state vector through a pre-trained hierarchical reinforcement learning model to generate dynamic control parameters. The control module is used to adjust the heating power, ventilation rate, and humidity compensation of the baking equipment according to the dynamic control parameters.
8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-6.
Citation Information
Cited By
Deep learning flue-cured tobacco intelligent control method based on image recognition and moisture
CN121143052A
Coffee baking process accurate control method and system based on digital twinning
CN122308153A