Multi-layer instruction adaptive control method, device, system and storage medium

Through a multi-layer instruction adaptive control method, real-time control instructions are generated using Kalman filtering, deep learning and reinforcement learning networks, which solves the problem that traditional tea baking equipment cannot adapt to environmental changes, achieves the stability and consistency of tea baking effects, and improves tea quality and production efficiency.

CN120295149BActive Publication Date: 2025-09-23QUANZHOU INST OF EQUIP MFG +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510793856.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-23
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

Traditional tea roasting equipment has difficulty adapting to real-time changes in environmental parameters during the roasting process, resulting in unstable and inconsistent roasting results.

Method used

A multi-layer instruction adaptive control method is adopted to obtain the current baking parameters, generate real-time control instructions using Kalman filtering, deep learning prediction model, model predictive control algorithm and reinforcement learning network, dynamically adjust the baking parameters and achieve adaptive control.

Benefits of technology

Ensure the stability and consistency of the baking effect, improve the consistency of tea quality and taste, and reduce manual intervention and operation difficulty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295149B_ABST
    Figure CN120295149B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of tea roasting technology, and provides a multi-layer instruction adaptive control method, device, system, and storage medium. The method comprises obtaining a first roasting parameter at the current moment in a roasting machine; using the first roasting parameter to predict a second roasting parameter at the next moment; using the first roasting parameter and the second roasting parameter to generate a first real-time control instruction and a third roasting parameter within a preset time period; using the first real-time control instruction and the third roasting parameter, applying a reinforcement learning network to generate a second real-time control instruction; and using the first real-time control instruction and the second real-time control instruction to control a roasting parameter adjustment mechanism. The method adaptively determines a control instruction based on the first roasting parameter, the second roasting parameter, and the third roasting parameter, and adaptively and dynamically controls the roasting parameter adjustment mechanism in the roasting machine to achieve an optimal roasting effect, ensure the stability and consistency of the roasting effect, and ensure consistent quality across different tea types and roasting batches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tea baking, and in particular to a multi-layer instruction adaptive control method, device, system and storage medium. Background Art

[0002] Tea roasting plays a vital role in the entire tea production process, and the precision of its control directly determines the quality and taste of the tea.

[0003] Traditional tea roasting equipment relies primarily on preset temperature control curves and relies heavily on manual adjustments during the roasting process. However, in the actual roasting process, environmental parameters are constantly changing. Traditional tea roasting equipment struggles to adapt to these real-time changes, and cannot guarantee stable and consistent roasting results. For example, temperature fluctuations can cause the tea leaves to burn or become under-roasted; unstable air quality can introduce odors or affect the tea's color.

[0004] Based on this, there is an urgent need to provide a multi-layer instruction adaptive control method that can be applied to tea roasting control. Summary of the Invention

[0005] The present invention provides a multi-layer instruction adaptive control method, device, system and storage medium to solve the defects in the prior art.

[0006] The present invention provides a multi-layer instruction adaptive control method, comprising:

[0007] Get the first baking parameter in the baking machine at the current moment;

[0008] Based on the first roasting parameter, predicting a second roasting parameter at a next moment in the roasting machine;

[0009] generating a first real-time control instruction and a third baking parameter within a preset time period based on the first baking parameter and the second baking parameter;

[0010] Based on the first real-time control instruction and the third baking parameter, applying a reinforcement learning network to generate a second real-time control instruction;

[0011] Based on the first real-time control instruction and the second real-time control instruction, a baking parameter adjustment mechanism in the baking machine is controlled.

[0012] According to a multi-layer instruction adaptive control method provided by the present invention, predicting a second baking parameter at a next moment in the baking machine based on the first baking parameter includes:

[0013] Performing Kalman filtering on the first baking parameter to obtain a filtering result;

[0014] Inputting the filtering result into a deep learning prediction model to obtain the second baking parameter output by the deep learning prediction model;

[0015] The deep learning prediction model is trained based on historical baking parameters and baking effect indicators.

[0016] According to a multi-layer instruction adaptive control method provided by the present invention, generating a first real-time control instruction and a third baking parameter within a preset time period based on the first baking parameter and the second baking parameter, comprising:

[0017] Establishing a system model of the tea roasting process based on a model predictive control algorithm, and predicting the third roasting parameter based on the system model and the first roasting parameter;

[0018] The first real-time control instruction is solved based on the third baking parameter and the optimization target.

[0019] According to a multi-layer instruction adaptive control method provided by the present invention, the method generates a second real-time control instruction based on the first real-time control instruction and the third baking parameter by applying a reinforcement learning network, including:

[0020] extracting control decision information from the first real-time control instruction;

[0021] Inputting the control decision information and the third baking parameter into the reinforcement learning network to obtain the second real-time control instruction output by the reinforcement learning network based on the state space, the action space, and the reward function;

[0022] The reinforcement learning network learns the action with the highest value based on the state and reward value of the tea roasting process. The state space includes roasting parameters, the action space includes control instructions, and the reward function is determined based on the roasting effect.

[0023] According to a multi-layer instruction adaptive control method provided by the present invention, controlling a baking parameter adjustment mechanism in the baking machine based on the first real-time control instruction and the second real-time control instruction includes:

[0024] generating a target control instruction based on the first real-time control instruction and the second real-time control instruction;

[0025] Based on the target control instruction, a closed-loop control algorithm is applied to control the baking parameter adjustment mechanism.

[0026] According to a multi-layer instruction adaptive control method provided by the present invention, the first baking parameter includes at least one of humidity, temperature, carbon dioxide concentration, carbon monoxide concentration and air quality.

[0027] The present invention also provides a multi-layer instruction adaptive control device, comprising:

[0028] A baking parameter acquisition module, used to obtain the first baking parameter in the baking machine at the current moment;

[0029] a baking parameter prediction module, configured to predict a second baking parameter at a next moment in the baking machine based on the first baking parameter;

[0030] a first instruction generating module, configured to generate a first real-time control instruction and a third baking parameter within a preset time period based on the first baking parameter and the second baking parameter;

[0031] a second instruction generating module, configured to generate a second real-time control instruction based on the first real-time control instruction and the third baking parameter by applying a reinforcement learning network;

[0032] The baking control module is configured to control a baking parameter adjustment mechanism in the baking machine based on the first real-time control instruction and the second real-time control instruction.

[0033] The present invention also provides a multi-layer instruction adaptive control system, comprising: a controller, a baking machine, and a baking parameter acquisition device, wherein the baking parameter acquisition device is connected to the controller;

[0034] The roasting machine is used to load tea leaves and roast the tea leaves;

[0035] The baking parameter acquisition device is provided in the baking machine and is used to acquire the first baking parameter in real time;

[0036] The controller is used to receive the first baking parameter and execute the above-mentioned multi-layer instruction adaptive control method.

[0037] According to a multi-layer instruction adaptive control system provided by the present invention, the baking parameter acquisition device includes at least one of a temperature sensor, a humidity sensor, a carbon monoxide sensor, a carbon dioxide sensor, and an air quality sensor.

[0038] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the multi-layer instruction adaptive control methods described above.

[0039] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the multi-layer instruction adaptive control methods described above.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] The multi-layered instruction adaptive control method, device, system, and storage medium provided by the present invention first obtain a first roasting parameter in a roaster at the current moment; then use the first roasting parameter to predict a second roasting parameter in the roaster at the next moment; then use the first roasting parameter and the second roasting parameter to generate a first real-time control instruction and a third roasting parameter within a preset time period; then use the first real-time control instruction and the third roasting parameter to generate a second real-time control instruction using a reinforcement learning network; and finally use the first real-time control instruction and the second real-time control instruction to control a roasting parameter adjustment mechanism in the roaster. This method can adaptively determine a control instruction based on the first roasting parameter, the second roasting parameter, and the third roasting parameter, and then use the control instruction to implement adaptive dynamic control of the roasting parameter adjustment mechanism in the roaster to achieve an optimal roasting effect, thereby ensuring the stability and consistency of the roasting effect and ensuring consistent quality across different tea types and roasting batches. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on the drawings in the following description without any creative work.

[0043] Figure 1 It is a flow chart of the multi-layer instruction adaptive control method provided by the present invention;

[0044] Figure 2 It is a structural diagram of the multi-layer instruction adaptive control device provided by the present invention;

[0045] Figure 3 This is one of the structural diagrams of the multi-layer instruction adaptive control system provided by the present invention;

[0046] Figure 4 This is the second structural diagram of the multi-layer instruction adaptive control system provided by the present invention;

[0047] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0048] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0049] During the actual tea roasting process, environmental parameters are constantly changing. Traditional tea roasting equipment relies heavily on manual adjustments, making it difficult to adapt to these real-time changes in environmental parameters during the roasting process. Consequently, the stability and consistency of the roasting results cannot be guaranteed. Therefore, an embodiment of the present invention provides a tea roasting control method.

[0050] Figure 1 FIG. 1 is a flow chart of a multi-layer instruction adaptive control method provided in an embodiment of the present invention, as shown in FIG. Figure 1 As shown, the method includes:

[0051] S1, obtaining the first baking parameter in the baking machine at the current moment;

[0052] S2, predicting a second roasting parameter at a next moment in the roasting machine based on the first roasting parameter;

[0053] S3, generating a first real-time control instruction and a third baking parameter within a preset time period based on the first baking parameter and the second baking parameter;

[0054] S4, applying a reinforcement learning network based on the first real-time control instruction and the third baking parameter to generate a second real-time control instruction;

[0055] S5: Controlling a baking parameter adjustment mechanism in the baking machine based on the first real-time control instruction and the second real-time control instruction.

[0056] Specifically, the multi-layer instruction adaptive control method provided in the embodiment of the present invention is executed by a controller, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., which is not specifically limited here.

[0057] First, step S1 is performed to obtain a first baking parameter within the roaster at the current moment. The first baking parameter refers to a baking parameter within the roaster at the current moment, and may be an environmental parameter, such as at least one of humidity, temperature, carbon dioxide concentration, carbon monoxide concentration, and air quality. The first baking parameter may be acquired by a baking parameter acquisition device, which may include at least one of a humidity sensor, a temperature sensor, a carbon dioxide sensor, a carbon monoxide sensor, and an air quality sensor.

[0058] The baking parameter acquisition device can communicate with the controller via a wired connection (such as RS485) to ensure fast transmission speed and high reliability. The position of the baking parameter acquisition device is optimized to ensure that the first baking parameters of each key position can be collected.

[0059] Temperature sensors can be high-precision, high-temperature-resistant thermocouples, which can accurately measure the temperature inside the roaster even in high-temperature environments. Temperature sensors are placed in various locations within the roaster, including the top, bottom, and sides, to ensure comprehensive monitoring of the temperature distribution within the roaster.

[0060] Humidity sensors with high accuracy, fast response, and suitability for high-temperature environments should be selected. Ideally, the humidity sensor should be strategically placed inside the roaster, near the tea placement area and in a location with high air circulation, such as the central area of ​​the roaster. This allows for accurate monitoring of humidity changes at various locations during the roasting process, providing real-time humidity parameters. This helps ensure more precise control of the roasting environment, ensuring tea is roasted under optimal humidity conditions, and ultimately improving tea quality and taste.

[0061] Carbon dioxide (CO) sensors, carbon monoxide (CO) sensors, and air quality sensors should be high-precision and have a wide range. This allows them to accurately detect changes in gas and solid particulate matter (carbon content) concentrations during the roasting process and provide accurate air quality parameters. These sensors can be placed near the roaster's air inlet and outlet to monitor CO2 and CO concentrations, as well as air quality. Air quality sensors can include PM2.5 and PM10 sensors.

[0062] Then, step S2 is executed, using the first roasting parameters to predict the second roasting parameters at the next moment in the roaster. For example, the first roasting parameters can be directly input into a deep learning prediction model, which then outputs the second roasting parameters. This deep learning prediction model can be a Gated Recurrent Unit (GRU) neural network, which uses update and reset gates to control the flow of information. The update gate determines how much information in the current state is derived from the previous state, while the reset gate determines how much information is forgotten. The GRU neural network has a relatively simple structure, fast training speed, and is capable of effectively processing time series data.

[0063] Deep learning prediction models can be trained using historical baking parameters and baking performance indicators. These historical baking parameters can include temperature, humidity, carbon monoxide concentration, carbon dioxide concentration, air quality, and other parameters. After collecting historical baking parameters, they can also be preprocessed, including data cleaning and normalization.

[0064] Roasting performance indicators are supervised learning labels that indicate the actual results achieved by tea leaves under historical roasting parameters. Roasting performance indicators are a set of parameters used to measure the quality and condition of tea leaves during the roasting process. They comprehensively reflect the appearance, aroma, taste, and chemical properties of roasted tea leaves and are a key indicator for evaluating tea roasting control performance.

[0065] In an embodiment of the present invention, the baking effect indicators may include tea color, tea aroma, tea moisture content, tea morphology, tea taste, etc. The tea color can be evaluated by image analysis or comparison with a standard color sample to see whether the color of the tea meets expectations. The tea aroma can be evaluated using a professional odor sensor or by a professional tea taster to provide the intensity and quality grade of the aroma. The tea moisture content can be measured using a moisture meter. The tea morphology can be determined by observing whether the tea leaves are intact after baking and whether they are broken or deformed. Image recognition technology can also be used to automatically detect the integrity and degree of breakage of the tea leaves. By training a tea morphology evaluation model, tea leaves of different morphologies can be identified and corresponding evaluation results can be given. The taste of the tea can be indirectly evaluated by analyzing the chemical composition of the tea leaves using an automated detection method. For example, the content of tea polyphenols, caffeine, and other ingredients in the tea leaves can be measured, which are closely related to the taste of the tea leaves.

[0066] The historical baking parameters are divided into training data and test data, and the training data is input into the initial prediction model for training. The structural parameters of the initial prediction model are adjusted using the backpropagation algorithm and an optimization algorithm (such as the Adam optimizer) to minimize the prediction error.

[0067] The trained deep learning prediction model is then used to predict the test data and compared with the baking effect indicators corresponding to the test data to evaluate the performance of the deep learning prediction model. Metrics such as root mean square error (RMSE) and mean absolute error (MAE) can be used to measure the prediction accuracy of the deep learning prediction model.

[0068] Next, step S3 is executed to generate a first real-time control instruction and a third roasting parameter for a preset future time period using the first and second roasting parameters. This can be achieved using a Model Predictive Control (MPC) algorithm, an advanced control algorithm that uses models to predict future states. The MPC algorithm establishes a system model of the tea roasting process to predict the system state for a period of time in the future, i.e., the third roasting parameter for a preset future time period. The algorithm then generates an optimal control sequence, i.e., the first real-time control instruction, based on the optimization objective function. Thus, the MPC algorithm can predict the changing trend of the third roasting parameter for a preset future time period based on the first and second roasting parameters and generate the optimal first real-time control instruction.

[0069] The length of the preset time period can be set as needed and is not specifically limited here.

[0070] Finally, step S4 is executed. A reinforcement learning network (Deep Q-Network, DQN) is applied using the first real-time control instruction and the third baking parameter to generate a second real-time control instruction. The DQN is introduced to further optimize and adjust the first real-time control instruction. The DQN input can include the key parameters in the first real-time control instruction and the third baking parameter, or the third baking parameter and control decision information determined by the key parameters. The output is the second real-time control instruction. It is understood that if the first real-time control instruction controls the temperature to a set value, then the set value is the key parameter. The control decision information is the range of the temperature control device's opening level, such as high, medium, and low, rather than the specific value of the key parameter.

[0071] During DQN training, the initial network outputs control instructions, which are then used to control the roasting parameter adjustment mechanism, achieving intelligent control of the tea roasting process. Subsequently, the strategy is continuously adjusted based on the reward value of the tea roasting effect to learn more optimal control instructions, further optimizing DQN's performance under different roasting parameters and improving its adaptability and intelligence.

[0072] In the embodiment of the present invention, the baking parameter adjustment mechanism may include a temperature adjustment device, a fan, an air quality regulator and other devices. The temperature adjustment device may be a heater, and the air quality regulator may be an air purification system.

[0073] The temperature inside the roaster can be adjusted by the temperature regulating device, the humidity, carbon monoxide concentration and carbon dioxide concentration inside the roaster can be adjusted by the fan, and the air quality inside the roaster can be adjusted by the air quality regulator.

[0074] Finally, step S5 is executed, where the first and second real-time control instructions are used to control the baking parameter adjustment mechanism within the roaster. Here, the first and second real-time control instructions can be integrated using methods such as weighted averaging and fuzzy logic to obtain a target control instruction. This target control instruction is then used to control the baking parameter adjustment mechanism and adjust the first baking parameter.

[0075] The multi-layered instruction adaptive control method provided in an embodiment of the present invention first obtains a first roasting parameter at the current moment in a roaster; then uses the first roasting parameter to predict a second roasting parameter at the next moment in the roaster; thereafter uses the first roasting parameter and the second roasting parameter to generate a first real-time control instruction and a third roasting parameter within a preset time period; thereafter uses the first real-time control instruction and the third roasting parameter to generate a second real-time control instruction using a reinforcement learning network; and finally uses the first real-time control instruction and the second real-time control instruction to control a roasting parameter adjustment mechanism in the roaster. This method can adaptively determine a control instruction based on the first roasting parameter, the second roasting parameter, and the third roasting parameter, and then use the control instruction to implement adaptive dynamic control of the roasting parameter adjustment mechanism in the roaster to achieve the optimal roasting effect, thereby ensuring the stability and consistency of the roasting effect and ensuring consistent quality across different tea types and roasting batches.

[0076] Based on the above embodiment, predicting the second roasting parameter at the next moment in the roasting machine based on the first roasting parameter includes:

[0077] Performing Kalman filtering on the first baking parameter to obtain a filtering result;

[0078] Inputting the filtering result into a deep learning prediction model to obtain the second baking parameter output by the deep learning prediction model;

[0079] The deep learning prediction model is trained based on historical baking parameters and baking effect indicators.

[0080] Specifically, when predicting the second baking parameter at the next moment in the roaster, a Kalman filter can be first performed on the first baking parameter to obtain a filtering result. Here, the process of performing Kalman filtering on the first baking parameter is the process of smoothing the first baking parameter using the Kalman filter algorithm.

[0081] The Kalman filter algorithm is a recursive algorithm based on linear minimum variance estimation. It estimates and predicts the system state to eliminate noise interference in the first roasting parameters. During the tea roasting process, the Kalman filter algorithm can process the first roasting parameters in real time, obtaining more accurate and stable first roasting parameters.

[0082] During the Kalman filter process, the roasting parameters can be used as the system state. First, the current system state is predicted based on the tea roasting process model and the system state at the previous moment. The first roasting parameter is then compared with the predicted current system state, and the second roasting parameter is calculated using the Kalman gain. These steps are repeated to achieve real-time estimation and prediction of the second roasting parameter.

[0083] When there is noise interference in the first baking parameters, the Kalman filter can eliminate the interference through state estimation, making the obtained filtering results more stable and accurate.

[0084] Thereafter, the filtering result may be input into a deep learning prediction model, and the second baking parameter may be outputted through the deep learning prediction model.

[0085] In an embodiment of the present invention, before predicting the second baking parameters using a deep learning prediction model, a Kalman filter may be performed on the first baking parameters to eliminate noise interference in the first baking parameters, thereby making the obtained second baking parameters more accurate.

[0086] Based on the above embodiment, the step of generating a first real-time control instruction and a third baking parameter within a preset time period based on the first baking parameter and the second baking parameter includes:

[0087] Establishing a system model of the tea roasting process based on a model predictive control algorithm, and predicting the third roasting parameter based on the system model and the first roasting parameter;

[0088] The first real-time control instruction is solved based on the third baking parameter and the optimization target.

[0089] Specifically, when generating the first real-time control instruction and the third roasting parameter within a preset time period, an MPC algorithm can be used. The MPC algorithm can first establish a system model of the tea roasting process based on the physical characteristics and control requirements of the tea roasting process. The system model can be a mathematical model. The dynamic characteristics of the system model can be described using a state space model, a transfer function model, or the like.

[0090] Furthermore, the baking parameter is used as the system state. By using the system model and the first baking parameter, a rolling time domain prediction method can be adopted to predict the system state within a preset time period in the future, that is, the third baking parameter.

[0091] The optimal control sequence can be solved based on the third baking parameter and the optimization goal. This optimization goal can take into account multiple control objectives such as temperature, humidity, and air quality, as well as constraints such as energy consumption and equipment life.

[0092] Finally, the first control instruction in the optimal control sequence may be used as the first real-time control instruction and sent to the baking parameter adjustment mechanism to adjust the first baking parameter.

[0093] In this embodiment of the present invention, a model predictive control algorithm is applied to the first and second roasting parameters to generate a first real-time control instruction and a third roasting parameter within a preset time period, providing a basis for generating the second real-time control instruction. Furthermore, by combining the model predictive control algorithm with a reinforcement learning network, adaptive control can be achieved, automatically controlling the roasting parameter adjustment mechanism to adjust the first roasting parameter. This significantly improves the precision and stability of tea roasting, enhances the intelligent level of tea roasting, reduces manual intervention, and reduces the difficulty and cost of manual operation.

[0094] Based on the above embodiment, the method of applying a reinforcement learning network to generate a second real-time control instruction based on the first real-time control instruction and the third baking parameter includes:

[0095] extracting control decision information from the first real-time control instruction;

[0096] Inputting the control decision information and the third baking parameter into the reinforcement learning network to obtain the second real-time control instruction output by the reinforcement learning network based on the state space, the action space, and the reward function;

[0097] The reinforcement learning network learns the action with the highest value based on the state and reward value of the tea roasting process. The state space includes roasting parameters, the action space includes control instructions, and the reward function is determined based on the roasting effect.

[0098] Specifically, in an embodiment of the present invention, after the first real-time control instruction is determined, control decision information in the first real-time control instruction may be extracted, such as the opening degree range of the temperature adjustment device.

[0099] Thereafter, the control decision information and the third baking parameter are input into the DQN, and the DQN may output a second real-time control instruction in a control instruction generation phase.

[0100] DQN can be divided into several phases: environment modeling, policy learning, policy updating, and control instruction generation. The environment modeling, policy learning, and policy updating phases are all training phases of DQN, while the control instruction generation phase is the application phase of DQN.

[0101] During the environment modeling phase, the state space, action space, and reward function can be defined separately. The state space can include roasting parameters such as temperature, humidity, and air quality. It can also include secondary roasting parameters as well as primary roasting parameters. Furthermore, it may include statistical characteristics of historical roasting parameters, such as average temperature and humidity trends. These roasting parameters together form a description of the current state of the tea roasting process, providing a basis for the DQN's decision-making.

[0102] The action space includes control instructions for baking parameter control mechanisms such as heaters, fans, and air quality regulators. For example, a heater control instruction could be to increase, decrease, or maintain the temperature; a fan control instruction could be to increase, decrease, or turn off the fan; and an air quality regulator control instruction could be to increase, decrease, or turn off the air purification intensity. The combination of these control instructions constitutes the action space that the DQN can select.

[0103] The reward function can be determined based on roasting quality metrics such as the tea's color, aroma, and moisture content. For example, if the roasted tea has the desired color, a rich aroma, and an appropriate moisture content, a higher reward value is given; if the tea is burnt, lacks aroma, or has an excessively high or low moisture content, a lower reward value is given. The reward function is designed to guide the DQN to learn the control strategy that produces the best roasting results.

[0104] During the policy learning phase, the DQN, acting as an intelligent agent, selects an action from the action space based on the current state information (i.e., control decision information) and the third baking parameter. This selection process is typically based on the DQN's estimated value (i.e., Q-value) for different state-action pairs. The DQN estimates Q-values ​​using a neural network, taking the current state as input and outputting a Q-value for each possible action. The action with the highest Q-value is selected as the current control decision. After the baking parameter adjustment mechanism executes the selected action, the agent observes the feedback from the environment, which results in a reward value. This reward is calculated based on a reward function and reflects the impact of the selected action on the baking result. For example, if the action of increasing the heater temperature is selected, and subsequent sensor data indicates an increase in temperature, resulting in improved baking results, a higher reward value may be obtained. Conversely, if the temperature increase causes the tea leaves to burn, the reward value will be lower.

[0105] During the policy update phase, a deep neural network is used to approximate the Q-function. By continuously adjusting the neural network parameters, the estimated Q-values ​​are brought closer and closer to the true Q-values. In each iteration, the DQN randomly draws a batch of samples (including state, action, reward, and next state) from the experience replay buffer and then uses these samples for training. The neural network parameters are updated by minimizing the loss function (typically the mean squared error between the predicted Q-values ​​and the target Q-values).

[0106] Policy parameters are adjusted based on the reward value and the estimated value of the Q function. The reward value plays a key role in policy updates. If an action receives a high reward value, the DQN adjusts the policy parameters to make it more likely to be chosen in similar states. Conversely, if an action receives a low reward value, the DQN reduces the probability of choosing this action in similar states. At the same time, the estimated value of the Q function is also used to guide policy updates. The DQN strives to make the estimated Q value more accurate so that it can make more informed decisions when choosing an action.

[0107] During the control instruction generation phase, after multiple iterations and policy updates, DQN gradually learns the optimal control strategy. When a new state is generated, DQN selects the action with the highest Q value as the second real-time control instruction based on the current state information.

[0108] In the embodiment of the present invention, the optimal roasting strategy is learned based on the current state and reward value of the tea roasting process through a reinforcement learning network to improve the quality and production efficiency of the tea.

[0109] Based on the above embodiment, controlling the baking parameter adjustment mechanism in the baking machine based on the first real-time control instruction and the second real-time control instruction includes:

[0110] generating a target control instruction based on the first real-time control instruction and the second real-time control instruction;

[0111] Based on the target control instruction, a closed-loop control algorithm is applied to control the baking parameter adjustment mechanism.

[0112] Specifically, when controlling the baking parameter adjustment mechanism in the baking machine, the first real-time control instruction and the second real-time control instruction can be used to generate a target control instruction. For example, the first real-time control instruction and the second real-time control instruction can be integrated by weighted averaging, fuzzy logic, etc. to obtain the target control instruction, so as to give full play to the advantages of the two algorithms.

[0113] Subsequently, a closed-loop control algorithm is applied to the roasting parameter adjustment mechanism using the target control instructions. This closed-loop control algorithm, which can be a proportional-integral-differential (PID) control algorithm, continuously optimizes the control strategy based on real-time feedback of the first roasting parameter, improving system stability and accuracy. This ensures the roaster is always operating in optimal conditions, enhances roasting consistency, and improves roasting control stability and accuracy.

[0114] like Figure 2 As shown, based on the above embodiment, an embodiment of the present invention provides a multi-layer instruction adaptive control device, including:

[0115] A baking parameter acquisition module 21 is used to obtain a first baking parameter in the baking machine at the current moment;

[0116] a baking parameter prediction module 22, configured to predict a second baking parameter at a next moment in the baking machine based on the first baking parameter;

[0117] A first instruction generating module 23 is configured to generate a first real-time control instruction and a third baking parameter within a preset time period based on the first baking parameter and the second baking parameter;

[0118] a second instruction generating module 24 for applying a reinforcement learning network to generate a second real-time control instruction based on the first real-time control instruction and the third baking parameter;

[0119] The baking control module 25 is configured to control a baking parameter adjustment mechanism within the baking machine based on the first real-time control instruction and the second real-time control instruction.

[0120] On the basis of the above embodiments, in the multi-layer instruction adaptive control device provided in the embodiments of the present invention, the baking parameter prediction module is specifically used to:

[0121] Performing Kalman filtering on the first baking parameter to obtain a filtering result;

[0122] Inputting the filtering result into a deep learning prediction model to obtain the second baking parameter output by the deep learning prediction model;

[0123] The deep learning prediction model is trained based on historical baking parameters and baking effect indicators.

[0124] On the basis of the above embodiment, in the multi-layer instruction adaptive control device provided in the embodiment of the present invention, the first instruction generation module is specifically configured to:

[0125] Establishing a system model of the tea roasting process based on a model predictive control algorithm, and predicting the third roasting parameter based on the system model and the first roasting parameter;

[0126] The first real-time control instruction is solved based on the third baking parameter and the optimization target.

[0127] On the basis of the above embodiment, in the multi-layer instruction adaptive control device provided in the embodiment of the present invention, the second instruction generating module is specifically configured to:

[0128] extracting control decision information from the first real-time control instruction;

[0129] Inputting the control decision information and the third baking parameter into the reinforcement learning network to obtain the second real-time control instruction output by the reinforcement learning network based on the state space, the action space, and the reward function;

[0130] The reinforcement learning network learns the action with the highest value based on the state and reward value of the tea roasting process. The state space includes roasting parameters, the action space includes control instructions, and the reward function is determined based on the roasting effect.

[0131] On the basis of the above embodiments, in the multi-layer instruction adaptive control device provided in the embodiments of the present invention, the baking control module is specifically configured to:

[0132] generating a target control instruction based on the first real-time control instruction and the second real-time control instruction;

[0133] Based on the target control instruction, a closed-loop control algorithm is applied to control the baking parameter adjustment mechanism.

[0134] Based on the above embodiment, in the multi-layer instruction adaptive control device provided in the embodiment of the present invention, the first baking parameter includes at least one of humidity, temperature, carbon dioxide concentration, carbon monoxide concentration and air quality.

[0135] Specifically, the functions of each module in the multi-layer instruction adaptive control device provided in the embodiment of the present invention correspond one-to-one to the operating procedures of each step in the above-mentioned method embodiment, and the effects achieved are also consistent. Please refer to the above-mentioned embodiment for details, and no further details will be given in the embodiment of the present invention.

[0136] like Figure 3 As shown, based on the above embodiment, a multi-layer instruction adaptive control system is further provided in an embodiment of the present invention, including: a controller 31, a baking machine 32 and a baking parameter acquisition device 33, and the baking parameter acquisition device 33 can be connected to the controller 31 by wire.

[0137] The roaster 32 can be loaded with tea leaves and roast the tea leaves.

[0138] The baking parameter acquisition device 33 is disposed in the baking machine 32 and is used to acquire the first baking parameter in real time.

[0139] The controller 31 may receive the first baking parameter collected by the baking parameter collection device 33 and execute the multi-layer instruction adaptive control method provided in the above embodiments.

[0140] The multi-layer instruction adaptive control system provided in the embodiment of the present invention can collect the first roasting parameter in real time through the roasting parameter acquisition device, and can realize automatic control of the environment in the roaster through the controller, thereby ensuring the roasting effect of tea.

[0141] Based on the above embodiment, the baking parameter acquisition device includes at least one of a temperature sensor, a humidity sensor, a carbon monoxide sensor, a carbon dioxide sensor and an air quality sensor.

[0142] Specifically, if Figure 4 As shown, after the baking parameter acquisition device acquires the first baking data, it performs Kalman filtering to obtain a filtering result.

[0143] On the one hand, the filtering result is input into the deep learning prediction model, and the second baking parameter is output through the deep learning prediction model.

[0144] On the other hand, the first real-time control instruction and the third baking parameter are generated by using the filtering result and the second baking parameter with the help of a model predictive control algorithm.

[0145] The first real-time control instruction and the third baking parameter are used to generate a second real-time control instruction by applying a reinforcement learning network.

[0146] By combining the first real-time control instruction and the second real-time control instruction, a target control instruction can be obtained, and the target control instruction is used to control the baking parameter adjustment mechanism in the baking machine.

[0147] During the operation of the baking parameter adjustment mechanism, the baking parameter acquisition device realizes closed-loop control by acquiring the first baking parameter in real time.

[0148] Figure 5 An example of a physical structure diagram of an electronic device is shown below. Figure 5 As shown, the electronic device may include: a processor (Processor) 510, a communication interface (Communications Interface) 520, a memory (Memory) 530, and a communication bus 540. The processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the multi-layer instruction adaptive control method provided in the above embodiments.

[0149] Furthermore, the logic instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0150] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-layer instruction adaptive control method provided in the above embodiments.

[0151] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to execute the multi-layer instruction adaptive control method provided in the above embodiments.

[0152] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0153] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multi-layer instruction adaptive control method, characterized in that: include: Obtaining a first baking parameter in the baking machine at a current moment; the first baking parameter includes at least one of humidity, temperature, carbon dioxide concentration, carbon monoxide concentration, and air quality; Based on the first roasting parameter, predicting a second roasting parameter at a next moment in the roasting machine; generating a first real-time control instruction and a third baking parameter within a preset time period based on the first baking parameter and the second baking parameter; Based on the first real-time control instruction and the third baking parameter, applying a reinforcement learning network to generate a second real-time control instruction; controlling a baking parameter adjustment mechanism in the baking machine based on the first real-time control instruction and the second real-time control instruction; The step of predicting a second baking parameter at a next moment in the baking machine based on the first baking parameter includes: Performing Kalman filtering on the first baking parameter to obtain a filtering result; Inputting the filtering result into a deep learning prediction model to obtain the second baking parameter output by the deep learning prediction model; Among them, the deep learning prediction model is trained based on historical roasting parameters and roasting effect indicators. The roasting effect indicators are used to indicate the actual roasting effect that can be achieved by tea under the environment of historical roasting parameters. It is a set of parameters used to measure the quality and state achieved by tea during the roasting process.

2. The multi-layer instruction adaptive control method according to claim 1, characterized in that: The generating, based on the first baking parameter and the second baking parameter, a first real-time control instruction and a third baking parameter within a preset time period includes: Establishing a system model of the tea roasting process based on a model predictive control algorithm, and predicting the third roasting parameter based on the system model and the first roasting parameter; The first real-time control instruction is solved based on the third baking parameter and the optimization target.

3. The multi-layer instruction adaptive control method according to claim 1, characterized in that: The step of applying a reinforcement learning network to generate a second real-time control instruction based on the first real-time control instruction and the third baking parameter includes: extracting control decision information from the first real-time control instruction; Inputting the control decision information and the third baking parameter into the reinforcement learning network to obtain the second real-time control instruction output by the reinforcement learning network based on the state space, the action space, and the reward function; The reinforcement learning network learns the action with the highest value based on the state and reward value of the tea roasting process. The state space includes roasting parameters, the action space includes control instructions, and the reward function is determined based on the roasting effect.

4. The multi-layer instruction adaptive control method according to any one of claims 1 to 3, characterized in that: The controlling of the baking parameter adjustment mechanism in the baking machine based on the first real-time control instruction and the second real-time control instruction includes: generating a target control instruction based on the first real-time control instruction and the second real-time control instruction; Based on the target control instruction, a closed-loop control algorithm is applied to control the baking parameter adjustment mechanism.

5. A multi-layer instruction adaptive control device, characterized in that: include: a baking parameter acquisition module, configured to acquire a first baking parameter in the baking machine at a current moment; the first baking parameter comprising at least one of humidity, temperature, carbon dioxide concentration, carbon monoxide concentration, and air quality; a baking parameter prediction module, configured to predict a second baking parameter at a next moment in the baking machine based on the first baking parameter; a first instruction generating module, configured to generate a first real-time control instruction and a third baking parameter within a preset time period based on the first baking parameter and the second baking parameter; a second instruction generating module, configured to generate a second real-time control instruction based on the first real-time control instruction and the third baking parameter by applying a reinforcement learning network; a baking control module, configured to control a baking parameter adjustment mechanism within the baking machine based on the first real-time control instruction and the second real-time control instruction; The baking parameter prediction module is specifically used for: Performing Kalman filtering on the first baking parameter to obtain a filtering result; Inputting the filtering result into a deep learning prediction model to obtain the second baking parameter output by the deep learning prediction model; Among them, the deep learning prediction model is trained based on historical roasting parameters and roasting effect indicators. The roasting effect indicators are used to indicate the actual roasting effect that can be achieved by tea under the environment of historical roasting parameters. It is a set of parameters used to measure the quality and state achieved by tea during the roasting process.

6. A multi-layer instruction adaptive control system, characterized in that: include: A controller, a baking machine, and a baking parameter acquisition device, wherein the baking parameter acquisition device is connected to the controller; The roasting machine is used to load tea leaves and roast the tea leaves; The baking parameter acquisition device is provided in the baking machine and is used to acquire the first baking parameter in real time; The controller is configured to receive the first baking parameter and execute the multi-layer instruction adaptive control method according to any one of claims 1 to 4.

7. The multi-layer instruction adaptive control system according to claim 6, characterized in that: The baking parameter acquisition device includes at least one of a temperature sensor, a humidity sensor, a carbon monoxide sensor, a carbon dioxide sensor, and an air quality sensor.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-layer instruction adaptive control method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Manufacturing process control using constrained reinforcement machine learning

    US20210247744A1

  • Method for displaying the residual time until a cooking process has been finished

    WO2009026887A2