Multi-layer instruction adaptive control method, device and system and storage medium

Through the multi-layer instruction adaptive control method, the tea baking parameters are dynamically adjusted using Kalman filtering, deep learning and reinforcement learning networks, which solves the problem that traditional tea baking equipment cannot adapt to environmental changes and achieves the stability and consistency of baking effects.

CN120295149AActive Publication Date: 2025-07-11QUANZHOU INST OF EQUIP MFG +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510793856.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-11
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

Traditional tea baking equipment is difficult to adapt to the real-time changes in environmental parameters during the baking process, resulting in unstable baking effect and poor consistency.

Method used

The multi-layer instruction adaptive control method is adopted to obtain the current baking parameters, and use Kalman filtering, deep learning prediction model, model prediction control algorithm and reinforcement learning network to generate real-time control instructions to dynamically adjust the baking parameters.

Benefits of technology

The stability and consistency of tea baking effects are achieved, ensuring the consistency of quality of different tea types and baking batches, and reducing manual intervention and operation difficulty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295149A_ABST
    Figure CN120295149A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of tea baking, and provides a multi-layer instruction adaptive control method, device and system and a storage medium, and the method comprises the steps: obtaining a first baking parameter at a current moment in a baking machine; predicting a second baking parameter at the next moment by using the first baking parameter; generating a first real-time control instruction and a third baking parameter in a preset time period by using the first baking parameter and the second baking parameter; generating a second real-time control instruction by using the first real-time control instruction and the third baking parameter and applying a reinforcement learning network; and controlling the baking parameter adjusting mechanism by using the first real-time control instruction and the second real-time control instruction. According to the method, through the first baking parameter, the second baking parameter and the third baking parameter, the control instruction is adaptively determined, and the baking parameter adjusting mechanism in the baking machine is adaptively and dynamically controlled, so that the optimal baking effect is achieved, the stability and consistency of the baking effect are ensured, and the quality of different tea types and baking batches is consistent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tea baking, and in particular, to a multi-layer instruction adaptive control method, device, system, and storage medium. Background Art

[0002] Tea baking plays a crucial role in the entire tea production process, and the accuracy of its control directly determines the quality and taste of tea.

[0003] Traditional tea baking equipment mainly relies on a preset temperature control curve and highly depends on manual operation for adjustment during the adjustment process. However, during the actual baking process, the environmental parameters are constantly changing, and it is difficult for traditional tea baking equipment to adapt to the real-time changes of environmental parameters during the baking process, and it is impossible to ensure the stability and consistency of the baking effect. For example, temperature fluctuations may cause the tea to be burnt or under-baked; unstable air quality may introduce odors or affect the color of the tea.

[0004] Based on this, there is an urgent need to provide a multi-layer instruction adaptive control method that can be applied to tea baking control. Summary of the Invention

[0005] The present invention provides a multi-layer instruction adaptive control method, device, system, and storage medium to solve the defects existing in the prior art.

[0006] The present invention provides a multi-layer instruction adaptive control method, including: Obtain the first baking parameter at the current moment in the baking machine; Based on the first baking parameter, predict the second baking parameter at the next moment in the baking machine; Based on the first baking parameter and the second baking parameter, generate a first real-time control instruction and the third baking parameter within a preset time period; Based on the first real-time control instruction and the third baking parameter, apply a reinforcement learning network to generate a second real-time control instruction; Based on the first real-time control instruction and the second real-time control instruction, control the baking parameter adjustment mechanism in the baking machine.

[0007] According to the multi-layer instruction adaptive control method provided by the present invention, the predicting the second baking parameter at the next moment in the baking machine based on the first baking parameter includes: Perform Kalman filtering on the first baking parameter to obtain a filtering result; Input the filtering result into a deep learning prediction model to obtain the second baking parameter output by the deep learning prediction model; Among them, the deep learning prediction model is trained based on historical baking parameters and baking effect indicators.

[0008] According to a multi-layer instruction adaptive control method provided by the present invention, generating a first real-time control instruction and a third baking parameter within a preset time period based on the first baking parameter and the second baking parameter includes: Based on a model predictive control algorithm, establish a system model of the tea baking process, and predict the third baking parameter based on the system model and the first baking parameter; Solve the first real-time control instruction based on the third baking parameter and the optimization objective.

[0009] According to a multi-layer instruction adaptive control method provided by the present invention, generating a second real-time control instruction by applying a reinforcement learning network based on the first real-time control instruction and the third baking parameter includes: Extract the control decision information in the first real-time control instruction; Input the control decision information and the third baking parameter into the reinforcement learning network to obtain the second real-time control instruction output by the reinforcement learning network based on the state space, action space, and reward function; The reinforcement learning network learns the action with the highest value based on the state and reward value of the tea baking process. The state space includes baking parameters, the action space includes control instructions, and the reward function is determined based on the baking effect.

[0010] According to a multi-layer instruction adaptive control method provided by the present invention, controlling a baking parameter adjustment mechanism in the baking machine based on the first real-time control instruction and the second real-time control instruction includes: Generate a target control instruction based on the first real-time control instruction and the second real-time control instruction; Based on the target control instruction, apply a closed-loop control algorithm to control the baking parameter adjustment mechanism.

[0011] According to a multi-layer instruction adaptive control method provided by the present invention, the first baking parameter includes at least one of humidity, temperature, carbon dioxide concentration, carbon monoxide concentration, and air quality.

[0012] The present invention also provides a multi-layer instruction adaptive control device, including: A baking parameter acquisition module for acquiring a first baking parameter at the current moment in the baking machine; A baking parameter prediction module for predicting a second baking parameter at the next moment in the baking machine based on the first baking parameter; A first instruction generation module, configured to generate a first real-time control instruction and third baking parameters within a preset time period based on the first baking parameters and the second baking parameters; A second instruction generation module, configured to apply a reinforcement learning network to generate a second real-time control instruction based on the first real-time control instruction and the third baking parameters; A baking control module, configured to control a baking parameter adjustment mechanism in the baking machine based on the first real-time control instruction and the second real-time control instruction.

[0013] The present invention further provides a multi-layer instruction adaptive control system, including: a controller, a baking machine, and a baking parameter acquisition device, where the baking parameter acquisition device is connected to the controller; The baking machine is configured to load tea leaves and bake the tea leaves; The baking parameter acquisition device is disposed in the baking machine and is configured to acquire first baking parameters in real time; The controller is configured to receive the first baking parameters and execute the multi-layer instruction adaptive control method described above.

[0014] According to the multi-layer instruction adaptive control system provided by the present invention, the baking parameter acquisition device includes at least one of a temperature sensor, a humidity sensor, a carbon monoxide sensor, a carbon dioxide sensor, and an air quality sensor.

[0015] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the multi-layer instruction adaptive control method described in any one of the above is implemented.

[0016] The present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the multi-layer instruction adaptive control method described in any one of the above is implemented.

[0017] Compared with the prior art, the present invention has the following beneficial effects: The multi-layer instruction adaptive control method, device, system and storage medium provided by the present invention first obtain the first baking parameter at the current moment in the baking machine; then use the first baking parameter to predict the second baking parameter at the next moment in the baking machine; thereafter, use the first baking parameter and the second baking parameter to generate the first real-time control instruction and the third baking parameter within a preset time period; thereafter, use the first real-time control instruction and the third baking parameter to apply a reinforcement learning network to generate the second real-time control instruction; finally, use the first real-time control instruction and the second real-time control instruction to control the baking parameter adjustment mechanism in the baking machine. Through the first baking parameter, the second baking parameter and the third baking parameter, this method can adaptively determine the control instruction, and then realize the adaptive dynamic control of the baking parameter adjustment mechanism in the baking machine through this control instruction, so as to achieve the best baking effect, ensure the stability and consistency of the baking effect, and make the quality of different tea varieties and baking batches consistent. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below can also be obtained by those of ordinary skill in the art without creative efforts based on these drawings.

[0019] Figure 1 is a flowchart of the multi-layer instruction adaptive control method provided by the present invention; Figure 2 is a structural diagram of the multi-layer instruction adaptive control device provided by the present invention; Figure 3 is one of the structural diagrams of the multi-layer instruction adaptive control system provided by the present invention; Figure 4 is another structural diagram of the multi-layer instruction adaptive control system provided by the present invention; Figure 5 is a structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0021] In the actual tea baking process, the environmental parameters are constantly changing. Since traditional tea baking equipment highly relies on manual operation for adjustment, it is difficult to adapt to the real-time changes of environmental parameters during the baking process, and thus it is impossible to ensure the stability and consistency of the baking effect. Based on this, an embodiment of the present invention provides a tea baking control method.

[0022] Figure 1 It is a schematic flow chart of a multi-layer instruction adaptive control method provided in an embodiment of the present invention. As Figure 1 shown, the method includes: S1, obtaining the first baking parameter at the current moment in the baking machine; S2, predicting the second baking parameter at the next moment in the baking machine based on the first baking parameter; S3, generating a first real-time control instruction and a third baking parameter within a preset time period based on the first baking parameter and the second baking parameter; S4, applying a reinforcement learning network based on the first real-time control instruction and the third baking parameter to generate a second real-time control instruction; S5, controlling the baking parameter adjustment mechanism in the baking machine based on the first real-time control instruction and the second real-time control instruction.

[0023] Specifically, for the multi-layer instruction adaptive control method provided in an embodiment of the present invention, the execution entity is a controller, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., and no specific limitation is made here.

[0024] First, step S1 is executed to obtain the first baking parameter at the current moment in the baking machine. The first baking parameter refers to the baking parameter in the baking machine at the current moment, which can be an environmental parameter. For example, it can include at least one of humidity, temperature, carbon dioxide concentration, carbon monoxide concentration, and air quality. The first baking parameter can be collected by a baking parameter collection device, and the baking parameter collection device can include at least one of a humidity sensor, a temperature sensor, a carbon dioxide sensor, a carbon monoxide sensor, and an air quality sensor.

[0025] The baking parameter collection device can communicate with the controller through a wired connection (such as RS485) to ensure fast transmission speed and high reliability. The position of the baking parameter collection device is optimized to ensure that the first baking parameter at each key position can be collected.

[0026] A temperature sensor can select a thermocouple sensor with high precision and high temperature resistance, which can accurately measure the temperature inside the baking machine in a high-temperature environment. The temperature sensors are arranged at different positions of the baking machine, including the top, bottom, side, etc., to ensure that the temperature distribution inside the baking machine can be comprehensively monitored.

[0027] A humidity sensor can select a humidity sensor with high precision, fast response speed and suitable for high-temperature environments. The humidity sensor is reasonably arranged inside the baking machine. It can be considered to be placed in areas close to the tea placement area and where air circulation is relatively frequent, such as the middle space of the baking machine. This can accurately monitor the humidity changes at different positions during the baking process, provide real-time humidity parameters, help to more precisely control the baking environment, ensure that the tea is baked under suitable humidity conditions, and thus improve the quality and taste of the tea.

[0028] Carbon dioxide sensors, carbon monoxide sensors, and air quality sensors can be selected with high precision and large measurement ranges, so as to accurately detect the concentration changes of gases and solid particles (carbon content) during the baking process and provide accurate air quality parameters. Carbon dioxide sensors, carbon monoxide sensors, and air quality sensors can be arranged near the air inlet and outlet of the baking machine to monitor the carbon dioxide concentration, carbon monoxide concentration, and air quality of the inlet and outlet air. The air quality sensor can include a PM2.5 sensor and a PM10 sensor.

[0029] Then step S2 is executed. Using the first baking parameter, the second baking parameter at the next moment inside the baking machine is predicted. For example, the first baking parameter can be directly input into the deep learning prediction model, and the second baking parameter is output through the deep learning prediction model. The deep learning prediction model can be a Gated Recurrent Unit (GRU) neural network, which controls the flow of information by introducing an update gate and a reset gate. The update gate determines how much information from the previous state is included in the current state, and the reset gate determines how much information is forgotten. The structure of the GRU neural network is relatively simple, with a fast training speed, and can effectively process time series data.

[0030] The deep learning prediction model can be trained using historical baking parameters and baking effect indicators. The historical baking parameters can include parameters such as temperature, humidity, carbon monoxide concentration, carbon dioxide concentration, and air quality. After collecting the historical baking parameters, preprocessing of the historical baking parameters can also be performed, including operations such as data cleaning and normalization.

[0031] The baking effect index is the label of supervised learning, which is used to indicate the actual achievable effect of tea under the environment of historical baking parameters. The baking effect index is a set of parameters used to measure the quality and state of tea during the baking process, which can comprehensively reflect the performance of tea in terms of appearance, smell, taste, and chemical characteristics after baking, and is the key basis for evaluating the baking control performance of tea.

[0032] In the embodiments of the present invention, the baking effect index may include tea color, tea aroma, tea water content, tea morphology, tea taste, etc. The tea color can be evaluated by image analysis or comparison with a standard color sample to determine whether the color of the tea meets the expectation. The tea aroma can be evaluated using a professional odor sensor or by a professional tea taster, who gives the intensity and quality grade of the aroma. The tea water content can be measured using a moisture meter. The tea morphology can be determined by observing whether the baked tea is intact, without breakage or deformation, or image recognition technology can be used to automatically detect the integrity and degree of breakage of the tea. By training a tea morphology evaluation model, different forms of tea can be recognized and corresponding evaluation results can be given. The tea taste can be evaluated indirectly by an automated detection method through analyzing the chemical components of the tea. For example, measuring the content of components such as tea polyphenols and caffeine in the tea, which are closely related to the taste of the tea.

[0033] The historical baking parameters are divided into training data and test data, and the training data is input into the initial prediction model for training. The backpropagation algorithm and optimization algorithms (such as the Adam optimizer) are used to adjust the structural parameters of the initial prediction model to minimize the prediction error.

[0034] Thereafter, the trained deep learning prediction model is used to predict the test data and compared with the baking effect index corresponding to the test data to evaluate the performance of the deep learning prediction model. Here, metrics such as Root Mean Square Error (RMSE) and Mean Absolute Error (MAE) can be used to measure the prediction accuracy of the deep learning prediction model.

[0035] Subsequently, step S3 is executed to generate a first real-time control instruction and third baking parameters within a future preset time period by using the first baking parameter and the second baking parameter. Among them, it can be implemented by using the Model Predictive Control (MPC) algorithm. The MPC algorithm is an advanced control algorithm based on predicting future states. It establishes a system model of the tea baking process, predicts the system state within a future period of time, that is, the third baking parameters within the preset time period, and generates an optimal control sequence, that is, the first real-time control instruction, according to the optimization objective function. Thus, the MPC algorithm can predict the change trend of the third baking parameters within the future preset time period based on the first baking parameter and the second baking parameter, and generate an optimal first real-time control instruction.

[0036] The length of the preset time period can be set as needed and is not specifically limited here.

[0037] Finally, step S4 is executed to apply the Deep Q-Network (DQN) by using the first real-time control instruction and the third baking parameters to generate a second real-time control instruction. The DQN is introduced to further optimize and adjust the first real-time control instruction. The input of the DQN can include both the key parameters in the first real-time control instruction and the third baking parameters, or the third baking parameters and the control decision information determined by the key parameters, etc. The output is the second real-time control instruction. It can be understood that if the first real-time control instruction is to control the temperature to a set value, then the set value is the key parameter. The control decision information is the range of the opening degree of the temperature adjustment device, such as the indication information of high, medium, and low gears, rather than the specific value of the key parameter.

[0038] During the process of training the DQN, the initial network can be used to output a control instruction first, and the baking parameter adjustment mechanism can be controlled by the control instruction to realize the intelligent control of the tea baking process. Subsequently, the strategy is continuously adjusted according to the reward value of the baking effect of the tea to learn a better control instruction, so as to further optimize the performance of the DQN under different baking parameters and improve the adaptability and intelligence of the DQN.

[0039] In the embodiment of the present invention, the baking parameter adjustment mechanism may include devices such as a temperature adjustment device, a fan, and an air quality regulator. The temperature adjustment device may be a heater, and the air quality regulator may be an air purification system.

[0040] The temperature inside the baking machine can be adjusted by the temperature adjustment device, the humidity, carbon monoxide concentration, and carbon dioxide concentration inside the baking machine can be adjusted by the fan, and the air quality inside the baking machine can be adjusted by the air quality regulator.

[0041] Finally, step S5 is executed to control the baking parameter adjustment mechanism in the baking machine by using the first real-time control instruction and the second real-time control instruction. Here, methods such as weighted average and fuzzy logic can be used to fuse the first real-time control instruction and the second real-time control instruction to obtain a target control instruction, and then the baking parameter adjustment mechanism is controlled through the target control instruction to adjust the first baking parameter.

[0042] In the multi-layer instruction adaptive control method provided in the embodiment of the present invention, first, the first baking parameter at the current moment in the baking machine is obtained; then, the second baking parameter at the next moment in the baking machine is predicted by using the first baking parameter; thereafter, the first real-time control instruction and the third baking parameter within a preset time period are generated by using the first baking parameter and the second baking parameter; thereafter, the second real-time control instruction is generated by applying a reinforcement learning network by using the first real-time control instruction and the third baking parameter; finally, the baking parameter adjustment mechanism in the baking machine is controlled by using the first real-time control instruction and the second real-time control instruction. Through the first baking parameter, the second baking parameter, and the third baking parameter, this method can adaptively determine the control instruction, and then adaptively and dynamically control the baking parameter adjustment mechanism in the baking machine through this control instruction to achieve the best baking effect, ensure the stability and consistency of the baking effect, and make the quality of different tea varieties and baking batches consistent.

[0043] Based on the above embodiment, predicting the second baking parameter at the next moment in the baking machine based on the first baking parameter includes: Performing Kalman filtering on the first baking parameter to obtain a filtering result; Inputting the filtering result into a deep learning prediction model to obtain the second baking parameter output by the deep learning prediction model; Wherein, the deep learning prediction model is trained based on historical baking parameters and baking effect indicators.

[0044] Specifically, when predicting the second baking parameter at the next moment in the baking machine, Kalman filtering can be first performed on the first baking parameter to obtain a filtering result. Here, the process of performing Kalman filtering on the first baking parameter is a process of smoothing the first baking parameter by using the Kalman filtering algorithm.

[0045] The Kalman filtering algorithm is a recursive algorithm based on linear minimum variance estimation. By estimating and predicting the system state, the noise interference in the first baking parameter is eliminated. During the tea baking process, the Kalman filtering algorithm can perform real-time processing on the first baking parameter to obtain a more accurate and stable first baking parameter.

[0046] In the process of performing Kalman filtering, the baking parameters can be used as the system state. First, based on the system model of the tea baking process and the system state at the previous moment, the system state at the current moment is predicted. Then, the first baking parameter is compared with the predicted system state at the current moment, and the second baking parameter is calculated through the Kalman gain. By repeating the above steps, the real-time estimation and prediction of the second baking parameter are achieved.

[0047] In the case where there is noise interference in the first baking parameter, Kalman filtering can eliminate these interferences through state estimation, making the obtained filtering result more stable and accurate.

[0048] After that, the filtering result can be input into the deep learning prediction model, and the second baking parameter is output through the deep learning prediction model.

[0049] In the embodiment of the present invention, before predicting the second baking parameter through the deep learning prediction model, Kalman filtering can be performed on the first baking parameter to eliminate the noise interference existing in the first baking parameter, making the obtained second baking parameter more accurate.

[0050] Based on the above embodiments, generating the first real-time control instruction and the third baking parameter within a preset time period based on the first baking parameter and the second baking parameter includes: Based on the model predictive control algorithm, a system model of the tea baking process is established, and based on the system model and the first baking parameter, the third baking parameter is predicted; Based on the third baking parameter and the optimization objective, the first real-time control instruction is solved.

[0051] Specifically, when generating the first real-time control instruction and the third baking parameter within a preset time period, the MPC algorithm can be used to implement. The MPC algorithm can first establish a system model of the tea baking process according to the physical characteristics and control requirements of the tea baking process. This system model can be a mathematical model. The dynamic characteristics of the system model are described in the form of a state space model, a transfer function model, etc.

[0052] Furthermore, taking the baking parameter as the system state, using the system model and the first baking parameter, the method of rolling horizon prediction can be used to predict the system state within a future preset time period, that is, the third baking parameter.

[0053] According to the third baking parameter and the optimization objective, the optimal control sequence can be solved. This optimization objective can consider multiple control objectives such as temperature, humidity, air quality, etc., as well as constraint conditions such as energy consumption and equipment life.

[0054] Finally, the first control instruction in the optimal control sequence can be used as the first real-time control instruction and sent to the baking parameter adjustment mechanism to adjust the first baking parameter.

[0055] In the embodiment of the present invention, based on the first baking parameter and the second baking parameter, the model predictive control algorithm is applied to generate the first real-time control instruction and the third baking parameter within a preset time period, providing a basis for the generation of the second real-time control instruction. Moreover, by combining the model predictive control algorithm with the reinforcement learning network, adaptive control can be achieved, enabling the automatic control of the baking parameter adjustment mechanism to adjust the first baking parameter, significantly improving the accuracy and stability of tea baking, enhancing the intelligent level of tea baking, reducing manual intervention, and lowering the difficulty and cost of manual operation.

[0056] Based on the above embodiment, applying the reinforcement learning network based on the first real-time control instruction and the third baking parameter to generate the second real-time control instruction includes: Extracting the control decision information in the first real-time control instruction; Inputting the control decision information and the third baking parameter into the reinforcement learning network to obtain the second real-time control instruction output by the reinforcement learning network based on the state space, action space, and reward function; The reinforcement learning network learns the action with the highest value based on the state and reward value of the tea baking process. The state space includes baking parameters, the action space includes control instructions, and the reward function is determined based on the baking effect.

[0057] Specifically, in the embodiment of the present invention, after determining the first real-time control instruction, the control decision information in the first real-time control instruction can be extracted, such as the opening degree range of the temperature adjustment device.

[0058] Thereafter, the control decision information and the third baking parameter are input into the DQN, and the DQN can output the second real-time control instruction during the control instruction generation stage.

[0059] The DQN can include stages such as environment modeling, policy learning, policy update, and control instruction generation. The environment modeling stage, policy learning stage, and policy update stage are all training stages of the DQN, and the control instruction generation stage is the application stage of the DQN.

[0060] In the environmental modeling stage, the state space, action space, and reward function can be defined separately. The state space can include baking parameters such as temperature, humidity, and air quality, which can include the second baking parameters and also the first baking parameters. In addition, it may also include statistical features of historical baking parameters, such as average temperature, humidity change trend, etc. These baking parameters together constitute the current state description of the tea baking process and provide a basis for decision-making for DQN.

[0061] The action space includes control instructions for baking parameter adjustment mechanisms such as heaters, fans, and air quality regulators. For example, the control instruction for the heater can be to increase the temperature, decrease the temperature, or keep the temperature unchanged; the control instruction for the fan can be to increase the rotation speed, decrease the rotation speed, or turn off; the control instruction for the air quality regulator can be to increase the air purification intensity, decrease the purification intensity, or turn off, etc. The combination of these control instructions constitutes the action space that DQN can choose from.

[0062] The reward function can be determined according to baking effect indicators such as the color, aroma, and moisture content of the tea. For example, if the color of the baked tea meets the expectations, has a strong aroma, and the moisture content is appropriate, then a higher reward value is given; if the tea is burnt, has insufficient aroma, or the moisture content is too high / low, then a lower reward value is given. The design of the reward function aims to guide DQN to learn control strategies that can produce the best baking effect.

[0063] In the policy learning stage, as an agent, DQN selects an action from the action space based on the current state information, that is, the control decision information and the third baking parameter. This selection process is usually based on DQN's value estimation (i.e., Q-value) of different state-action pairs. DQN estimates the Q-value through a neural network, with the current state as the input and the Q-values of each possible action as the output. The action with the highest Q-value is selected as the current control decision. After the baking parameter adjustment mechanism executes the selected action, the agent observes the feedback of the environment, that is, obtains the reward value. This reward value is calculated according to the reward function and reflects the impact of the selected action on the baking effect. For example, if the action of increasing the heater temperature is selected, and then the sensor data shows that the temperature rises and the final baking effect is improved, then a relatively high reward value may be obtained; conversely, if the temperature rise causes the tea to be burnt, the reward value will be lower.

[0064] In the policy update stage, a deep neural network is used to approximate the Q function. By continuously adjusting the parameters of the neural network, the estimated Q-value gets closer and closer to the true Q-value. In each iteration, DQN randomly samples a batch of samples (including state, action, reward, next state) from the experience replay buffer, and then uses these samples for training. The parameters of the neural network are updated by minimizing the loss function (usually the mean square error between the predicted Q-value and the target Q-value).

[0065] The policy parameters are adjusted according to the reward value and the estimated value of the Q function, and the reward value plays a key role in policy update. If an action obtains a high reward value, then DQN will adjust the policy parameters so that this action is more likely to be selected in a similar state. On the contrary, if an action obtains a low reward value, DQN will reduce the probability of selecting this action in a similar state. At the same time, the estimated value of the Q function is also used to guide policy update. DQN endeavors to make the estimated Q value more accurate so as to make more informed decisions when selecting actions.

[0066] In the stage of generating control instructions, after multiple iterations and policy updates, DQN gradually learns the optimal control policy. When a new state is generated, DQN selects the action with the highest Q value as the second real-time control instruction according to the current state information.

[0067] In the embodiment of the present invention, the reinforcement learning network learns the optimal baking strategy according to the current state and the reward value of the tea baking process to improve the quality and production efficiency of tea.

[0068] Based on the above embodiments, controlling the baking parameter adjusting mechanism in the baking machine based on the first real-time control instruction and the second real-time control instruction includes: Generating a target control instruction based on the first real-time control instruction and the second real-time control instruction; Based on the target control instruction, applying a closed-loop control algorithm to control the baking parameter adjusting mechanism.

[0069] Specifically, when controlling the baking parameter adjusting mechanism in the baking machine, the first real-time control instruction and the second real-time control instruction can be used to generate a target control instruction first. For example, methods such as weighted average and fuzzy logic can be used to fuse the first real-time control instruction and the second real-time control instruction to obtain the target control instruction, so as to give full play to the advantages of the two algorithms.

[0070] After that, through the target control instruction, a closed-loop control algorithm is applied to control the baking parameter adjusting mechanism. This closed-loop control algorithm can be a Proportional Integral Differential (PID) control algorithm. According to the first baking parameter fed back in real time, the control strategy is continuously optimized to improve the stability and accuracy of the system, which can ensure that the baking machine is always in the best operating state, improve the consistency of the baking effect, and improve the stability and accuracy of baking control.

[0071] Such as Figure 2As shown in the figure, on the basis of the above embodiments, an adaptive multi-layer instruction control device is provided in an embodiment of the present invention, including: A baking parameter acquisition module 21, configured to acquire first baking parameters at the current moment in the baking machine; A baking parameter prediction module 22, configured to predict second baking parameters at the next moment in the baking machine based on the first baking parameters; A first instruction generation module 23, configured to generate first real-time control instructions and third baking parameters within a preset time period based on the first baking parameters and the second baking parameters; A second instruction generation module 24, configured to generate second real-time control instructions by applying a reinforcement learning network based on the first real-time control instructions and the third baking parameters; A baking control module 25, configured to control a baking parameter adjustment mechanism in the baking machine based on the first real-time control instructions and the second real-time control instructions.

[0072] On the basis of the above embodiments, for the adaptive multi-layer instruction control device provided in an embodiment of the present invention, the baking parameter prediction module is specifically configured to: Perform Kalman filtering on the first baking parameters to obtain a filtering result; Input the filtering result into a deep learning prediction model to obtain the second baking parameters output by the deep learning prediction model; Wherein, the deep learning prediction model is trained based on historical baking parameters and baking effect indicators.

[0073] On the basis of the above embodiments, for the adaptive multi-layer instruction control device provided in an embodiment of the present invention, the first instruction generation module is specifically configured to: Based on a model predictive control algorithm, establish a system model for the tea baking process, and predict the third baking parameters based on the system model and the first baking parameters; Solve the first real-time control instructions based on the third baking parameters and the optimization objective.

[0074] On the basis of the above embodiments, for the adaptive multi-layer instruction control device provided in an embodiment of the present invention, the second instruction generation module is specifically configured to: Extract control decision information in the first real-time control instructions; Input the control decision information and the third baking parameters into the reinforcement learning network to obtain the second real-time control instructions output by the reinforcement learning network based on a state space, an action space, and a reward function; The reinforcement learning network learns the action with the highest value based on the state and reward value of the tea baking process. The state space includes baking parameters, the action space includes control instructions, and the reward function is determined based on the baking effect.

[0075] Based on the above embodiments, in the multi-layer instruction adaptive control device provided in the embodiments of the present invention, the baking control module is specifically configured to: Generate a target control instruction based on the first real-time control instruction and the second real-time control instruction; Based on the target control instruction, apply a closed-loop control algorithm to control the baking parameter adjustment mechanism.

[0076] Based on the above embodiments, in the multi-layer instruction adaptive control device provided in the embodiments of the present invention, the first baking parameters include at least one of humidity, temperature, carbon dioxide concentration, carbon monoxide concentration, and air quality.

[0077] Specifically, the functions of the modules in the multi-layer instruction adaptive control device provided in the embodiments of the present invention correspond one-to-one to the operation processes of the steps in the above method embodiments, and the achieved effects are also the same. For details, please refer to the above embodiments, and the embodiments of the present invention will not be elaborated herein.

[0078] As Figure 3 shown, based on the above embodiments, the embodiments of the present invention further provide a multi-layer instruction adaptive control system, including: a controller 31, a baking machine 32, and a baking parameter acquisition device 33. The baking parameter acquisition device 33 can be connected to the controller 31 in a wired manner.

[0079] The baking machine 32 can load tea and bake the tea.

[0080] The baking parameter acquisition device 33 is arranged in the baking machine 32 and is used to collect the first baking parameters in real time.

[0081] The controller 31 can receive the first baking parameters collected by the baking parameter acquisition device 33 and execute the multi-layer instruction adaptive control method provided in the above embodiments.

[0082] In the multi-layer instruction adaptive control system provided in the embodiments of the present invention, the first baking parameters can be collected in real time through the baking parameter acquisition device, and the environment inside the baking machine can be automatically controlled through the controller, thereby ensuring the baking effect of the tea.

[0083] Based on the above embodiments, the baking parameter acquisition device includes at least one of a temperature sensor, a humidity sensor, a carbon monoxide sensor, a carbon dioxide sensor, and an air quality sensor.

[0084] Specifically, as Figure 4 shown, after the baking parameter acquisition device acquires the first baking data, Kalman filtering is performed to obtain a filtering result.

[0085] On the one hand, the filtering result is input into the deep learning prediction model, and the second baking parameter is output through the deep learning prediction model.

[0086] On the other hand, using the filtering result and the second baking parameter, with the help of the model predictive control algorithm, the first real-time control instruction and the third baking parameter are generated.

[0087] Using the first real-time control instruction and the third baking parameter, the reinforcement learning network is applied to generate the second real-time control instruction.

[0088] Combining the first real-time control instruction and the second real-time control instruction, the target control instruction can be obtained, and the target control instruction is used to control the baking parameter adjustment mechanism in the baking machine.

[0089] During the working process of the baking parameter adjustment mechanism, the baking parameter acquisition device realizes closed-loop control by collecting the first baking parameter in real time.

[0090] Figure 5 An entity structure diagram of an electronic device is exemplified. As Figure 5 shown, the electronic device may include: a processor (Processor) 510, a communication interface (Communications Interface) 520, a memory (Memory) 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 complete mutual communication through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute the multi-layer instruction adaptive control method provided in the above-mentioned embodiments.

[0091] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention. And the foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk, or an optical disc that can store program codes.

[0092] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-layer instruction adaptive control method provided in each of the above embodiments.

[0093] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the multi-layer instruction adaptive control method provided in each of the above embodiments.

[0094] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0095] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-layer instruction adaptive control method, characterized in that Including: Obtain the first baking parameter at the current moment in the baking machine; Based on the first baking parameter, predict the second baking parameter at the next moment in the baking machine; Based on the first baking parameter and the second baking parameter, generate a first real-time control instruction and a third baking parameter within a preset time period; Based on the first real-time control instruction and the third baking parameter, apply a reinforcement learning network to generate a second real-time control instruction; Based on the first real-time control instruction and the second real-time control instruction, control the baking parameter adjustment mechanism in the baking machine.

2. The multi-layer instruction adaptive control method according to claim 1, wherein The predicting the second baking parameter at the next moment in the baking machine based on the first baking parameter includes: Perform Kalman filtering on the first baking parameter to obtain a filtering result; Input the filtering result into a deep learning prediction model to obtain the second baking parameter output by the deep learning prediction model; Wherein, the deep learning prediction model is trained based on historical baking parameters and baking effect indicators.

3. The multi-layer instruction adaptive control method according to claim 1, characterized in that The generating the first real-time control instruction and the third baking parameter within a preset time period based on the first baking parameter and the second baking parameter includes: Based on a model predictive control algorithm, establish a system model for the tea baking process, and predict the third baking parameter based on the system model and the first baking parameter; Solve the first real-time control instruction based on the third baking parameter and the optimization objective.

4. The multi-layer instruction adaptive control method according to claim 1, wherein The applying a reinforcement learning network to generate a second real-time control instruction based on the first real-time control instruction and the third baking parameter includes: Extract the control decision information in the first real-time control instruction; Input the control decision information and the third baking parameter into the reinforcement learning network to obtain the second real-time control instruction output by the reinforcement learning network based on the state space, action space, and reward function; The reinforcement learning network learns the action with the highest value based on the state and reward value of the tea baking process, the state space includes baking parameters, the action space includes control instructions, and the reward function is determined based on the baking effect.

5. The multi-layer instruction adaptive control method according to any one of claims 1-4, characterized in that The controlling the baking parameter adjustment mechanism in the baking machine based on the first real-time control instruction and the second real-time control instruction includes: Generate a target control instruction based on the first real-time control instruction and the second real-time control instruction; Based on the target control instruction, apply a closed-loop control algorithm to control the baking parameter adjustment mechanism.

6. The multi-layer instruction adaptive control method according to any one of claims 1-4, characterized in that, The first baking parameter includes at least one of humidity, temperature, carbon dioxide concentration, carbon monoxide concentration, and air quality.

7. A multi-layer instruction adaptive control device, characterized in that, Including: A baking parameter acquisition module for obtaining the first baking parameter at the current moment in the baking machine; A baking parameter prediction module for predicting the second baking parameter at the next moment in the baking machine based on the first baking parameter; A first instruction generation module for generating a first real-time control instruction and a third baking parameter within a preset time period based on the first baking parameter and the second baking parameter; A second instruction generation module, configured to apply a reinforcement learning network based on the first real-time control instruction and the third baking parameter to generate a second real-time control instruction; A baking control module, configured to control a baking parameter adjustment mechanism in the baking machine based on the first real-time control instruction and the second real-time control instruction.

8. A multi-layer instruction adaptive control system, characterized in that, Comprising: A controller, a baking machine, and a baking parameter acquisition device, where the baking parameter acquisition device is connected to the controller; The baking machine is used to load tea leaves and bake the tea leaves; The baking parameter acquisition device is arranged in the baking machine and is used to acquire first baking parameters in real time; The controller is used to receive the first baking parameters and execute the multi-layer instruction adaptive control method according to any one of claims 1-6.

9. The multi-layer instruction adaptive control system according to claim 8, characterized in that, The baking parameter acquisition device includes at least one of a temperature sensor, a humidity sensor, a carbon monoxide sensor, a carbon dioxide sensor, and an air quality sensor.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the multi-layer instruction adaptive control method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Subway train tracking control method and system based on multi-target DMC prediction control algorithm

    CN111176297A

  • Adjusting method and device of cut tobacco drying system and medium

    CN118805931A

  • Quality prediction method for food and drink using deep learning, and food and drink

    JP2018018354A

  • Method, apparatus and electronic device for constructing reinforcement learning model and medium

    US20210216686A1

  • Manufacturing process control using constrained reinforcement machine learning

    US20210247744A1