Control method and equipment for fly ash autoclaved brick forming process

Through MDP modeling and reinforcement learning algorithm, the parameter changes in fly ash autoclave brick forming process is predicted, and the process parameter drift caused by equipment wear and environmental interference of the PID controller is solved, and the safety and robustness of the intelligent control system is improved.

CN120386178AActive Publication Date: 2025-07-29CHIPING XINYUAN ENVIRONMENTAL PROTECTION BUILDING MATERIALS CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510581248.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-29
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

During the fly ash autoclaved brick forming process, the PID controller drifts in time due to dynamic wear of equipment and environmental interference, resulting in time-varying drifts in the relationship between process parameters and mass, resulting in a lag in pressure adjustment and a longer brick density stability time.

Method used

MDP modeling and reinforcement learning algorithm are used, combined with statistical analysis and Euclidean distance comparison, to predict the change trend of parameters and adjust the PID gain in advance, and to drive the PID controller to adjust the parameter by calculating the parameter difference value to build an intelligent control system.

Benefits of technology

It significantly improves the safety and robustness of the fly ash autoclaved brick forming control system, reduces the gain drift of the PID controller, and improves the safety of the production process and fault positioning accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386178A_ABST
    Figure CN120386178A_ABST
Patent Text Reader

Abstract

The invention provides a control method and system for a fly ash autoclaved brick forming process, and relates to the technical field of fly ash autoclaved bricks, and the method comprises the steps: firstly, collecting historical parameters of each process step in different batches of forming processes, taking the historical parameters as a state basis for constructing a state space, constructing the state space, and taking the state space as a state basis; adjusting variables corresponding to historical parameters are integrated into an action space, a state transition probability is estimated based on a statistical analysis algorithm, and meanwhile, an MDP model is constructed by combining a green brick qualification rate and an energy consumption target design reward function; parameter values of the current process step are collected in real time and serve as initial states to be input into a reinforcement learning algorithm, an MDP model is solved through iterative optimization, optimal parameters enabling long-term accumulated rewards to be maximized are output, the difference value between the optimal parameters and real-time parameters is calculated, the difference value serves as adjusting input of a PID controller, and the PID controller is controlled to be in a real-time state. Actuating mechanism outputs such as valve opening and pressure set values are dynamically adjusted, and closed-loop control over technological parameters is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of autoclaved fly ash bricks, and particularly to a control method and device for the forming process of autoclaved fly ash bricks. Background Art

[0002] As a new type of wall material, autoclaved fly ash bricks play an important role in the process of building industrialization due to their resource utilization of industrial solid waste (fly ash) and low-carbon environmental protection characteristics. The forming process of autoclaved fly ash bricks involves strong coupling of multiple parameters, including raw material moisture content, mixing time, forming pressure, autoclaving temperature, etc., and the process has non-linear dynamic characteristics. The PID controller adjusts the device parameters through proportional (P), integral (I), and derivative (D) gains to achieve closed-loop control of key parameters such as forming pressure and mixing speed. However, due to the dynamic wear of autoclaved fly ash brick equipment and environmental interference (such as the expansion of die clearance and the decline of autoclave sealing performance), the relationship between process parameters and quality will change over time, and it is necessary to dynamically adjust the gains of the PID controller to maintain control performance.

[0003] For example, in terms of forming pressure control, the PID controller takes the deviation of brick blank density as the input and dynamically adjusts the pressure. The proportional term quickly responds to the density deviation, the integral term eliminates the steady-state error, and the derivative term suppresses overshoot. However, the dynamic wear and environmental interference of autoclaved fly ash brick equipment, such as the expansion of die clearance and the decline of autoclave sealing performance, will cause the relationship between process parameters and quality to change over time, and then it is necessary to dynamically adjust the PID controller gain to maintain control performance. However, the derivative term of the PID controller has a response delay to changes in die clearance or autoclave sealing performance, which will cause pressure regulation to lag and the time for the brick blank density to reach a stable value to be prolonged. Summary of the Invention

[0004] To predict the change trend of parameters such as die clearance and adjust the PID gain in advance, this application provides a control method and device for the forming process of autoclaved fly ash bricks.

[0005] In the first aspect, this application provides a control method for the forming process of autoclaved fly ash bricks, adopting the following technical solution: A control method for the forming process of autoclaved fly ash bricks includes the following steps: Modeling: Obtain the historical parameters of each process step in different autoclaved fly ash brick forming processes. Take the historical parameters of each process step in the same autoclaved fly ash brick forming process as states, integrate all states into a state space, take the adjustment amounts of each historical parameter as actions, integrate all actions into an action space, set the state transition probability based on all historical parameters using a statistical analysis algorithm, set the reward function according to the qualified rate and energy consumption of brick blanks, and establish an MDP model; Solution: Obtain the real-time parameters of the current process step, solve the MDP model using the reinforcement learning algorithm based on the real-time parameters, obtain the first parameter that maximizes the value of the reward function, and calculate the difference between the first parameter and the real-time parameters; Control: Adjust the real-time parameters by controlling the PID controller according to the difference.

[0006] This application realizes the intelligent control of various parameters in the autoclaved brick forming process of fly ash through MDP modeling and reinforcement learning. This application first takes the historical parameters (such as pressure and temperature) of each process step as states and the parameter adjustment amount as actions, implicitly captures the process dynamic law by combining statistical analysis algorithms (such as Markov chain), and guides the optimization through a multi-objective reward function (taking into account both the qualification rate and energy consumption); Subsequently, select the reinforcement learning algorithm (such as DDPG / DQN) according to the characteristics of the state space (continuous / discrete), input the real-time parameters into the MDP model to output the predicted optimal control strategy (i.e., the first parameter); Finally, drive the PID controller to adjust by calculating the difference between the first parameter and the real-time parameters. This application can predict the difference of each parameter based on historical data and real-time feedback, and control the PID controller to adjust the parameters in advance according to the difference to continuously improve the control strategy and reduce the drift problem of the PID controller gain caused by non-linearity.

[0007] Optionally, the method further includes: Calculate the Euclidean distance between the real-time parameters and the first parameter, denoted as the first data; obtain the second parameter that minimizes the value of the reward function, calculate the Euclidean distance between the real-time parameters and the second parameter, denoted as the second data; Judge whether the first data is greater than the second data. If so, execute the control step; if not, output an alarm signal.

[0008] This application significantly enhances the safety and robustness of the autoclaved brick forming control system of fly ash by introducing the Euclidean distance comparison and alarm decision mechanism between real-time parameters and optimal / worst parameters: By calculating the distance difference between the current parameters and the optimal solution of reinforcement learning (the first parameter) and the worst solution of the reward function (the second parameter), this application can indirectly judge whether the working condition requires an alarm, and trigger an alarm when deviating from the optimal solution and deteriorating significantly to reduce the occurrence of wrong control.

[0009] Optionally, before executing the control step, the method further includes: Calculate the attenuation rate: Calculate the attenuation rate of the reward function and set the normal value range of the attenuation rate, judge whether the attenuation rate at the current moment is within the normal value range. If so, execute the control step; if not, output a warning signal; The attenuation rate at the current moment The calculation model is as follows: ; Among them, is the value of the reward function at the current moment; is the average value of the reward function over a preset time period before the current moment.

[0010] This application calculates the attenuation rate, which characterizes the relative performance change of the current control strategy relative to the historical average level. If the attenuation rate is within the preset normal range, it indicates that the control strategy is effective, and the control steps are continued to adjust the real-time parameters. If the attenuation rate exceeds the normal value range, a warning signal is triggered, indicating that the MDP model may fail or the process environment has changed suddenly. By quantifying the performance change of the MDP model, this application significantly improves the robustness of the MDP model.

[0011] Optionally, after performing the control steps, the method further includes: Count the number of warning signals issued within the time period T after adjusting the real-time parameters, denoted as the third data; respectively count the number of warning signals issued within n time periods T before adjusting the real-time parameters, denoted as the fourth data, and use rule to determine whether the third data is normal. If so, no processing is performed; if not, a warning signal is output.

[0012] By comparing the frequency changes of the warning signals before and after adjustment, this application quantitatively evaluates the effectiveness of the adjustment action and reduces the dependence on the single adjustment result. This application adds a secondary warning mechanism on the basis of the original parameter adjustment to form a closed loop of "adjustment - evaluation - re-warning", which is especially suitable for high-risk processes and improves the safety of the production process.

[0013] Optionally, the third data and the fourth data are the number of warning times in the same process step during the autoclaved brick forming process of fly ash.

[0014] By adopting the above technical solution, this application can compare the number of warning times before and after adjustment within the same process step, and thus can immediately evaluate the local effect of the adjustment action, avoid the risk diffusion caused by the global statistical delay as much as possible, and improve the fault location accuracy and risk intervention timeliness.

[0015] Optionally, the method further includes: Decouple the MDP model into multiple MDP sub-models, each MDP sub-model corresponding to a different process step in the autoclaved brick forming process of fly ash. Use the Monte Carlo tree search algorithm to select the MDP sub-model to be activated, integrate the MDP sub-models to be activated in the order of appearance in the autoclaved brick forming process of fly ash into a new MDP model, and perform the solution steps.

[0016] This application decomposes the complex control tasks in the autoclaved fly ash brick forming process (such as raw material mixing, pressing forming, steam curing, etc.) into multiple independent MDP sub-models. Each MDP sub-model corresponds to a single process step and includes a state space (such as humidity, pressure, temperature), an action space (such as parameter adjustment), and a reward function (such as quality qualification rate, energy consumption). The MDP sub-model of this application only focuses on local process constraints (such as the dynamic of the hydraulic system in the pressing step), which can reduce the situation of global state space explosion (such as the combined state dimension of the mixing + pressing + drying steps being too high). Subsequently, this application uses the Monte Carlo tree search algorithm to dynamically select the sub-model to be activated according to the current process state (such as humidity deviation in the mixing step, pressure fluctuation in the pressing step) (such as preferentially activating the humidity control sub-model rather than the steam curing sub-model), preferentially activating the sub-model that has the greatest impact on the current process state (such as preferentially activating the mixing step control sub-model when humidity is abnormal), and reducing ineffective calculations. Subsequently, this application reorganizes the activated sub-models into a new MDP model according to the physical time sequence of the autoclaved fly ash brick forming process (such as mixing → pressing → drying), so that the control logic meets the requirements of the process flow. Perform reinforcement learning to solve the reorganized MDP model (such as policy gradient method, Q-learning) to generate an optimal control strategy.

[0017] Optionally, the method further includes: formulating multiple adjustment strategies according to the difference, optimizing the adjustment strategies according to the genetic algorithm, constructing a fitness function of the genetic algorithm based on the autoclaved fly ash brick forming efficiency, the qualified rate of brick blanks, and the energy consumption, and iteratively screening the adjustment strategy that maximizes the fitness function value, which is denoted as the optimal condition strategy; the adjustment strategy includes: the adjustment amplitude, adjustment order, and adjustment period corresponding to the PID controller, and the magnitude of the adjustment amplitude is equal to the magnitude of the difference; In the control step, based on the optimal adjustment strategy, control the corresponding PID controller to adjust the real-time parameters.

[0018] This application can generate a multi-dimensional adjustment strategy based on the deviation between real-time monitoring parameters (such as humidity, pressure, temperature) and target values (such as humidity difference ΔH = target humidity - actual humidity), covering the adjustment amplitude, adjustment sequence, and adjustment period of the PID controller. Subsequently, by combining adjustment parameters of different dimensions (such as humidity adjustment amplitude ±5%, pressure adjustment sequence advanced, adjustment period shortened to 5 seconds), a strategy pool is formed. This application can iteratively optimize the strategy pool through selection, crossover, and mutation, retain high-fitness strategies, and eliminate inefficient strategies. After multiple generations of genetic iteration, the strategy with the largest fitness function value is selected as the optimal condition strategy (such as humidity adjustment amplitude ±3%, pressure adjustment sequence backward, adjustment period 8 seconds), and the adjustment amplitude, sequence, and period parameters of the optimal strategy are written into the PID controller to adjust the actuator (such as variable-frequency motor, steam valve) in real time, forming a closed-loop control of "deviation → strategy → PID → execution".

[0019] Optionally, the control step further includes: During the process of adjusting real-time parameters, the energy consumption of each PID controller is detected in real time. If the energy consumption of any PID controller exceeds the preset energy consumption threshold, an alarm signal is issued; otherwise, this step is executed again.

[0020] By adopting the above technical solution, this application can give early warning of potential faults and reduce unplanned shutdowns when the energy consumption of the hydraulic system or motor increases due to aging.

[0021] Optionally, the method further includes: obtaining the adjustment step size of the process step corresponding to the real-time parameter by using a statistical analysis algorithm based on historical parameters; In the control step, based on the optimal adjustment strategy, the corresponding PID controller is controlled to adjust the real-time parameter according to the adjustment step size.

[0022] By analyzing historical parameters, this application can obtain the adjustment step sizes suitable for different process steps, reduce overshoot or oscillation caused by a fixed step size. In scenarios of raw material fluctuations or equipment noise, the adjustment step size can dynamically adapt to the parameter change rate and reduce the overshoot amount of PID control. By limiting the adjustment step size, energy consumption waste caused by high-frequency large-amplitude adjustment is reduced, and the adjustment amplitude and step size in the optimal adjustment strategy complement each other to balance the adjustment speed and stability.

[0023] In a second aspect, this application provides a control device for the autoclaved brick forming process of fly ash, adopting the following technical solution: A control device for the autoclaved brick forming process of fly ash, the control device is used to execute the control method, and the control device further includes: The modeling module is used to obtain the historical parameters of each process step in the molding process of different fly ash autoclaved bricks, use the historical parameters of each process step in the same fly ash autoclaved brick molding process as states, integrate all states into a state space, use the adjustment amount of each historical parameter as an action, and integrate all actions into an action space. Based on all historical parameters, a statistical analysis algorithm is used to set the state transition probability, and a reward function is set according to the brick pass rate and energy consumption to establish an MDP model. A solution module is used to obtain real-time parameters of the current process step, solve the MDP model based on the real-time parameters using a reinforcement learning algorithm, obtain a first parameter that maximizes the value of the reward function, and calculate the difference between the first parameter and the real-time parameter; The control module is used to control the PID controller to adjust the real-time parameters according to the difference.

[0024] In summary, the present application includes at least one of the following beneficial technical effects: 1. This application realizes intelligent control of various parameters in the fly ash autoclaved brick forming process through MDP modeling and reinforcement learning. This application first takes the historical parameters of each process step (such as pressure and temperature) as the state and the parameter adjustment amount as the action, combines statistical analysis algorithms (such as Markov chains) to implicitly capture the dynamic laws of the process, and guides the optimization through a multi-objective reward function (taking into account both the pass rate and energy consumption); then, based on the state space characteristics (continuous / discrete), a reinforcement learning algorithm (such as DDPG / DQN) is selected, and the optimal control strategy (i.e., the first parameter) predicted by the MDP model output is input using real-time parameters; finally, the PID controller is driven to adjust by calculating the difference between the first parameter and the real-time parameter. This application can predict the difference between each parameter based on historical data and real-time feedback, and control the PID controller in advance to adjust the parameters according to the difference, so as to continuously improve the control strategy and reduce the drift problem of the PID controller gain caused by nonlinearity.

[0025] 2. This application significantly enhances the safety and robustness of the fly ash autoclaved brick forming control system by introducing the Euclidean distance comparison and alarm decision mechanism between real-time parameters and optimal / worst parameters: by calculating the distance difference between the current parameters and the optimal solution of reinforcement learning (first parameter) and the worst solution of the reward function (second parameter), this application can indirectly determine whether the working conditions require an alarm, and trigger an alarm when they deviate from the optimal solution and deteriorate significantly to reduce the occurrence of erroneous control. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a flowchart from S1 modeling to S3 control in Example 1 of the present application. DETAILED DESCRIPTION

[0027] The present application is further described in detail below with reference to the accompanying drawings.

[0028] Example 1: This example discloses a control method for the autoclaved brick forming process of fly ash. Refer to Figure 1 , the control method includes: S1 Modeling, S2 Solving, and S3 Control. First, collect the historical parameters of each process step during the forming process of different batches, use them as states to construct a state space, and integrate the adjustment amounts of the corresponding historical parameters into an action space. Estimate the state transition probability based on a statistical analysis algorithm, and at the same time design a reward function in combination with the qualified rate of brick blanks and the energy consumption target to construct a Markov decision process (MDP) model; collect the parameter values of the current process step in real time, use them as the initial state and input them into the reinforcement learning algorithm, solve the MDP model through iterative optimization, output the optimal parameters that maximize the long-term cumulative reward, calculate the difference between them and the real-time parameters, and use this difference as the adjustment input of the PID controller to dynamically adjust the output of actuators such as valve opening and pressure set value to achieve closed-loop control of process parameters. This example includes the following steps: S1 Modeling, collect the historical parameters of each process step in the entire process of autoclaved brick forming of fly ash, including but not limited to: Mixing stage: composition of fly ash, ratio of fly ash to binder, stirring speed, mixing time; Pressing stage: pressing pressure, pressing speed, holding pressure time; Curing stage: steam temperature, steam pressure, curing duration, humidity gradient.

[0029] The sources of the historical parameters include PLC records, sensor logs, quality inspection reports, etc.

[0030] Regard the parameter combination of each step in the forming process of the same batch as a state vector, map all possible state combinations to the state space, regard the adjustment amount of each historical parameter as an action, and integrate all actions into an action space.

[0031] Based on the historical parameter sequence, the following methods can be used in this example to estimate the state transition probability: Frequency statistics method: Count the frequency of the state transferring to the next state after performing an action in the historical data, and the result after normalization is recorded as the probability.

[0032] Hidden Markov model (HMM): If the state cannot be fully observed (such as some parameters are missing), model the potential state transition through HMM.

[0033] Time series analysis: Perform ARIMA modeling on the continuous parameter sequence (such as pressure) to predict its dynamic change law.

[0034] Set up a reward function according to the qualified rate of brick blanks and energy consumption, and establish an MDP model.

[0035] Solve S2, obtain real-time parameters corresponding to the historical parameter categories in S1 modeling, and solve the MDP model using the reinforcement learning algorithm based on the real-time parameters.

[0036] The Markov Decision Process (MDP) model is a mathematical model used to describe problems with uncertainty and sequential decision-making. It consists of State, Action, State Transition Probability, and Reward Function. In the scenario of a process step, the state is the state space formed by the current real-time parameters, the action is the adjustment operation taken on the process parameters, the state transition probability describes the possibility of transitioning to other states after taking a certain action in the current state, and the reward function is used to evaluate the benefit or reward obtained after taking a certain action in a certain state.

[0037] Reinforcement learning is a machine learning method whose purpose is to enable an agent to find the optimal behavioral strategy in an environment through continuous trial and learning to maximize long-term rewards. Common reinforcement learning algorithms include Q-learning, Deep Q-Network (DQN), Policy Gradient algorithm, etc. Based on the real-time parameters, the agent will select an action (adjustment plan for the process parameters) according to the current state (the state formed by the real-time parameters), and then observe the feedback of the environment (the process system), that is, the change of the state and the obtained reward. By continuously iterating this process, the agent gradually learns the optimal strategy, enabling the value of the reward function to be maximized in the MDP model.

[0038] During the learning process of the reinforcement learning algorithm, the agent will continuously try different actions (adjustments to the process parameters) and record the rewards brought by each action. As the learning progresses, the agent will gradually find a set of optimal action sequences, and the process parameter values corresponding to these actions are the parameter values that maximize the value of the reward function, that is, the first parameter.

[0039] After obtaining the first parameter that maximizes the reward function, compare it with the real-time parameters of the current process step. Calculate the difference between each parameter. For example, if the temperature value in the first parameter is T1 and the temperature value in the real-time parameter is T2, then the temperature difference is T1 - T2. By calculating the difference, we can understand the gap between the current process parameters and the optimal parameters, providing a basis for subsequent process adjustment. If the difference is large, it indicates that there is still a large room for optimization in the current process, and the adjustment direction and amplitude of the process parameters can be determined according to the sign and magnitude of the difference to gradually approach the optimal process state.

[0040] S3 control, adjust the real-time parameters by controlling the PID controller according to the difference obtained in S2 solution.

[0041] The following elaborates on the solution provided in this embodiment in combination with specific cases.

[0042] S1 Modeling Historical parameter collection: In the entire process of autoclaved brick forming from fly ash, various historical parameters are collected from different process steps.

[0043] In the mixing stage, collect the composition of fly ash (such as the content of silicon dioxide, alumina, etc.), the ratio of fly ash to binder (such as a ratio of 10:1), the stirring speed (such as revolutions per minute), and the mixing time (in minutes). In the pressing stage, focus on the pressing pressure (the unit may be megapascals), the pressing speed (the displacement per unit time), and the holding pressure time (in minutes).

[0044] In the curing stage, record the steam temperature (in degrees Celsius), the steam pressure (in megapascals), the curing duration (in hours), and the humidity gradient (the humidity change per meter). These parameters come from a wide range of sources. The PLC (Programmable Logic Controller) records the parameters during the operation of the equipment, the sensor logs detail the real-time data measured by the sensors, and the quality inspection reports contain the parameter information related to product quality.

[0045] Regard the parameter combination of each step in the forming process of the same batch as a state vector. For example, a state vector may be [Fly ash composition: Silicon dioxide 60%, Alumina 20%...; Ratio of fly ash to binder 8:1; Stirring speed 100 revolutions per minute; Mixing time 5 minutes; Pressing pressure 10 megapascals; Pressing speed 0.5 millimeters per second; Holding pressure time 3 minutes; Steam temperature 120 degrees Celsius; Steam pressure 0.8 megapascals; Curing duration 10 hours; Humidity gradient 0.5% per meter]. All such possible state combinations constitute the state space. Take the adjustment amount of each historical parameter (such as an increase of 0.1 in the ratio of fly ash to binder, an increase of 1 megapascal in the pressing pressure, etc.) as an action, and all these actions are integrated to form the action space.

[0046] An example of the state transition probability estimation method is as follows: Frequency statistics method: By counting the frequency of a certain state transferring to the next state after performing a certain action in historical data, and then performing normalization processing to obtain the state transition probability. For example, in 100 historical records, state A transfers to state B 20 times after performing action a, then the probability of state A transferring to state B by performing action a is 20÷100 = 0.2.

[0047] Hidden Markov Model (HMM): When there are missing partial parameters resulting in incomplete observability of states, HMM is used to model potential state transitions. For example, in some records, the humidity gradient parameter in the curing stage is missing, but the state transitions can be inferred by HMM based on other observable parameters (such as steam temperature, pressure, etc.).

[0048] Time series analysis: For a continuous parameter sequence like pressure, ARIMA (Autoregressive Integrated Moving Average Model) is used for modeling to predict its dynamic change pattern. For example, by performing ARIMA modeling on historical pressing pressure data, the change trend of pressure in future pressing processes can be predicted.

[0049] Set up a reward function based on the qualified rate and energy consumption of the brick blanks. If the qualified rate of the brick blanks is high and the energy consumption is low, a higher reward value is given; otherwise, a lower reward value is given. For example, when the qualified rate of the brick blanks reaches over 90% and the energy consumption is lower than a certain standard, the reward value is 10; if the qualified rate is lower than 80% or the energy consumption is too high, the reward value is -5. Combining the state, action, state transition probability, and reward function, a Markov Decision Process (MDP) model is established.

[0050] S2 Solving Obtain real-time parameters corresponding to the historical parameter categories in S1 modeling. These real-time parameters reflect the actual state of the current autoclaved fly ash brick forming process. For example, the real-time fly ash composition in the current mixing stage, the real-time ratio of fly ash to binder, etc.

[0051] Use common reinforcement learning algorithms such as Q-learning, DQN, or policy gradient algorithms, etc., to solve the MDP model based on the real-time parameters. The agent selects an action (such as adjusting the pressing pressure) from the action space according to the state composed of the current real-time parameters, and then observes the feedback of the process system, that is, the change of the state (such as the change of the density of the brick blank after pressing) and the obtained reward (calculate the reward value according to the qualified rate and energy consumption of the brick blank in the new state). By continuously repeating this process of selecting actions, observing feedback, and learning, the agent gradually finds a set of optimal action sequences, and the process parameter values corresponding to these actions are the first parameters that maximize the value of the reward function. For example, after multiple iterative learning, it is found that when the parameter combination such as the pressing pressure is 12 MPa and the pressing speed is 0.4 mm / s, the value of the reward function is the largest.

[0052] Compare the first parameters that maximize the reward function with the current real-time parameters. Calculate the difference between each parameter. For example, if the pressing pressure in the first parameter is 12 MPa and the pressing speed is 0.4 mm / s, and the pressing pressure in the real-time parameters is 10 MPa and the pressing speed is 0.4 mm / s, then the difference in pressing pressure is 2 MPa.

[0053] S3 control, based on the difference obtained in the solution of S2, controls the PID (Proportional-Integral-Derivative) controller to adjust the real-time parameters.

[0054] The PID controller, based on the magnitude and change trend of the difference, etc., outputs a suitable control signal to adjust the process parameters through proportional, integral, and derivative operations. For example, if the difference in the pressing pressure is positive, it indicates that the real-time pressing pressure is lower than the optimal pressing pressure. The PID controller will output a control signal to increase the pressing pressure, gradually making the real-time parameters approach the optimal parameters, achieving optimized control of the autoclaved fly ash brick forming process.

[0055] In other embodiments, the method further includes: Calculating the Euclidean distance between the real-time parameter and the first parameter, denoted as the first data; through reverse exploration (such as modifying the reward function to a negative incentive) or sampling strategy, obtaining the second parameter that minimizes the value of the reward function, and calculating the Euclidean distance between the real-time parameter and the second parameter, denoted as the second data.

[0056] Judging whether the first data is greater than the second data. If so, execute S3 control; if not, output an alarm signal.

[0057] In other embodiments, although the real-time parameter is close to the positive ideal parameter, but the attenuation rate of the reward function is very large. At this time, it is determined that there is a phenomenon of process mismatch. Therefore, before executing S3 control, the method further includes: S4 calculates the attenuation rate, calculates the attenuation rate of the reward function and sets the normal value range of the attenuation rate, and judges whether the attenuation rate at the current moment is within the normal value range. If so, execute S3 control; if not, output a warning signal.

[0058] The attenuation rate at the current moment The calculation model is as follows: ; Wherein, is the value of the reward function at the current moment; is the average value of the reward function at a preset time duration before the current moment.

[0059] A positive attenuation rate means that the current reward is lower than the historical average, indicating a decline in process performance (such as a decrease in the qualified rate or an increase in energy consumption); a negative attenuation rate means that the current reward is higher than the historical average, indicating an improvement in process performance; a zero attenuation rate means that the current reward is equal to the historical average.

[0060] After executing S3 control, the method further includes: After statistically adjusting the real-time parameters, count the number of warning signals sent within the duration T (the duration T is a time interval set according to the production process characteristics and actual requirements of autoclaved fly ash bricks. For example, considering that it may take some time for the impact of process parameter adjustment on the production process to become apparent, T can be set to 30 minutes, 1 hour, etc. The length of T should be able to reasonably reflect the change in the process state after parameter adjustment), which is recorded as the third data; maintain a historical warning queue with a length of n×T. When a new duration T ends, add the data of the number of warning signals within this duration T to the end of the queue, and at the same time remove the first data in the queue (i.e., the warning count data within the earliest duration T). In this way, the queue always maintains the latest information on the number of warning signals within the latest n duration Ts. For example, initially, the warning counts within the past 3 one-hour periods in the queue are 2, 1, and 3 respectively. When a new one-hour period has passed and the warning count within this one-hour period is 4, then the queue is updated to 1, 3, 4.

[0061] Respectively count the number of warning signals sent within the previous n duration Ts before each adjustment of the real-time parameters, which is recorded as the fourth data. The third data and the fourth data are the warning counts in the same process step during the autoclaved fly ash brick forming process.

[0062] Based on the fourth data, use rules to determine whether the third data is normal. If it is, no processing is performed; if not, a warning signal is output.

[0063] Embodiment 2: The difference between this embodiment and Embodiment 1 is that the method further includes: Decouple the MDP model into multiple MDP sub-models, and each MDP sub-model corresponds to a different process step in the autoclaved fly ash brick forming process. According to the technological process of autoclaved fly ash brick forming, clearly divide the process steps corresponding to each MDP sub-model. For example, take all relevant parameters in the mixing stage (such as the composition of fly ash, the ratio of fly ash to binder, stirring speed, mixing time, etc.) as the state variables of an MDP sub-model, and take the operations that can be performed in this stage (such as adjusting the ratio, changing the stirring speed, etc.) as the action variables to construct the MDP sub-model of the mixing stage; similarly, construct the MDP sub-models of the pressing stage and the curing stage respectively.

[0064] After that, use the Monte Carlo tree search algorithm to select the MDP sub-model to be activated, including: Selection stage: Based on the UCT (Upper Confidence Bound for Trees) algorithm, preferentially explore high-value nodes (such as the steam pressure volatility in the curing link); Expansion stage: When the state of the sub-model deviates from the expectation (such as the deviation of the brick blank density > 3%), dynamically expand the sub-model; Simulation stage: Evaluate the cumulative rewards of different sub-model combinations through fast Monte Carlo simulation; Backtracking stage: Update the access times and value estimates of each sub-model to provide a basis for decision-making at the next moment.

[0065] Connect the activated sub-models in series in the order of MDP1→MDP2→MDP3→MDP4 to form a global MDP model, and execute S2 for solution.

[0066] Formulate multiple adjustment strategies according to the difference, optimize the adjustment strategies according to the genetic algorithm, construct a fitness function of the genetic algorithm based on the forming efficiency, brick blank qualification rate and energy consumption of autoclaved fly ash bricks, and iteratively screen the adjustment strategy that maximizes the fitness function value, which is recorded as the optimal condition strategy; the adjustment strategies include: the adjustment amplitude, adjustment order and adjustment period corresponding to the PID controller, and the magnitude of the adjustment amplitude is equal to the magnitude of the difference.

[0067] Based on historical parameters, use a statistical analysis algorithm to obtain the adjustment step size of the process step corresponding to the real-time parameters.

[0068] In S3 control, based on the optimal adjustment strategy, control the corresponding PID controller to adjust the real-time parameters according to the adjustment step size.

[0069] In other embodiments, the S3 control further includes: during the process of adjusting the real-time parameters, detect the energy consumption of each PID controller in real time. If the energy consumption of any PID controller exceeds the preset energy consumption threshold, an alarm signal is sent, otherwise, this step is re-executed.

[0070] In this embodiment, through hierarchical MDP modeling and Monte Carlo tree search (MCTS) dynamic decision-making, the forming process of autoclaved fly ash bricks is decoupled into 4 independent sub-models: raw material pretreatment, forming pressing, steam curing, and finished product inspection. They are dynamically activated and integrated into a global MDP model according to the process time sequence. An initial adjustment strategy is generated based on the target-actual difference. The genetic algorithm is used to optimize the strategy parameters (such as PID adjustment amplitude, order, and period) with the forming efficiency, brick blank qualification rate, and unit energy consumption as the fitness function. The real-time adjustment step size is determined by combining historical data statistical analysis. The optimal strategy drives the PID controller to achieve closed-loop parameter adjustment. At the same time, a real-time energy consumption monitoring mechanism is introduced to trigger an alarm and roll back the control for any PID energy consumption exceeding the limit, and finally form an adaptive intelligent control system covering the entire process. Embodiment 3: This embodiment discloses a control device for the forming process of autoclaved fly ash bricks, and the control device includes: A modeling module, which is used to obtain the historical parameters of each process step in the autoclaved brick forming process of different fly ashes, take the historical parameters of each process step in the autoclaved brick forming process of the same fly ash as states, integrate all states into a state space, take the adjustment amounts of each historical parameter as actions, integrate all actions into an action space, set state transition probabilities based on all historical parameters using a statistical analysis algorithm, set a reward function according to the qualified rate and energy consumption of the brick blank, and establish an MDP model.

[0071] A solving module, which is used to obtain the real-time parameters of the current process step, solve the MDP model using a reinforcement learning algorithm based on the real-time parameters, obtain the first parameter that maximizes the value of the reward function, and calculate the difference between the first parameter and the real-time parameters.

[0072] A control module, which is used to control a PID controller to adjust the real-time parameters according to the difference.

[0073] The above are all the preferred embodiments of the present application. The protection scope of the present application is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present application should be covered within the protection scope of the present application.

Claims

1. A control method for the autoclaved brick forming process using fly ash, characterized in that, Including: Modeling: Obtain the historical parameters of each process step during the autoclaved brick forming process of different fly ashes. Take the historical parameters of each process step during the autoclaved brick forming process of the same fly ash as states, integrate all states into a state space, take the adjustment amount of each historical parameter as actions, integrate all actions into an action space, set the state transition probability based on all historical parameters using a statistical analysis algorithm, set a reward function according to the qualified rate and energy consumption of the brick blank, and establish an MDP model. Solving: Obtain the real-time parameters of the current process step, solve the MDP model using a reinforcement learning algorithm based on the real-time parameters, obtain the first parameter that maximizes the value of the reward function, and calculate the difference between the first parameter and the real-time parameters. Controlling: Control the PID controller to adjust the real-time parameters according to the difference.

2. The control method for the autoclaved brick forming process of fly ash according to claim 1, characterized in that, The method further includes: Calculate the Euclidean distance between the real-time parameters and the first parameter, denoted as the first data; obtain the second parameter that minimizes the value of the reward function, calculate the Euclidean distance between the real-time parameters and the second parameter, denoted as the second data. Judge whether the first data is greater than the second data. If so, execute the control step; if not, output an alarm signal.

3. The control method for the autoclaved brick forming process of fly ash according to claim 2, characterized in that, Before executing the control step, the method further includes: Calculate the attenuation rate: Calculate the attenuation rate of the reward function and set the normal value range of the attenuation rate, judge whether the attenuation rate at the current moment is within the normal value range. If so, execute the control step; if not, output a warning signal. The attenuation rate at the current moment has the following calculation model: ; Among them, is the value of the reward function at the current moment; is the average value of the reward function over a preset duration before the current moment.

4. The control method for the autoclaved brick forming process of fly ash according to claim 3, characterized in that, After executing the control step, the method further includes: Count the number of warning signals sent within the duration T after statistically adjusting the real-time parameters, which is recorded as the third data; respectively count the number of warning signals sent within n durations T before adjusting the real-time parameters, which is recorded as the fourth data, and based on the fourth data, use the rule to judge whether the third data is normal. If it is, no processing is required; if not, a warning signal is output.

5. The control method for the autoclaved brick forming process of fly ash according to claim 4, characterized in that, The third data and the fourth data are the number of warning times in the same process step during the autoclaved brick forming process of fly ash.

6. The control method for the autoclaved brick forming process of fly ash according to claim 1, characterized in that The method further includes: Decouple the MDP model into multiple MDP sub-models, each MDP sub-model corresponding to a different process step during the autoclaved brick forming process of fly ash. Use the Monte Carlo tree search algorithm to select the MDP sub-model to be activated, integrate the MDP sub-models to be activated in the order of appearance during the autoclaved brick forming process of fly ash into a new MDP model, and execute the solving step.

7. The control method for the autoclaved brick forming process of fly ash according to claim 6, characterized in that, The method further includes: Formulate multiple adjustment strategies according to the difference, optimize the adjustment strategies according to the genetic algorithm, construct a fitness function of the genetic algorithm according to the forming efficiency of autoclaved fly ash bricks, the qualified rate of brick blanks, and energy consumption, and iteratively screen out the adjustment strategy that maximizes the fitness function value, denoted as the optimal condition strategy; the adjustment strategy includes: the adjustment amplitude, adjustment order, and adjustment period corresponding to the PID controller, and the magnitude of the adjustment amplitude is equal to the magnitude of the difference. In the control step, control the corresponding PID controller to adjust the real-time parameters based on the optimal adjustment strategy.

8. The control method for the autoclaved brick forming process of fly ash according to claim 7, characterized in that, The control step further includes: During the process of adjusting the real-time parameters, detect the energy consumption of each PID controller in real time. If the energy consumption of any PID controller exceeds the preset energy consumption threshold, send out an alarm signal; otherwise, re-execute this step.

9. The control method for the autoclaved brick forming process of fly ash according to claim 8, characterized in that, The method further includes: Based on the historical parameters, use a statistical analysis algorithm to obtain the adjustment step size of the process step corresponding to the real-time parameters. In the control step, based on the optimal adjustment strategy, the corresponding PID controller is controlled to adjust the real-time parameters according to the adjustment step size.

10. A control device for the autoclaved brick forming process of fly ash, characterized in that, The control device is used to execute the control method according to any one of claims 1-9. The control device further includes: A modeling module, configured to obtain historical parameters of each process step in the autoclaved brick forming process of different fly ashes, use the historical parameters of each process step in the same autoclaved brick forming process as states, integrate all states into a state space, use the adjustment amounts of each historical parameter as actions, integrate all actions into an action space, set state transition probabilities based on all historical parameters using a statistical analysis algorithm, set a reward function according to the qualified rate and energy consumption of the brick blank, and establish an MDP model; A solving module, configured to obtain the real-time parameters of the current process step, solve the MDP model using a reinforcement learning algorithm based on the real-time parameters, obtain a first parameter that maximizes the value of the reward function, and calculate the difference between the first parameter and the real-time parameters; A control module, configured to control the PID controller to adjust the real-time parameters according to the difference.

Citation Information

Patent Citations

  • Air conditioner temperature control method based on deep reinforcement learning

    CN116358114A

  • Cloth setting machine energy-saving optimization method based on artificial intelligence

    CN119668090A

  • Discrete element modeling method and system for compaction behavior of printing material

    CN119673340A

  • Quality monitoring method for thermotechnical control loop of thermal power plant

    CN119828646A

  • Control system, method and equipment for preparing dust reducing agent based on reinforcement learning

    CN119861573A