A flakeboard hot-pressing pressure-maintaining time self-adaptive decision algorithm and system
Patent Information
- Application Number
- CN202610844005.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-09-01
AI Technical Summary
然而,热压保压过程存在严格的质量安全约束,传统强化学习方法以最大化期望回报为目标,难以在优化能耗的同时严格保证质量约束的满足
(1)本发明通过构建约束马尔可夫决策过程模型,将保压决策形式化为在质量违约概率约束下的能耗最小化问题,利用拉格朗日乘子法实现约束与目标的平衡,能够在保证板材质量合格率满足预设严格阈值的前提下,最大化降低过程能耗;
Smart Images

Figure CN122672318A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of particleboard production technology, specifically relating to an adaptive decision-making algorithm and system for hot pressing and holding time of particleboard. Background Technology
[0002] Hot pressing is a crucial step in particleboard production. During this process, the adhesive in the board must be fully cured under high temperature and pressure to achieve the specified physical and mechanical properties. The holding time is one of the core process parameters, directly affecting the product's internal bond strength, thickness deviation, and other quality indicators, as well as production energy consumption.
[0003] Currently, determining the holding time in industrial production mainly relies on fixed process procedures or the experience and judgment of operators. Fixed process procedures cannot adapt to changes in raw material characteristics and environmental conditions, such as fluctuations in moisture content and differences in rubber types. They often use the most conservative parameter settings to ensure quality, leading to over-holding and energy waste. Manual experience-based judgment, on the other hand, suffers from strong subjectivity and poor consistency.
[0004] Reinforcement learning, as a data-driven adaptive decision-making method, has shown promising application prospects in sequential decision-making problems. However, the hot-pressing and holding process has strict quality and safety constraints. Traditional reinforcement learning methods, which aim to maximize expected returns, struggle to ensure that quality constraints are strictly met while optimizing energy consumption.
[0005] Therefore, there is an urgent need to design an adaptive decision algorithm for the hot pressing and holding time of particleboard that can optimize energy consumption while strictly limiting the risk of quality default to solve the current technical problems. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides an adaptive decision-making algorithm and system for hot pressing and holding time of particleboard that can optimize energy consumption while strictly limiting the risk of quality breach.
[0007] The technical solution of this invention is: an adaptive decision algorithm for hot pressing and holding time of particleboard, comprising the following steps: S1: Obtain the state information during the hot pressing process, including at least the slab temperature, pressure, thickness, holding time, and raw material characteristic parameters, to form a state vector; S2: Construct a constrained Markov decision process model, determine that the action space consists of continuing to hold pressure and stopping to hold pressure, and set a discount factor. Reward function The cost function is set based on process energy consumption. The system is designed based on events of non-compliance with board quality standards, and a maximum allowable expected probability threshold for default is set. ; S3: Process the state vector using a distributed deep Q-network to obtain the future reward quantile distribution and future cost quantile distribution for each action; S4: Based on the cost quantile distribution and the maximum allowed expected default probability threshold, filter out actions that cause the expected default risk to exceed the maximum allowed expected default probability threshold to obtain a set of allowed actions; S5: In the set of allowed actions, utilize the reward quantile distribution and the current Lagrange multiplier. Calculate the overall action value and select the action to be executed accordingly; S6: Execute the selected action and obtain the next state. Update the distributed deep Q-network using the generated energy reward and quality default cost, and update the Lagrange multipliers using the cumulative default cost of the holding pressure process. ; S7: Using the updated state as the new current state, repeat S3~S6 and update the network and Lagrange multipliers. The process continues until the convergence condition is met, resulting in a well-trained distributed deep Q-network and Lagrange multipliers. ; S8: Input the real-time acquired state vector into the trained distributed deep Q-network, and determine the output based on the quantile distribution and the Lagrange multipliers. Make a decision to maintain pressure or stop.
[0008] Furthermore, the raw material characteristic parameters include the moisture content of the substrate, the curing kinetic parameters of the adhesive, the target density, and the estimated degree of curing calculated online based on the temperature history and the adhesive kinetic model.
[0009] Furthermore, the cost function is: After the pressure holding is stopped, if the internal bond strength of the sheet is lower than the qualified threshold or the thickness deviation exceeds the allowable range, a cost value of 1 is assigned; otherwise, the cost value is 0. The cost value of each process step is 0.
[0010] Furthermore, the distributed deep Q-network includes a shared feature extraction network, a reward distribution output branch, and a cost distribution output branch, wherein the reward distribution output branch and the cost distribution output branch each output a preset number of quantiles.
[0011] Furthermore, the filtering in S4 includes the following sub-steps: S41: Take the confidence level of the cost quantile distribution as... The corresponding upper quantile value; S42: If the upper quantile value is greater than 0, then the expected default risk of the action is determined to exceed the maximum permissible expected default probability threshold. This action will be disabled.
[0012] Furthermore, the calculation of the comprehensive action value in S5 includes the following sub-steps: S511: Calculate the expected value of the quantile distribution of returns; S512: Calculate the expected value of the cost quantile distribution; S513: Subtract the current Lagrange multiplier from the expected value of the said return quantile distribution. The difference between the product of the cost quantile distribution and the expected value is used as the action value.
[0013] Furthermore, during the training of the distributed deep Q-network, step S5 selects the action to be executed using a masked ε-greedy strategy, including the following sub-steps: S521: Based on probability Select the action with the highest overall action value from the set of allowed actions; S522: Based on probability A random action is generated. If the action does not belong to the set of allowed actions, a new action is randomly selected from the set of allowed actions. S523: When the set of allowed actions is empty, force the selection to continue holding pressure.
[0014] Furthermore, in step S6, the cumulative default cost of the pressure holding process is used to update the Lagrange multiplier. Includes the following sub-steps: S61: Calculate the discounted total cost of a complete pressure holding process. Assume the process takes a total of Step, cost of each step is The total cost of the discount is Since the cost of each process step is 0, the total cost of this discount is equal to , The cost value for the termination step; S62: Calculate update amount in The learning rate of the Lagrange multipliers; S63: Set the current Lagrange multiplier Updated to And the updated Lagrange multipliers Apply a preset upper limit and exponential moving average smoothing.
[0015] Furthermore, the maximum allowable expected probability of default threshold is set to 0.005.
[0016] An adaptive decision-making system for hot pressing and holding time of particleboard includes: The data acquisition module is used to acquire the status information during the hot pressing process; The constraint reinforcement learning decision module has a built-in distributed deep Q-network trained according to the adaptive decision algorithm described above, which is used to output a pressure holding or stop action based on the real-time state vector. The security filtering module filters out high-risk actions based on the cost quantile distribution and the maximum allowable expected default probability threshold. Lagrange multiplier storage module, for storing the Lagrange multipliers; The training and update module is used to update the distributed deep Q network and the Lagrange multipliers offline or online using production data.
[0017] The beneficial effects of this invention are: (1) This invention constructs a constrained Markov decision process model, formalizes the pressure holding decision into an energy consumption minimization problem under the constraint of quality default probability, and uses the Lagrange multiplier method to achieve a balance between constraints and objectives, which can maximize the reduction of process energy consumption while ensuring that the board quality pass rate meets the preset strict threshold. (2) This invention uses a distributed deep Q network to output the quantile distribution of returns and costs, instead of just outputting the expected value. By accurately characterizing the tail risk of the cost quantile distribution, it achieves fine filtering of the default risk of actions, ensuring that the expected default probability of the selected action does not exceed the allowable threshold. (3) This invention uses the adaptive update mechanism of Lagrange multipliers to enable the algorithm to automatically adjust the penalty weight of quality default cost during the training process, avoiding the difficulty of manual parameter tuning, and making the algorithm have good adaptability to different raw material conditions and different equipment states. (4) The present invention deploys the trained network and multipliers in a real-time decision system, which can output pressure holding or stop decisions online based on the real-time collected status information, thereby realizing adaptive and precise control of the hot pressing process. Attached Figure Description
[0018] Figure 1 This is a flowchart of the adaptive decision-making algorithm for hot pressing and holding time of particleboard in this invention.
[0019] Figure 2 This is a flowchart of the filtering sub-step in step S2 of the present invention.
[0020] Figure 3 This is a flowchart of the sub-step in step S5 of the present invention, which calculates the comprehensive action value for each action in the filtered set of allowed actions.
[0021] Figure 4 This is a flowchart of the shielded ε-greedy strategy sub-step in step S5 of the present invention. Detailed Implementation
[0022] Various exemplary embodiments of the invention will now be described in detail with reference to the accompanying drawings. The descriptions of the exemplary embodiments are merely illustrative and are in no way intended to limit the invention or its application or use. The invention can be embodied in many different forms and is not limited to the embodiments described herein. These embodiments are provided to make the invention thorough and complete, and to fully express the scope of the invention to those skilled in the art. It should be noted that, unless otherwise specifically stated, the relative arrangement of components and steps, the composition of materials, numerical expressions, and values set forth in these embodiments should be interpreted as merely exemplary and not as limiting.
[0023] The terms "first," "second," and similar words used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different parts. Words such as "including" or "comprising" mean that the element preceding the word encompasses the element listed after it, without excluding the possibility of encompassing other elements. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0024] like Figure 1 As shown, an adaptive decision algorithm for hot pressing and holding time of particleboard is disclosed, which is implemented by steps S1 to S8.
[0025] In step S1, the state information during the hot pressing process is obtained, including at least the slab temperature, pressure, thickness, holding time, and raw material characteristic parameters, which constitute a state vector.
[0026] In the specific implementation process, the core layer temperature of the slab is collected in real time by a temperature sensor, the current pressure value applied is obtained by a hot press, the current thickness of the slab is obtained by a displacement sensor, and the cumulative time from the start of the pressure holding stage is the pressure holding time.
[0027] The raw material characteristic parameters include the moisture content of the substrate, the curing kinetic parameters of the adhesive, the target density, and the estimated degree of curing calculated online based on the temperature history and the adhesive kinetic model.
[0028] The moisture content of the board is obtained through an online moisture content meter; the curing kinetic parameters of the adhesive, including the activation energy and frequency factor of the curing reaction, are pre-stored in the system database; the target density is the target density setting value for this batch of boards; the degree of curing estimate can be calculated online based on the temperature history and the adhesive kinetic model, and is used to characterize the current degree of curing of the adhesive.
[0029] In one example of online calculation of the degree of cure estimate based on temperature history and a kinetic model of the adhesive, the kinetic model adopts the autocatalytic model commonly used in engineering:
[0030] in, The degree of curing (value range 0~1); Time (s); The absolute temperature of the core layer of the slab (K); and This refers to the reaction order, which is usually... ; Let be the reaction rate constant, which obeys the Arrhenius equation:
[0031] in, Frequency factor ; The activation energy of the curing reaction (J / mol); is the ideal gas constant.
[0032] For the adhesives used, the above parameters are calibrated using differential scanning calorimetry (DSC) and pre-stored in the system database. Multiple sets of parameters can be pre-stored for different adhesive types, and the system retrieves the corresponding parameters according to the current production task.
[0033] Thermo-pressure control system with fixed decision cycle s performs the following steps: Read the core temperature of the slab at the current decision moment ; Get the degree of curing saved at the previous moment Initial time hour, Retrieve a pre-stored value, such as 0.05; Calculate the reaction rate constant at the current temperature using the Arrhenius equation described above. ; Will , , , Substitute the above autocatalytic model into the equation and calculate the current reaction rate. ; The degree of curing was updated by discretizing the first-order Euler method:
[0034] Will Limited to Interval: ; The calculated This value serves as an estimate of the solidification degree of the current state vector and is saved for use in the next cycle.
[0035] In step S2, a constrained Markov decision process model is constructed, the action space is determined to consist of continuing to hold pressure and stopping to hold pressure, and a discount factor is set. Reward function The cost function is set based on process energy consumption. The system is designed based on events of non-compliance with board quality standards, and a maximum allowable expected probability threshold for default is set. .
[0036] The pressure holding time decision problem is modeled as a constrained Markov decision process (CMDP). It includes two discrete actions: continuing pressure holding and stopping pressure holding. Discount factor. It is set to 0.99 to balance the importance of current and future gains.
[0037] reward function The settings are based on process energy consumption. Specifically, when choosing to continue holding pressure, the reward value is the negative of the energy consumed at the current time step, that is, the higher the energy consumption, the smaller the reward. This drives the algorithm to learn to stop holding pressure as early as possible while ensuring quality to save energy. When choosing to stop holding pressure, the process enters a terminated state, no longer generates energy consumption, and the subsequent reward value is 0.
[0038] The specific quantification method for energy consumption rewards is as follows: Let the current decision cycle length be... (s), the pressure applied by the hot press is (MPa), current slab thickness is (m), the thickness at the previous moment was Then the change in thickness The mechanical work done by the press on the slab during this cycle is approximately:
[0039] in, The area of the slab (m²) 2 The slab dimensions are pre-stored in the system database.
[0040] Actual power consumption also needs to take into account the compressor efficiency. and standby power :
[0041] The unit is J or kWh. For numerical stability, energy consumption can be normalized, for example, by dividing it by a baseline energy consumption. Receive a dimensionless reward value:
[0042] When the action is to stop holding pressure .
[0043] Baseline energy consumption It can be taken as the average energy consumption of each step in all pressure holding processes in historical data, or set as a fixed value to ensure that the reward value is maintained at a certain level. Magnitude.
[0044] Cost function Based on the setting of sheet material quality non-conformity events, the cost value of all process steps during the entire pressure holding process is 0, except for the final termination state when pressure holding is stopped. After pressure holding stops, the system performs quality inspection on the formed sheet material: if the internal bond strength of the sheet material is lower than the qualified threshold or the thickness deviation exceeds the allowable range, it is judged as quality non-conformity and assigned a cost value of 1; otherwise, the cost value is 0. Through this Bernoulli definition of cost, the expected cost value is directly equal to the probability of quality failure.
[0045] For example, the maximum allowed expected probability of default threshold The value is set to 0.005, which means that the failure rate of the boards should not exceed 0.5% under long-term operation.
[0046] In step S3, the state vector is processed using a distributed deep Q-network to obtain the future reward quantile distribution and the future cost quantile distribution for each action.
[0047] Specifically, the distributed deep Q-network adopts a dual-branch structure, which includes a shared feature extraction network, a reward distribution output branch, and a cost distribution output branch.
[0048] The shared feature extraction network consists of several fully connected layers and is used to extract shared feature representations from the state vector.
[0049] The reward distribution output branch receives shared features, and the output dimension is... A tensor, where |A| is the number of actions. For the preset number of quantiles, each action corresponds to The quantile values represent the estimated quantile distribution of the future discount return for this action; an example of the preset number of quantiles. .
[0050] The cost distribution output branch has the same structure as the reward branch, and outputs the quantile distribution estimate of the future discount cost for each action.
[0051] The network is trained using quantile regression to minimize the quantile Huber loss between the predicted quantile distribution and the target distribution.
[0052] As an example of a distributed deep Q-network, the shared feature extraction network consists of three fully connected layers with the following numbers of neurons: The activation function used is ReLU. The input layer dimension is equal to the dimension of the state vector, and the output dimension is 32.
[0053] The reward distribution output branch receives shared features, then connects to a fully connected layer with 64 neurons and ReLU activation, and finally outputs a dimension of... The tensor, in which , The output of this branch represents the 200 quantile values corresponding to each action. quantiles .
[0054] In step S4, based on the cost quantile distribution and the maximum allowed expected default probability threshold, actions that cause the expected default risk to exceed the maximum allowed expected default probability threshold are filtered to obtain a set of allowed actions.
[0055] Action filtering is a key step in ensuring that quality constraints are met. For each action, the cost quantile distribution is obtained from the cost distribution output branch.
[0056] like Figure 2 As shown, the filtering process includes the following sub-steps: Step S41: Take the confidence level of the cost quantile distribution as... The corresponding upper quantile value; Step S42: If the upper quantile value is greater than 0, then the expected default risk of the action is determined to exceed the maximum permissible expected default probability threshold. This action will be disabled.
[0057] If the upper quantile value is greater than 0, it indicates that under this action, the probability of a future quality default exceeding the allowable threshold with a confidence level of 99.5%. This action will be blocked and excluded from the allowed action set. Since the cost value can only be 0 or 1, and the cost quantile value can only be 1 in the tail region of the distribution, this filtering mechanism can accurately identify and block high-risk actions. When the allowed action set is empty, continued pressure holding is forced.
[0058] In step S41 above, for each action The cost distribution output branch gives its value at the quantile. Quantile estimate at Since the cost value can only be 0 or 1, and The confidence level needs to be calculated as follows: The corresponding upper quantile value.
[0059] If there exists a certain Then take directly ; If 0.995 is not in the quantile set, then take the smallest quantile greater than 0.995. The corresponding quantile values are used as approximations, or linear interpolation is employed:
[0060] The upper quantile value is calculated.
[0061] When the allowed action set is empty, the system forces continued pressure holding. To prevent infinite pressure holding due to anomalies, a safety constraint is set: The absolute maximum holding time is equal to 1.2 times the maximum holding time based on process experience. If the current holding time exceeds the absolute maximum holding time, the holding time will be forcibly stopped. Simultaneously, an alarm will be triggered, the abnormal state will be recorded, and the operator will be notified to check for deviations in temperature and pressure sensors or model predictions.
[0062] In step S5, within the allowed action set, the reward quantile distribution and the current Lagrange multiplier are used. Calculate the overall action value and select the action to be executed accordingly.
[0063] For each action in the filtered set of allowed actions, calculate its composite action value, such as... Figure 3 As shown, the specific steps include the following: Step S511: Calculate the quantile distribution of returns. The expected value of the return quantile distribution is calculated by averaging the quantiles. Step S512: Calculate the cost quantile distribution The expected value of the cost quantile distribution is calculated by averaging the quantiles. Step S513: Subtract the current Lagrange multiplier from the expected value of the return quantile distribution. The difference between the product of the cost quantile distribution and the expected value is used as the action value.
[0064] When training the distributed deep Q-network, a masked ε-greedy strategy is used to select the action to be executed, such as... Figure 4 As shown, it includes the following sub-steps: Step S521: With probability Select the action with the highest overall action value from the set of allowed actions; Step S522: With probability A random action is generated. If the action does not belong to the set of allowed actions, a new action is randomly selected from the set of allowed actions. Step S523: When the set of allowed actions is empty, force the selection to continue holding pressure.
[0065] For example, A value of 0.1 can be used, and it should gradually decrease as training progresses, starting from 1.0.
[0066] By number of training rounds Exponential decay:
[0067] That is, multiply by 0.995 each round, until the lower limit of 0.01.
[0068] In step S6, the selected action is executed and the next state is obtained. The distributed deep Q-network is updated using the generated energy reward and quality default cost, and the Lagrange multipliers are updated using the cumulative default cost of the holding pressure process. .
[0069] After executing the selected action, obtain the reward and cost values from the environment and transition to the next state. Store the transition experience, including the current state, action, reward, cost, next state, and termination flag state, in the experience replay buffer. In the reward value, continuing to hold pressure results in a negative energy consumption value, while stopping holding pressure results in 0. In the cost value, the process step is 0, and the termination step is 0 or 1 depending on the quality inspection result.
[0070] A batch of empirical data is randomly sampled from the buffer to update the network. For the reward distribution branch, the target reward distribution is calculated and updated using quantile regression loss; for the cost distribution branch, the target cost distribution is calculated and updated using quantile regression loss as well.
[0071] After a complete holding pressure process is completed, the Lagrange multipliers are updated using the cumulative default costs of that process. Specifically, it includes the following sub-steps: Step S61: Calculate the discounted total cost of a complete pressure holding process. Assume the process takes a total of Step, cost of each step is The total cost of the discount is Since the cost of each process step is 0, the total cost of this discount is equal to , The cost value for the termination step; Step S62: Calculate the update amount in The learning rate of the Lagrange multipliers; Step S63: Set the current Lagrange multiplier Updated to And the updated Lagrange multipliers Apply a preset upper limit and exponential moving average smoothing.
[0072] Among these methods, the updated Lagrange multipliers are smoothed using an exponential moving average:
[0073] in, This is the original value after the upper limit is applied. The smoothed multipliers, which are stored and used, are initialized to 0; This is the smoothing coefficient. It can be set during the first 100 rounds of training. To expedite the response.
[0074] For example, the learning rate of the Lagrange multiplier Take 0.01; for the updated Lagrange multipliers The preset upper limit is set to 100.0.
[0075] The logic behind this update rule is: if the total cost of a pressure holding process exceeds the allowable threshold... Then increase the Lagrange multiplier. Increase the penalty weight for cost items in the next round to make the algorithm more inclined to stop early to avoid default; conversely, if the total discounted cost is below the threshold, decrease the Lagrange multiplier. This provides greater scope for energy consumption optimization.
[0076] During the algorithm training phase, a complete set of historical production records is collected. Each record contains complete time-series data on temperature, pressure, and thickness, along with corresponding laboratory quality inspection results. These results include internal bond strength and thickness deviation. During training, when the environment transitions to a termination state, the system, based on the complete time-series characteristics of the current pressure-holding process, uses the Dynamic Time Warping (DTW) algorithm to search the historical database for the most similar process. The matched actual quality inspection result is used as the cost value for this termination step. If a match fails, a conservative value of 1 is assigned.
[0077] In step S7, the updated state is used as the new current state. If the process has terminated, a new holding pressure process is re-initialized, and steps S3 to S6 are repeated while updating the network and Lagrange multipliers. The process continues until the convergence condition is met, resulting in a well-trained distributed deep Q-network and Lagrange multipliers. .
[0078] Specifically, the convergence condition can be set as follows: the average return value fluctuation over N consecutive training rounds is less than a preset threshold, or the number of training rounds reaches a preset maximum value.
[0079] One example of a convergence condition setting is: the average return value fluctuation over 100 consecutive training rounds is less than a preset threshold of 5%, or the number of training rounds reaches a preset maximum of 10,000 rounds.
[0080] During the offline training phase, at the start of each training round, the initial state of a complete holding pressure record is randomly selected from the historical dataset as the starting state for that round. If a simulation environment is used, initial conditions can be randomly generated according to the distribution of process parameters. The initial state includes the slab temperature, thickness, and degree of curing at the start of the holding pressure. The holding time is 0.
[0081] In step S8, the real-time acquired state vector is input into the trained distributed deep Q-network, based on the output quantile distribution and the Lagrange multipliers. Make a decision to maintain pressure or stop.
[0082] The trained network and multipliers are deployed to the actual production control system. During hot pressing, state vectors are acquired in real time and input into the network to obtain the reward and cost quantile distributions for each action. Action filtering and comprehensive action value calculation are then performed sequentially, selecting the action with the highest comprehensive value for output. If the output stops holding pressure, the hot press is controlled to release pressure and open the mold; if the output continues holding pressure, the system waits for the next decision cycle to make another judgment.
[0083] In some embodiments, an adaptive decision system for hot pressing and holding time of particleboard is also disclosed, comprising: The data acquisition module is connected to a temperature sensor, a pressure sensor, a displacement sensor, and a moisture content detector to acquire the status information during the hot pressing process. The constraint reinforcement learning decision module has a built-in distributed deep Q network trained according to the adaptive decision algorithm in any of the above embodiments, and executes the calculation process from S3 to S5 to output the pressure holding or stop action according to the real-time state vector. The security filtering module filters out high-risk actions based on the cost quantile distribution and the maximum allowable expected default probability threshold. The Lagrange multiplier storage module stores the Lagrange multiplier values for the decision module to call and receives updates from the training and update module. The training and update module is used to update the distributed deep Q network and the Lagrange multipliers offline or online using production data.
[0084] The various embodiments of the present invention have now been described in detail. To avoid obscuring the concept of the invention, some details known in the art have not been described. Those skilled in the art will fully understand how to implement the technical solutions disclosed herein based on the above description.
[0085] The embodiments described above only illustrate some implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. An adaptive decision-making algorithm for hot pressing and holding time of particleboard, characterized in that, Includes the following steps: S1: Obtain the state information during the hot pressing process, including at least the slab temperature, pressure, thickness, holding time, and raw material characteristic parameters, to form a state vector; S2: Construct a constrained Markov decision process model, determine that the action space consists of continuing to hold pressure and stopping to hold pressure, and set a discount factor. Reward function The cost function is set based on process energy consumption. The system is designed based on events of non-compliance with board quality standards, and a maximum allowable expected probability threshold for default is set. ; S3: Process the state vector using a distributed deep Q-network to obtain the future reward quantile distribution and future cost quantile distribution for each action; S4: Based on the cost quantile distribution and the maximum allowed expected default probability threshold, filter out actions that cause the expected default risk to exceed the maximum allowed expected default probability threshold to obtain a set of allowed actions; S5: In the set of allowed actions, utilize the reward quantile distribution and the current Lagrange multiplier. Calculate the overall action value and select the action to be executed accordingly; S6: Execute the selected action and obtain the next state. Update the distributed deep Q-network using the generated energy reward and quality default cost, and update the Lagrange multipliers using the cumulative default cost of the holding pressure process. ; S7: Using the updated state as the new current state, repeat S3~S6 and update the network and Lagrange multipliers. The process continues until the convergence condition is met, resulting in a well-trained distributed deep Q-network and Lagrange multipliers. ; S8: Input the real-time acquired state vector into the trained distributed deep Q-network, and determine the output quantile distribution and the Lagrange multiplier. Make a decision to maintain pressure or stop.
2. The adaptive decision-making algorithm for hot pressing and holding time of particleboard according to claim 1, characterized in that: The raw material characteristic parameters include the moisture content of the substrate, the curing kinetic parameters of the adhesive, the target density, and the estimated degree of curing calculated online based on the temperature history and the adhesive kinetic model.
3. The adaptive decision-making algorithm for hot pressing and holding time of particleboard according to claim 1, characterized in that, The cost function is: After the pressure holding is stopped, if the internal bond strength of the sheet is lower than the qualified threshold or the thickness deviation exceeds the allowable range, a cost value of 1 is assigned; otherwise, the cost value is 0. The cost value of each process step is 0.
4. The adaptive decision-making algorithm for hot pressing and holding time of particleboard according to claim 1, characterized in that: The distributed deep Q-network includes a shared feature extraction network, a reward distribution output branch, and a cost distribution output branch, each of which outputs a preset number of quantiles.
5. The adaptive decision-making algorithm for hot pressing and holding time of particleboard according to claim 1, characterized in that, The filtering in S4 includes the following sub-steps: S41: Take the confidence level of the cost quantile distribution as... The corresponding upper quantile value; S42: If the upper quantile value is greater than 0, then the expected default risk of the action is determined to exceed the maximum permissible expected default probability threshold. This action will be disabled.
6. The adaptive decision-making algorithm for hot pressing and holding time of particleboard according to claim 1, characterized in that, The calculation of the comprehensive action value in S5 includes the following sub-steps: S511: Calculate the expected value of the quantile distribution of returns; S512: Calculate the expected value of the cost quantile distribution; S513: Subtract the current Lagrange multiplier from the expected value of the aforementioned return quantile distribution. The difference between the product of the cost quantile distribution and the expected value is used as the action value.
7. The adaptive decision-making algorithm for hot pressing and holding time of particleboard according to claim 1, characterized in that, When training the distributed deep Q-network, step S5 selects the action to be executed using a masked ε-greedy strategy, including the following sub-steps: S521: Based on probability Select the action with the highest overall action value from the set of allowed actions; S522: Based on probability A random action is generated. If the action does not belong to the set of allowed actions, a new action is randomly selected from the set of allowed actions. S523: When the set of allowed actions is empty, force the selection to continue holding pressure.
8. The adaptive decision-making algorithm for hot pressing and holding time of particleboard according to claim 1, characterized in that, In step S6, the Lagrange multiplier is updated using the cumulative default cost of the holding pressure process. Includes the following sub-steps: S61: Calculate the discounted total cost of a complete pressure holding process. Assume the process takes a total of Step, cost of each step is The total cost of the discount is Since the cost of each process step is 0, the total cost of this discount is equal to... , The cost value for the termination step; S62: Calculate update amount in The learning rate of the Lagrange multipliers; S63: Set the current Lagrange multiplier Updated to And the updated Lagrange multipliers Apply a preset upper limit and exponential moving average smoothing.
9. The adaptive decision-making algorithm for hot pressing and holding time of particleboard according to claim 1, characterized in that: The maximum allowable expected probability of default threshold is set to 0.
005.
10. An adaptive decision-making system for hot pressing and holding time of particleboard, characterized in that, include: The data acquisition module is used to acquire the status information during the hot pressing process; The constrained reinforcement learning decision module has a built-in distributed deep Q-network trained by the adaptive decision algorithm according to any one of claims 1 to 9, which is used to output a pressure holding or stop action based on the real-time state vector. The security filtering module filters out high-risk actions based on the cost quantile distribution and the maximum allowable expected default probability threshold. Lagrange multiplier storage module, for storing the Lagrange multipliers; The training and update module is used to update the distributed deep Q network and the Lagrange multipliers offline or online using production data.