Fault Tolerant Control Method and Storage Medium for Humidifying Control System of Cigarette Cut Tobacco
By establishing a rule base for fault coding and processing solutions in the cigarette silk regeneration control system, and using OPC server and Actor network for action optimization, the fault tolerance control problem of cigarette silk regeneration equipment in the event of failure is solved, and a rapid and economical fault tolerance control effect is achieved.
Patent Information
- Application Number
- CN202211306278.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-24
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-10-24
AI Technical Summary
The prior art is difficult to effectively solve the fault-tolerant control problem of cigarette silk regeneration equipment in the event of failure, especially when the actuator fails completely.
By introducing a rule library of fault code, fault location and corresponding processing solutions in the cigarette silk reflux control system, combined with the OPC server and the Actor network, action selection and reward function optimization are realized to automatically issue control signals for fault tolerance control.
When a fault occurs, it can quickly identify the cause of the fault and select the minimum number of movements and the minimum valve movement to compensate, avoid parking maintenance losses, which is low cost, convenient and fast implementation.
Smart Images

Figure CN115586761B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a fault tolerance control method, and particularly to a fault tolerance control method for a cigarette cut tobacco conditioning control system and a storage medium. Background Art
[0002] During the production process of a cigarette cut tobacco line, conditioning equipment is mostly used to humidify and heat the leaves and cut tobacco, so as to increase the moisture content and temperature of the materials, enhance the toughness and processability of the materials, and meet the process requirements of subsequent processes. Drum-type conditioning equipment can be divided into leaf loosening and conditioning machines, leaf flavoring machines, leaf temperature and humidity increasing machines, cut tobacco super conditioning machines, cut tobacco flavoring machines, etc. according to their positions on the cut tobacco line and different process tasks.
[0003] CN114706356A discloses a fault tolerance control method for minimum-maximum optimization of an industrial process based on reinforcement learning, including: (1) establishing an augmented state space model containing a tracking error and a state increment based on the original system state space model with actuator faults and external disturbances, and proposing a performance index function according to the augmented state space model; (2) proposing a value function and a Q function according to the performance index function, and constructing expressions for the corresponding optimal control input, the worst external disturbance, the optimal control gain, and the worst external disturbance gain; (3) collecting data θj(k) and ρkj that can stabilize the system with the initial control gain and external disturbance gain, where θj(k) and ρkj are data containing system production information generated in the j-th iteration; (4) updating the control gain K1F and the external disturbance gain K2F through reinforcement learning; (5) if the iteration end condition is reached, the iteration ends, otherwise, return to step (4) to continue the iteration. The minimum-maximum optimization fault tolerance control method proposed by this invention can replace the traditional model-based fault tolerance control method, broaden the range of actuator faults that can be solved, and it aims at smaller faults such as actuator zero drift and external disturbances. If the actuator fails completely, including being unable to operate, fault tolerance control cannot be achieved. Essentially, it is a fault tolerance control method based on error compensation.
[0004] CN114448222A discloses a fault-tolerant control method, system, storage medium, as well as a transmission and drive control unit. When the first transmission control unit fails, this fault-tolerant control method can be applied to the second transmission control unit to perform fault-tolerant control on the first drive control unit originally controlled by the first transmission control unit, including the steps of: receiving and storing in real time the detection data of the first drive control unit by the first transmission control unit; when it is determined that the first transmission control unit fails, determining the parameter values of the first drive control unit; comparing the parameter values of the first drive control unit with the corresponding parameter values of the second drive control unit, and when the comparison result does not meet the preset conditions, and sending the adjusted control instruction to the first drive control unit to achieve the fault-tolerant control of the first drive control unit by the second transmission control unit. Through real-time communication, a cross-redundant system architecture is formed to achieve the fault-tolerant control of the single-car inverter system. This method forms a cross-redundant system through real-time communication, and the technical field belongs to the high-speed rail technology aspect.
[0005] CN114296350A discloses a fault-tolerant control method for an unmanned ship based on model reference reinforcement learning, including: analyzing the uncertain factors of the unmanned ship and constructing a nominal dynamic model of the unmanned ship; designing a nominal controller for the unmanned ship based on the nominal dynamic model of the unmanned ship; constructing a fault-tolerant controller based on model reference reinforcement learning according to the actual unmanned ship system, the difference in state variables between the nominal dynamic model of the unmanned ship and the output of the nominal controller of the unmanned ship by using the Actor-Critic method based on maximum entropy; building a reinforcement learning evaluation function and a control strategy model according to the control task requirements and training the fault-tolerant controller to obtain a trained control strategy. By using the present invention, the safety and reliability of the unmanned ship system can be significantly improved. As a fault-tolerant control method for an unmanned ship based on model reference reinforcement learning, the present invention can be widely applied to the field of unmanned ship control. This method is a model-based method, and the technical field belongs to unmanned ships. For the silk reeling and conditioning equipment, once a failure occurs, there must be a certain deviation in the previously established model, so the model-based method is not applicable to the situation of existing failures.
[0006] CN114253133A discloses a sliding mode fault-tolerant control method and device based on a dynamic event triggering mechanism, including: establishing a dynamic model of a discrete networked control system; designing a state sliding mode observer for the discrete networked control system based on the situation of sensor failures in the system; designing a sliding mode surface based on the observer estimation result; designing a dynamic event triggering mechanism; designing a sliding mode fault-tolerant controller based on the observer method. The present invention gives a sliding mode fault-tolerant control scheme considering sensor failures under a dynamic event triggering mechanism for a networked control system. This scheme is based on the situation of sensor failures in the system and is not a fault-tolerant control for control actuator failures.
[0007] CN114035523A discloses an industrial process fault-tolerant control method based on data-driven Q-learning, which includes the following steps: (1) Establish an equivalent state-space model with actuator faults including tracking error and state increment based on the state-space model of the original system, and propose a performance index function according to the new model; (2) Propose a value function and a Q function, and construct expressions for the corresponding optimal control input and control gain; (3) Initialize a stable control strategy K0 and collect data θj(k) and ρkj; (4) Update the controller gain K through a non-policy Q-learning algorithm; (5) If the iteration end condition is reached, the iteration ends, otherwise go back to step (4) to continue the iteration. The present invention can effectively address the problem that the system cannot be accurately modeled, reduce the system's dependence on the model, continuously learn by using the data generated in the actual production process of the system, and thus obtain the optimal control law, ultimately achieving good fault-tolerant control effects and tracking performance. The essence of this invention is to use the Q-learning algorithm to optimize the gain k of the controller in the fault state, but it does not consider the faults that the controller itself cannot control, which obviously exist, so the fault tolerance is not sufficient.
[0008] CN113954069A discloses a manipulator active fault-tolerant control method based on deep reinforcement learning, including: using deep learning methods for real-time fault detection, where the trained data-based dynamic model is used as the nominal model to generate the residual signal of the manipulator joint velocity, and fault detection and diagnosis are carried out according to the residual signal; when single or multiple actuator mutation faults occur, the fault joints are diagnosed and located; for the faulty joints, the auxiliary controller based on deep reinforcement learning works together with the system controller to output a compensation control torque to make up for the joint performance loss; among them, each joint of the manipulator is equipped with an auxiliary controller based on deep reinforcement learning. When a fault occurs, the auxiliary controller works in parallel with the nominal controller to autonomously estimate the fault degree of the actuator and output a compensation torque. The present invention can timely perform torque compensation when the manipulator has actuator mutation faults, realizing smooth active fault-tolerant control. This method performs active fault-tolerant control by assisting in outputting a compensation torque, which belongs to a form of redundancy. Summary of the Invention
[0009] The technical problem to be solved by the present invention is to overcome the above deficiencies and provide a fault-tolerant control method and storage medium for a cigarette cut tobacco conditioning control system.
[0010] The technical solution of the present invention is as follows:
[0011] A fault tolerance control method for a cigarette cut tobacco rehumidification control system. The control system is used to automatically control the hot air temperature in the drum of the cigarette cut tobacco rehumidification equipment, and includes the original equipment PLC and a fault tolerance control computer. The fault tolerance control computer is built-in with an OPC server and a fault signal alarm and fault code Faultnumber that can send out the root cause of the fault. The method includes the following steps:
[0012] Step 1, establish a rule base for fault code Faultnumber, fault location Faultlocation and corresponding processing solutions.
[0013] Step 2, according to the fault signal alarm and fault code Faultnumber of the root cause of the equipment fault sent by the fault tolerance control computer, query the fault rule base to obtain the fault location Faultlocation and corresponding processing solutions.
[0014] Step 3, establish a control strategy, including:
[0015] 3.1 The data detected by each sensor of the cut tobacco rehumidification equipment, the current opening degrees of each valve, the operating states and frequencies of each motor are used as the current state, represented by s, and the current moment is t. Represents the set of all previous, current, and future states:
[0016]
[0017] 3.2 The control action that the original equipment PLC is about to send to the control object is represented by a, and the current moment is t. Use To represent the set of all previous, current, and future actions:
[0018]
[0019] 3.3 The correlation probability between the current state of the cut tobacco rehumidification equipment and the actions taken previously is the state transition probability, denoted as p(s t+1 |s t , a t ). This probability is a Markov process, and those skilled in the art should know that its probability can be obtained from the historical production data stored in the database:
[0020] p(s t+1 |s 1 , a 1 , … s t , a t ) = P(s t+1 |s t , a t ) (3)
[0021] 3.4 Define a reward function. When a certain action is executed and the expected state is reached, a score is rewarded to evaluate the quality of the action taken by the controller:
[0022]
[0023] 3.5 Define a predictor
[0024] s′ t+1 = f(s t , a t ), t = 0, 1,..., N - 1, (5)
[0025] where: t is the current time, N is the time at the end of the task, used to given the current state s t , and the action a t , which can make a prediction about the state of the next stage according to formula (3). 3.5 Compose the Actor network, that is s′ t+1 = f(s t , a t ).
[0026] 3.6 Define the value function, whose meaning is the maximum score that can be obtained by the end of production under the current series of states,
[0027] where: 0 < γ < 1, γ is the attenuation coefficient, whose meaning is that the earlier actions are more meaningful and the value of taking actions later is lower. π represents a corresponding processing solution here, which is defined in the rule base in step 1 of the corresponding processing solution.
[0028] 3.7 Define the action value function,
[0029] whose meaning is the maximum score that can be obtained by the end of production for the current action; the difference from formula (6) is that the value function defined by formula (6) is only related to the state, and the action value function of formula (7) combines the current action and the current state. π represents a corresponding processing solution here, which is defined in the rule base in step 1 of the corresponding processing solution; compose the Crtic network with 3.6 and 3.8.
[0030] 3.8 Define the formula
[0031]
[0032] The purpose of formula (8) is that under the condition of a given state and action, the executed solution satisfies the maximum reward for formula (7).
[0033] Furthermore, the computer program formed by the fault-tolerant control method of the present invention is stored, run, and outputs results by an edge computing terminal. Under normal production conditions, it does not send control signals, but continuously reads and calculates corresponding results according to the fault-tolerant control method of the present invention by itself, which is called the offline training stage, and automatically updates the online Actor network according to the offline training results. When a fault signal is received, it enters the fault-tolerant control mode and starts to send control signals to the original device PLC through the OPC server. This control signal is an incremental signal, that is, based on the signal output by the original device PLC, addition or subtraction operations are performed.
[0034] Furthermore, it further includes step 4. When a fault signal alarm is received, the fault-tolerant control mode is started, including:
[0035] 4.1 If the fault alarm alarm is not received, it proceeds according to the normal working mode; if the fault alarm is received, it automatically enters the fault-tolerant control mode.
[0036] 4.2 Determine the control target of the fault-tolerant control according to the fault location Faultlocation and the fault code Faultnumbe, that is, the desired state s to be achieved.
[0037] 4.3 Select the best action to be executed and execute the fault-tolerant control scheme, including:
[0038] Set a loop until a satisfactory action is selected and ended:
[0039] Start of the loop:
[0040] 4.3.1 Select an action a t , input it into the prediction period, and obtain the predicted state after taking this action,
[0041] s′ t+1 = f(s t , a t ) (9)
[0042] 4.3.2 Evaluate the quality of the action to be taken by the reward function,
[0043]
[0044] where: s sp is the desired state to be achieved by the fault-tolerant control, d y is the dimension of the action space, that is, the number of non-zero outputs required to control the action, d α Set a coefficient of 0.005, which is the number of times required to reach the desired output, indicating that the smaller the number of control outputs, the better.
[0045] 4.3.3 Optimize the action Loss until it is less than a set value, which is set to 0.5 here
[0046]
[0047] Where: M is the maximum number of steps to be executed, which is set to 10000 here. After the action is executed, the deviation between the predictor and the actual execution result is observed to optimize the prediction accuracy of the predictor.
[0048]
[0049] Wherein: M is the maximum number of steps to be executed, which is set to 10000 here. The optimization method is the gradient descent method, which should be well known to those skilled in the art and will not be described in detail here.
[0050] 4.3.4 When performing an action a t After that, check whether the expected state s is reached sp If not achieved, execute 4.3.3 intermittently. If achieved, take no further action until deviation occurs and execute 4.3.3 again.
[0051] 4.3.5 When production is finished, the fault-tolerant control process is stopped, the fault location and fault code Pault number are output, and the technical staff is notified to actually handle the fault.
[0052] A computer-readable storage medium having a computer program stored thereon, characterized in that the computer program can be executed by a processor to implement the steps of the fault-tolerant control method of the cigarette shred rehumidification control system of the present invention.
[0053] Beneficial effects of the present invention:
[0054] (1) When a fault occurs, the control system gives the crux of the problem. After that, the fault-tolerant control method of the present invention selects the minimum number of actions and the least number of action valves according to the preset control strategy within the strategy range, compensates the hot air temperature, and automatically sends a control signal to the original equipment PLC to help the remaining production to be completed smoothly, avoiding the losses caused by stopping for maintenance and troubleshooting.
[0055] (2) Compared with other fault-tolerant control schemes, this method does not require modification of existing equipment pipelines or addition of redundant equipment, etc. It is lower in cost, more convenient to implement, and faster in implementation time.
[0056] (3) The present invention does not modify the existing control system. The original control system can still be used during normal production. The present invention only switches when a fault occurs, which will cause losses due to production stoppage, thereby avoiding losses caused by production stoppage as much as possible. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 : Flow chart of the fault tolerance control method for the cigarette cut tobacco rehumidifying equipment of the present invention.
[0058] Figure 2 : Schematic diagram of the fault tolerance control structure of the present invention.
[0059] Figure 3 : Schematic diagram of the fault tolerance control strategy.
[0060] Figure 4 : After a fault occurs, receive Faultnumber and Faultlocation and generate a fault tolerance control strategy diagram.
[0061] Figure 5 : Effect diagram generated by executing the fault tolerance control strategy. Specific implementation manner
[0062] As Figure 1 shown, a fault tolerance control method for a cigarette cut tobacco rehumidifying control system, the control system is used to realize the automatic control of the hot air temperature in the drum of the cigarette cut tobacco rehumidifying equipment, including the original equipment PLC and the fault tolerance control computer, and the fault tolerance control computer is built-in with an OPC server and a fault signal alarm and a fault code Faultnumber that can send out the root cause of the fault.
[0063] The method of the present invention includes the following steps:
[0064] Step 1, establish a fault code Faultnumber, a fault location Faultlocation and a corresponding processing solution rule library, and the fault rule library is shown in Table 1.
[0065] Table 1
[0066] Fault number Description Solution Code 1 Valve blockage Other influential position compensation 01 2 Plug or seat sediment Increase the valve opening and then close it slightly, repeat several times to try to flush it open 02 3 Plug or seat erosion Increase the valve opening 03 4 Increased valve or busbar friction Increase the supply pressure 04 5 External leakage Increase the supply pressure 04 6 Internal leakage (valve sealing) Increase the supply pressure 04 7 Excessive medium evaporation or critical flow rate Increase the supply pressure 04 8 Twist the servo motor piston rod Other influential position compensation 01 9 Servo motor housing or terminal sealing Increase the supply pressure 04 10 Servo motor diaphragm perforation Other influential position compensation 01 11 Servo motor spring failure Other influential position compensation 01 12 Electric sensor failure Control based on the actual flow rate and ignore this fault 05 13 Rod displacement sensor failure Control based on the actual flow rate and ignore this fault 05 14 Pressure sensor failure Control based on the actual flow rate and ignore this fault 05 15 Positioner feedback failure Control based on the actual flow rate and ignore this fault 05 16 Positioner supply pressure drop Increase the supply pressure 04 17 Unexpected pressure change between the entire valve Other influential position compensation 01 18 Fully or partially open bypass valve Reduce the supply pressure 04 19 Flow sensor failure Other influential position compensation 01
[0067] Step 2, according to the fault signal alarm and the fault code Faultnumber of the root cause of the equipment fault sent by the fault tolerance control computer, look up the fault rule library to obtain the fault location Faultlocation and the corresponding processing solution;
[0068] Step 3, establish a control strategy, including:
[0069] 3.1 The data detected by each sensor of the cut tobacco rehumidifying equipment, the current opening degree of each valve, the running state and frequency of each motor, etc., can be called the current state, represented by s, and the current moment is t, represents the set of all previous, current and future states,
[0070] 3.2 The control actions that the controller is about to issue to the controlled object, such as increasing the steam valve of the Y37 drum by 10% and closing the air supply of the regulating valve of the Y43 fresh air heat exchanger by 5%, are called actions, denoted by a. The current moment is t, and represents the set of all actions before, current, and after
[0071] 3.3 From a simple perspective, the current state of the silk reeling and moisture regain equipment is related to the actions taken before, but not deterministically related, because of raw materials, external factors, etc., there is a probability, denoted as p(s t+1 |s t , a t ). This probability is a Markov process, and those skilled in the art should know that p(s t+1 |s 1 , a 1 ,…s t , a t ) = p(s t+1 |s t , a t ). Its probability can be obtained from the historical production data stored in the database.
[0072] 3.4 Here, the present invention defines a reward function, whose meaning is that when a certain action is executed and the expected state is reached, a score is rewarded to evaluate the quality of the action taken by the controller.
[0073] 3.5 Define a predictor s′ t+1 = f(s t , a t ), t = 0, 1,..., N - 1, where t is the current moment and N is the moment when the task ends. The meaning of the predictor is that given the current state s t , and the action a t , it can predict the state of the next stage according to the formula in 3.3.
[0074] See Figure 2 shown, the Actor network is mainly the predictor s′ defined in this step t+1 = f(s t , a t ). Its function is to predict the new state s t+l that may be generated by taking this action before actually performing each action a. If the control requirements cannot be met, a new action a is generated until the new state generated by this action can meet the large reward calculated by the value function defined by the Crtic network composed of 3.6 and 3.8.
[0075] 3.6 Define the value function, which means the maximum score that can be obtained by the end of production under the current series of states. where 0 < γ < 1, γ is the attenuation coefficient, which means that earlier actions are more meaningful and the value of taking actions later is lower. π represents a certain policy here, and its policy is given in Table 1.
[0076] 3.7 Define the action value function. Its meaning is the maximum score that can be obtained by the end of production for the current action. The difference from 3.6 is that the value function defined in 3.6 is only related to the state, while the action value function in 3.7 combines the current action and the current state. π represents a certain policy here, and its policy is given in Table 1.
[0077] See Figure 2 As shown, 3.6 and 3.8 form the Crtic network.
[0078] 3.8 Therefore, the core of the implemented solution lies in making obtain the maximum reward.
[0079] As Figure 3 shown, the computer program formed by the fault-tolerant control method described in the present invention is stored, run, and outputs results by an edge computing terminal. Under normal production conditions, it does not send control signals, but always reads and calculates the corresponding results according to the fault-tolerant control method described in the present invention by itself, which is called the offline training stage, and automatically updates the online Actor network according to the offline training results. When receiving a fault signal, it switches to the fault-tolerant control mode and starts to send control signals to the original device PLC through the OPC server. The control signal is an incremental signal, that is, on the basis of the signal output by the original device PLC, addition or subtraction operations are performed.
[0080] See Figure 2 As shown, in step 4, when receiving the fault signal alarm, start the fault-tolerant control mode, including:
[0081] 4.1 If the fault alarm alarm is not received, operate according to the normal working mode; if the fault alarm is received, automatically switch to the fault-tolerant control mode.
[0082] 4.2 Determine the control objective of the fault-tolerant control according to the fault location Faultlocation and the fault code Faultnumbe, that is, the desired state s to be achieved.
[0083] 4.3 Select the best action to be executed and execute the fault-tolerant control scheme, including:
[0084] Set a loop until a satisfactory action is selected and end:
[0085] Loop start:
[0086] 4.3.1 Select an action a t , input it into the prediction period to obtain the predicted state s′ resulting from taking this action t+1 = f(s t , a t )
[0087] 4.3.2 Evaluate the action to be taken as good or bad by the reward function where s sp is the state expected to be achieved by fault tolerance control, d y is the dimension of the action space, i.e., the number of controls required for the action not to output 0, d α Set a coefficient of 0.005, which is the number of times required to reach the expected output, indicating that the smaller the number of control outputs, the better
[0088] 4.3.3 Optimize the action Loss until it is less than a set value, set to 0.5 here
[0089] where M is the maximum number of steps executed, set to 10000 here
[0090] After performing the action, observe the deviation between the predictor and the actual execution result, and optimize the prediction accuracy of the predictor where M is the maximum number of steps executed, set to 10000 here. The optimization method is the gradient descent method, which should be well-known to those skilled in the art and will not be elaborated here
[0091] 4.3.4 When performing an action a t , check whether the expected state s sp is reached. If not, continue to execute 4.3.3. If reached, do not take further action until a deviation occurs and then execute 4.3.3 again
[0092] 4.3.5 Production ends, stop the fault tolerance control process, output the fault location Faultlocation and the fault code Faultnumber, and notify the technician to actually handle the fault
[0093] The input-output relationships between the steps of the present invention are shown in the following table:
[0094]
[0095] Embodiment
[0096] According to the composition of the current cigarette silk conditioning equipment, the components involved include:
[0097] Y3. Feed end cleaning water valve 1, Y4. Discharge end cleaning water valve 2, Y5. Discharge end cleaning water valve 2, Y6. Water addition regulating valve, Y7. Circulating hot air heat exchanger regulating valve, Y8. Atomizing steam valve, Y9. Water addition regulating valve air supply, Y12. Drain compressed air valve, Y19 Discharge end purging valve, Y20 Screen cleaning and purging valve, Y25 Fresh air regulating damper air supply, Y26. Fresh air regulating damper, Y28. Circulating hot air regulating damper air supply, Y29. Circulating hot air regulating damper, Y30. Pipeline steam application valve, Y33. Differential pressure transmitter purging shut-off valve, Y34. Moisture exhaust regulating damper, Y35. Moisture exhaust regulating damper air supply, Y36. Differential pressure transmitter pipeline purging valve, Y37. Drum steam application valve, Y43. Fresh air heat exchanger regulating valve air supply, Y44. Fresh air heat exchanger regulating valve, Y45. Steam application regulating valve air supply, Y48. Steam application regulating valve.
[0098] The correspondence between the fault code Faultnumber issued by the fault-tolerant control computer and the root cause is shown in the following table.
[0099] Fault number Description of the root cause of the fault 1 Valve blockage 2 Plug or seat sediment 3 Plug or seat erosion 4 Increased valve or busbar friction 5 External leakage 6 Internal leakage (valve sealing) 7 Medium evaporation or critical flow rate 8 Twist the servo motor piston rod 9 Servo motor housing or terminal sealing 10 Servo motor diaphragm perforation 11 Servo motor spring failure 12 Electric sensor failure 13 Rod displacement sensor failure 14 Pressure sensor failure 15 Positioner feedback failure 16 Positioner supply pressure drop 17 Unexpected pressure change between the entire valve 18 Fully or partially open bypass valve 19 Flow sensor failure
[0100] The redundant control method includes:
[0101] 1. The fault signal alarm and fault code Faultnumber of the root cause of the equipment fault issued by the fault-tolerant control computer, as Figure 4 shown.
[0102] 2. According to the alarm code, match Y37. Drum steam application valve, and a valve blockage fault occurs. The solution is to "compensate for other affected parts". Y37. Drum steam application valve, its function is to directly inject steam into the control drum, thereby affecting the hot air temperature inside the drum. The hot air temperature needs to be stabilized within the range of 65°±1°C. Other parts that affect the hot air temperature are:
[0103] 1) Y1. Circulating hot air heat exchanger regulating valve air supply, its function is to control the steam volume entering the radiator and control the circulating air temperature, thereby affecting the hot air temperature inside the drum
[0104] 2) Y28. Circulating air regulating damper air supply, its function is to control the circulating air supply volume, that is, the air volume of the circulating air heated by the hot air heat exchanger entering the drum of Y1.
[0105] 3) Y7. Fresh air heat exchanger regulating valve air supply, its function is to control the steam volume entering the fresh air radiator to control the fresh air temperature, thereby affecting the hot air temperature inside the drum
[0106] 4) Y26. Fresh air regulating damper, its function is to control the fresh air supply volume, thereby affecting the hot air temperature inside the drum.
[0107] 5) Y34. Exhaust air regulating damper, whose function is to regulate the volume of exhaust air, thereby affecting the hot air temperature inside the drum.
[0108] 3. Read the current status a t = {{Y3. Feed end cleaning water valve, 1,0}, {Y4. Discharge end cleaning water valve 2,0}, {Y5. Discharge end cleaning water valve 2,0}, {Y6. Water addition regulating valve, 38%}, {Y7. Circulating hot air heat exchanger regulating valve, 43%}, {Y8. Atomizing steam valve 49%}, {Y9. Water addition regulating valve air supply 6}, {Y12. Drain compressed air valve, 0}, {Y19 Discharge end purging valve, 0}, {Y20 Screen cleaning valve, 0}, {Y25 Fresh air regulating damper air supply, 0}{Y26. Fresh air regulating damper, 0}, {Y28. Circulating hot air regulating damper air supply, 47}, {Y29. Circulating hot air regulating damper}, {Y30. Pipeline steam application valve, 38}, {Y33. Differential pressure transmitter purging shut-off valve, 0}, {Y34. Exhaust air regulating damper, 43}, {Y35. Exhaust air regulating damper air supply 36}, {Y36. Differential pressure transmitter pipeline purging valve, 0}, {Y37. Drum steam application valve, 0}, {Y43. Fresh air heat exchanger regulating valve air supply 0}, {Y44. Fresh air heat exchanger regulating valve 0}, {Y45. Fresh air steam application regulating valve air supply}, {Y48. Drum steam application application regulating valve, 67}, {Pre-loosening moisture content before flow, 6011.34}, {Pre-loosening water addition flow, 341.23}, {Pre-loosening total water addition 234.13}, {Pre-loosening steam flow 124.31}, {Pre-loosening total steam 201.31}, {Pre-loosening hot air temperature, 62.12}, {Pre-loosening hot air damper opening 34%}, {Pre-loosening drum motor frequency, 12}, {Loose outlet moisture content 17.81}, {Loose outlet temperature 54.3}}
[0109] Read the current action, a t= {{Y3. Feed end cleaning water valve 1,0}, {Y4. Discharge end cleaning water valve 2,0}, {Y5. Discharge end cleaning water valve 2,0}, {Y6. Water addition regulating valve, 38%}, {Y7. Circulating hot air heat exchanger regulating valve, 43%}, {Y8. Atomizing steam valve 49%}, {Y9. Water addition regulating valve air supply 6}, {Y12. Exhaust compressed air valve, 0}, {Y19 Discharge end purging valve, 0}, {Y20 Screen cleaning and purging valve, 0}, {Y25 Fresh air regulating damper air supply, 0} {Y26. Fresh air regulating damper, 0}, {Y28. Circulating hot air regulating damper air supply, 47}, {Y29. Circulating hot air regulating damper}, {Y30. Pipeline steam application valve, 38}, {Y33. Differential pressure transmitter purging shut-off valve, 0}, {Y34. Moisture exhaust regulating damper, 43}, {Y35. Moisture exhaust regulating damper air supply 36}, {Y36. Differential pressure transmitter pipeline purging valve, 0}, {Y37. Drum steam application valve, 0}, {Y43. Fresh air heat exchanger regulating valve air supply 0}, {Y44. Fresh air heat exchanger regulating valve 0}, {Y45. Fresh air steam application regulating valve air supply}, {Y48. Drum steam application and regulating valve, 67}}
[0110] For the convenience of understanding, the meanings of each value are added here. During the operation of the computer program, two bytes can be allocated to each variable, and the meaning of the fixed position can be specified.
[0111] 4. At this time, switch to the fault tolerance control program.
[0112] 4.3.1 The program starts to automatically select an action a t , and input it into the prediction period to obtain the predicted state of taking this action,
[0113] s′ t+1 = f(s t , a t )
[0114] 4.3.2 After that, the program evaluates the action taken. The reward function evaluates its quality, where s sp is the state expected to be achieved by the fault tolerance control, d y is the dimension of the action space, that is, the number of controls required for the action that is not output 0, d α Set a coefficient of 0.005, which is the number of times required to reach the expected output, indicating that the smaller the number of control outputs, the better.
[0115] Due to the requirements set in the formula, the smaller the number of control outputs and the fewer variables participating in the action, the more rewards can be obtained. Then use to optimize so that the reward function reaches the maximum value within a certain period.
[0116] The optimized result at this time is as follows:
[0117] α′ t ={{Y3. Feed end cleaning water valve, 0}, {Y4. Discharge end cleaning water valve 2, 0}, {Y5. Discharge end cleaning water valve 2, 0}, {Y6. Water addition regulating valve, 0}, {Y7. Circulating hot air heat exchanger regulating valve, 0}, {Y8. Atomizing steam valve 0}, {Y9. Water addition regulating valve air supply 0}, {Y12. Exhaust compressed air valve, 0}, {Y19 Discharge end purging valve, 0}, {Y20 Screen cleaning and blowing valve, 0}, {Y25 Fresh air regulating damper air supply, 0}, {Y26. Fresh air regulating damper, 0}, {Y28. Circulating hot air regulating damper air supply, 0}, {Y29. Circulating hot air regulating damper}, {Y30. Pipeline steam application valve, 0}, {Y33. Differential pressure transmitter purging shut-off valve, 0}, {Y34. Moisture exhaust regulating damper, 0}, {Y35. Moisture exhaust regulating damper air supply 0}, {Y36. Differential pressure transmitter pipeline purging valve, 0}, {Y37. Drum steam application valve, 0}, {Y43. Fresh air heat exchanger regulating valve air supply 0}, {Y44. Fresh air heat exchanger regulating valve 0}, {Y45. Fresh air steam application regulating valve air supply, 18}, {Y48. Drum steam application and regulating valve, 0}} It can be seen that after optimization, other valves in the program remain unchanged, only {Y45. Fresh air steam application regulating valve air supply, 18} is changed, and then execute.
[0118] After execution, the hot air temperature still does not reach the range of 65°±1°C. Continue to execute the steps in 4.3.1 until the production ends. Finally, the global loose rewetting hot air temperature curve is as Figure 5 shown.
[0119] The effects that can be produced by this embodiment include:
[0120] From Figure 5 it can be seen that when a fault occurs, the root cause reasoning method directly points out the crux of the problem. After that, the fault-tolerant control method described in the present invention selects the minimum number of action times and the fewest action valves within its strategy range according to the preset control strategy to compensate for the hot air temperature, and automatically sends a control signal to the original equipment PLC to help the remaining production be successfully completed, avoiding the losses caused by parking for maintenance and troubleshooting.
[0121] In addition, compared with other fault-tolerant control schemes, this method does not modify the pipelines of existing equipment or add redundant equipment, etc., with lower cost, more convenient implementation, and faster implementation time.
[0122] In addition, the present invention does not modify the existing control system, and the original control system can still be used during normal production. The present invention only switches when a fault occurs and is about to cause losses due to production stoppage, so as to avoid the losses caused by production stoppage as much as possible.
Claims
1. A fault-tolerant control method for a cigarette cut tobacco conditioning control system, characterized in that, the control system is used to achieve automatic control of the hot air temperature in the drum of the cigarette cut tobacco conditioning equipment, including the original equipment PLC and a fault-tolerant control computer. The fault-tolerant control computer is built-in with an OPC server and a fault signal alarm and a fault code Faultnumber that can issue the root cause of the fault. The method includes the following steps: Step 1, establish a rule library for fault code Faultnumber, fault location Faultlocation and corresponding processing solutions; Step 2, according to the fault signal alarm and fault code Faultnumber of the root cause of the equipment fault issued by the fault-tolerant control computer, query the fault rule library to obtain the fault location Faultlocation and corresponding processing solutions; Step 3, establish a control strategy, including: Step 3.1, regard the data detected by each sensor of the silk reeling and conditioning equipment, the current opening degrees of each valve, the operating states and frequencies of each motor as the current state, denoted by s, and the current moment is t. Denote the set of all previous, current, and future states: Step 3.2, the control action that the original equipment PLC is about to send to the control object is represented by a, the current moment is t, and is represented by the set of all actions before, current, and after: Step 3.3, the correlation probability between the current state of the silk reeling and moisture regain equipment and the actions taken previously is denoted as p(s t+1 |s t , a t ). This probability is a Markov process, and its probability can be obtained from the historical production data stored in the database: p(s t+1 |s 1 ,a 1 ,…s t ,a t ) = p(s t+1 |s t ,a t ) (3) Step 3.4, define a reward function. When a certain action is executed and the expected state is reached, a score is rewarded to evaluate the quality of the action taken by the controller: Step 3.5, define a predictor s′ t+1 = f(s t , a t ), t = 0, 1, ..., N-1, (5) Where: t is the current moment, and N is the moment at the end of the task, which is used to specify the current state s t and action a t , which makes a prediction about the state of the next stage according to formula (3); the predictor s' t+1 = f(s t , a t ) is used to predict the new state s t+1 that may be generated by taking this action before actually performing each action a. If the control requirements cannot be met, a new action a is generated until the new state generated by this action can meet the large reward calculated by the value function defined by the Crtic network composed of steps 3.6 and 3.8; Step 3.6, define the value function, which means the maximum score that can be obtained by the end of production under the current series of states. Wherein: 0 < γ < 1, where γ is the attenuation coefficient, which means that earlier actions are more meaningful and the value of taking actions later is lower. Here, π represents a corresponding processing solution, which is defined in the rule base in step 1 of the corresponding processing solution. Step 3.7, define the action value function, used to represent the maximum score that can be obtained from the current action until the end of production. π represents a corresponding processing solution here, and its corresponding processing solution is defined in the rule library in Step 1; Combine Step 3.6 and Step 3.8 to form a Crtic network; Step 3.8, define the formula The implemented solution meets the requirement of maximizing J(π θ ), i.e., it meets the set process and quality indicators.
2. The fault-tolerant control method for a cigarette cut tobacco conditioning control system according to claim 1, characterized in that, further includes: Step 4, calculate the action to be executed and predict the action result, including: Step 4.1, if the fault alarm alarm is not received, proceed according to the normal working mode; if the fault alarm is received, automatically switch to the fault-tolerant control mode; Step 4.2, determine the control target of the fault-tolerant control according to the fault location Faultlocation and fault code Faultnumbe, that is, the expected state s to be achieved; Step 4.3, select the best action to be executed and execute the fault-tolerant control plan, including: Set a loop until a satisfactory action is selected and ended: Start of loop: Step 4.3.1 Select an action a t , input it into the prediction period, and obtain the state of the prediction made by taking this action s′ t+1 = f(s t , a t ) (9) Step 4.3.2, evaluate the quality of the action to be taken by the reward function, Where: S sp is the state expected to be achieved by fault tolerance control, d y is the dimension of the action space, that is, the number of controls required for the action not to output 0, d α Set the coefficient of 0.005, which is the number of times required to achieve the desired output, indicating that the smaller the number of control outputs, the better. Step 4.3.3, optimize the action Loss until it is less than a set value where: M is the maximum number of steps executed; Step 4.3.4, when performing an action a t check whether the desired state S is reached sp if not, continue to execute Step 4.3.3; if reached, take no further action until a deviation occurs and then execute Step 4.3.3 again; Step 4.3.5, when production ends, stop the fault-tolerant control process, output the fault location Faultlocation and fault code Faultnumber, and notify the technical personnel to actually handle the fault.
3. The fault-tolerant control method for a cigarette cut tobacco conditioning control system according to claim 2, characterized in that, in Step 4.3.3: the optimized action Loss is set to 0.
5.
4. The fault-tolerant control method for a cigarette cut tobacco conditioning control system according to claim 2, characterized in that, in Step 4.3.3: the maximum number of steps M executed is set to 10000.
5. The fault tolerance control method of the cigarette cut tobacco conditioning control system according to claim 2, characterized in that, in step 4.3.3: The optimization method is the gradient descent method.
6. The fault tolerance control method of the cigarette cut tobacco conditioning control system according to any one of claims 1-5, characterized in that: The fault tolerance control method continuously reads and performs offline training, and automatically updates the online Actor network according to the offline training results. No control signal is sent in the normal production state. When a fault signal is received, it enters the fault tolerance control mode and starts to send a control signal to the original equipment PLC through the OPC server. This control signal is an incremental signal, that is, an addition or subtraction operation is performed on the basis of the signal output by the original equipment PLC.
7. A computer-readable storage medium, on which a computer program is stored, characterized in that, the computer program can be executed by a processor to implement the steps of the fault tolerance control method of the cigarette cut tobacco conditioning control system according to any one of claims 1-6.
Citation Information
Patent Citations
Mechanical arm active fault-tolerant control method based on deep reinforcement learning
CN113954069A
Cigarette tobacco cutting process tobacco flake preprocessing stage on-line monitoring and fault diagnosis method
CN104865951A
Vacuum damping machine operation state monitoring method and operation system
CN115067531A