Degradable film bag blowing film parameter regulation method based on reinforcement learning
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI PROVINCE HUACHI PLASTIC IND CO LTD
- Filing Date
- 2026-06-30
- Publication Date
- 2026-08-04
AI Technical Summary
[0006]本申请实施例提供了一种基于强化学习的降解膜袋吹膜参数调控方法,解决三层共挤降解膜袋隐藏层间供料偏移误调控问题
本发明通过将各层运行变化量输入层流量估计模型获得各层供料偏移类别,并结合总膜厚归一化偏差、膜泡宽度归一化偏差、牵引速度归一化变化量及连续判定窗口内的供料偏移类别进行一致性判定,使层流量隐状态能够区分隐藏层间偏移状态和整体成型偏差状态;在总膜厚和膜泡宽度处于允许范围而某层供料偏移持续存在时,该供料偏移进入后续动作判定过程,使外层连续性约束、中层供料稳定约束和内层热封供料约束参与目标调控动作的选取。
Smart Images

Figure CN122500924A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent control technology for blown film forming of biodegradable film bags, and in particular to a method for adjusting the parameters of blown film forming of biodegradable film bags based on reinforcement learning. Background Technology
[0002] Biodegradable food bags and heat-sealing bags are typically produced using blown film molding processes with PBAT, PLA, starch-based modified materials, or their blends. To balance appearance continuity, support performance, and heat-sealing contact performance, a three-layer co-extrusion structure is commonly used in biodegradable film bag production. The outer, middle, and inner layers respectively handle surface continuity, material support, and heat-sealing contact. Control parameters during blown film production typically include screw speed for each layer, melt temperature, die temperature, traction speed, cooling airflow, and cooling air temperature. Production control often uses total film thickness feedback, bubble width, frost line position, and equipment operating status as adjustment criteria.
[0003] During the blown film production process of three-layer co-extruded biodegradable film bags, differences exist in the melt viscosity, hygroscopic state, heat sensitivity, and filler content of the different layers. When the middle layer uses a highly filled biodegradable modified material or a starch-based modified material, the melt pressure, drive load, and filtration pressure difference of the middle layer are prone to fluctuations, which affect the film bubble formation state through the co-extrusion die. When the inner layer uses a heat-sealing contact material, the relative material supply of the inner layer is also affected by changes in traction speed, melt temperature, and cooling state. Therefore, even if the total film thickness feedback data and film bubble width are still within the allowable range, a relative material supply offset may have already occurred between the outer, middle, and inner layers.
[0004] Current blown film parameter control methods often rely on adjustments based on total film thickness deviation or bubble width deviation. This can easily lead to interlayer relative material feeding offsets being attributed to overall forming deviations, triggering adjustments to traction speed, cooling status, or screw speeds for each layer. This type of adjustment lacks simultaneous determination of constraints on outer layer continuity, middle layer material feeding stability, and inner layer heat-sealing material feeding, which can easily cause fluctuations in inner layer heat-sealing contact, localized thinning of the outer layer, uneven middle layer support, and edge wrinkles.
[0005] Therefore, in the production of three-layer co-extruded biodegradable film bags, it is necessary to solve the technical problem of identifying the relative feeding offset between layers when the total film thickness and film bubble width are within the allowable range, and constraining the blowing film parameter control action accordingly. Summary of the Invention
[0006] This application provides a reinforcement learning-based method for controlling the parameters of blown film for three-layer co-extruded biodegradable film bags, which solves the problem of mis-control of material feeding offset between hidden layers in three-layer co-extruded biodegradable film bags.
[0007] This invention provides a reinforcement learning-based method for controlling the blown film parameters of a co-extruded biodegradable film bag. The co-extruded biodegradable film bag includes an outer layer, a middle layer, and an inner layer. The layer functional constraints include an outer layer continuity constraint, a middle layer material feeding stability constraint, and an inner layer heat-sealing material feeding constraint, including: Acquire multi-source operational data and material type identifiers; A layer functional state vector is generated based on the multi-source operational data and material type identifier; Extract the rotation speed change, load change, melt pressure change, melt temperature change and filtration pressure difference change of each layer from the layer functional state vector, input them into the layer flow estimation model to obtain the feeding offset category of each layer, and combine the total film thickness normalized deviation, film bubble width normalized deviation, traction speed normalized change and feeding offset category in the continuous judgment window to make consistency judgment, and generate the layer flow hidden state including hidden inter-layer offset state and overall forming deviation state. Based on the hidden state of the layer traffic and the layer functional constraints, candidate control actions are determined, and candidate control actions that weaken the layer functional constraints are written into the blocked candidate control actions. The set of allowed actions is output based on the unblocked candidate control actions. The layer functional state vector, the layer traffic hidden state, and the set of allowed actions are input into the reinforcement learning policy model to obtain the target control action. The target control action is converted into blown film parameter adjustment commands and sent to the control system; The reinforcement learning strategy model is updated based on the multi-source running data after execution.
[0008] In some embodiments, the multi-source operating data includes screw speed, drive load, melt pressure, melt temperature status, filtration pressure difference, die temperature status, traction speed, cooling air volume, cooling air temperature, film bubble width, frost line position, total film thickness feedback data, winding tension, ambient temperature, and ambient humidity of each extrusion unit. After acquiring the multi-source operating data, source markers are set according to the outer extrusion unit, middle extrusion unit, inner extrusion unit, co-extrusion die, film bubble forming position and traction winding position, and time alignment is performed according to the material transfer lag time from the extrusion unit to the thickness measurement position.
[0009] In some embodiments, generating the layer functional state vector includes: The material type identifiers of the outer, middle and inner layers are respectively bound to the screw speed, drive load, melt pressure, melt temperature state and filtration pressure difference of the corresponding extrusion unit to form a layered operation sub-state, and the change direction of each layer at adjacent sampling times is recorded. Write the membrane bubble width, frost line position, total membrane thickness feedback data, traction speed, cooling air volume and cooling air temperature into the shared molding sub-state; Write the startup, material change, steady-state production, and anomaly recovery status into the production stage status; The layered operation sub-state, the shared forming sub-state, and the production stage state are combined to obtain the layer functional state vector.
[0010] In some embodiments, the layer flow estimation model is trained using historical blown film process data, which includes material type identification, extrusion operation data, film bubble formation data, total film thickness feedback data, offline layer thickness detection data, and heat sealing detection data. During training, offline layer thickness detection data and heat sealing detection data are used as the annotation source for material feeding offset categories. Material change transition data, filter pressure difference increase data and long-term production drift data are written into the training samples, and the corresponding production stage status is retained.
[0011] In some embodiments, the candidate control action refers to a single action in the set of blown film parameter adjustment actions that the control system can execute, and each candidate control action includes the adjustment object, adjustment direction, adjustment amplitude category and execution order. The adjustment objects include the outer extrusion unit, the middle extrusion unit, the inner extrusion unit, the co-extrusion die, the traction mechanism, and the cooling mechanism; Determining the candidate control action includes: Before the target control action is output, candidate control actions that weaken the layer functional constraints are excluded based on the layer flow hidden state and layer functional constraints. When the layer flow rate latent state indicates that the inner layer feed is lower than the inner layer heat sealing feed constraint, candidate control actions that reduce the inner layer screw speed, increase the traction speed, or enhance cooling and cause the relative inner layer feed to decrease are written into the masked candidate control actions.
[0012] In some embodiments, the candidate control actions include adjusting the speed of the middle layer screw and adjusting the temperature of the middle layer melt; The determination of the candidate control action also includes: When the layer flow rate latent state characterizes the transmission of the middle layer feed fluctuation to the bubble forming process, the presence of the middle layer feed fluctuation transmission is determined by the middle layer same-direction fluctuation retention degree determined by the middle layer melt pressure and the middle layer drive load, as well as the forming follow indicator of the bubble width or frost line position relative to the changing direction of the middle layer feed side. When the mid-layer feeding fluctuation is transmitted, and the mid-layer screw speed increment or mid-layer melt temperature increment is in the same direction as the corresponding increment in the previously executed control action, the corresponding candidate control action is written into the masked candidate control action.
[0013] In some embodiments, the reinforcement learning policy model uses the layer functional state vector and the layer flow hidden state as state inputs, actions from the set of allowed actions as optional actions, and layered reward items as the basis for updating. The tiered reward items include layer functional constraint satisfaction items, membrane bubble stability items, total membrane thickness deviation items, thermal state constraint items, and motion smoothing items. The update priority of the layer functional constraint satisfaction items is higher than that of the total membrane thickness deviation items. When a target control action reduces the total film thickness deviation and causes any layer's functional constraint to be unsatisfied, the target control action is marked as a low priority action. When a target control action maintains the functional constraints of the membrane layer and causes the membrane bubble width to change continuously, the target control action is marked as a reusable action.
[0014] In some embodiments, the production stage states include the start-up stage, material change stage, steady-state production stage, and anomaly recovery stage; During the start-up or material change phase, the set of actions can be output in the following order: temperature, screw speed, traction speed, and cooling status. During the steady-state production phase, the set of actions that can be performed includes screw speed adjustment actions and melt temperature adjustment actions for each layer. During the abnormal recovery phase, if the melt pressure, filtration differential pressure, membrane bubble width, or frost line position meet the abnormal recovery triggering conditions, the output of new target control actions will be paused, and the safety control actions corresponding to the current production stage will be invoked. The abnormal recovery triggering conditions are determined by the pressure threshold, differential pressure threshold, width fluctuation threshold, and frost line fluctuation threshold.
[0015] In some embodiments, updating the reinforcement learning policy model includes offline pre-training and online updating; Offline pre-training uses historical layer functional state vectors, historical layer traffic hidden states, historical regulation actions, and historical layer functional constraint satisfaction states to generate offline training samples. Online updates are initiated when the target control action belongs to the set of allowed actions and the state of the layer function constraints satisfied after execution is obtained. When the layer functional state vector obtained online exceeds the coverage of the offline training samples, online updates are stopped, and a safe adjustment action is output, or the current adjustment action is used as the target adjustment action.
[0016] In some embodiments, the method further includes generating a control traceability record, which includes the closed-loop control time, layer function state vector, layer flow hidden state, masked candidate control actions, set of allowed actions, target control action, blown film parameter adjustment instruction, layer function constraint satisfaction state, and model update state. When abnormal heat sealing contact, edge wrinkles, abnormal winding, or interlayer feeding deviation occurs in the subsequent production process, the corresponding sample record is extracted according to the control traceability record, and the sample record is written into the training sample of the layer flow estimation model or reinforcement learning strategy model.
[0017] Through the above technical solution, the present invention can achieve at least the following beneficial effects: This invention obtains the material supply offset category of each layer by inputting the operational changes of each layer into the layer flow estimation model. It then combines the normalized deviation of the total film thickness, the normalized deviation of the bubble width, the normalized change of the traction speed, and the material supply offset category within the continuous judgment window to make a consistency judgment. This allows the latent state of the layer flow to distinguish between the hidden interlayer offset state and the overall forming deviation state. When the total film thickness and bubble width are within the allowable range, but the material supply offset of a certain layer persists, the material supply offset enters the subsequent action judgment process, allowing the outer layer continuity constraint, the middle layer material supply stability constraint, and the inner layer heat sealing material supply constraint to participate in the selection of the target control action.
[0018] By judging candidate control actions based on the hidden state of layer flow and layer functional constraints, and writing candidate control actions that weaken layer functional constraints into the masked candidate control actions, the action selection of the reinforcement learning strategy model is restricted to the set of allowed actions, so that control actions that reduce the relative material supply of the inner layer, continue the material supply fluctuation transmission of the middle layer, or weaken the continuity of the outer layer do not enter the target control action output.
[0019] By inputting the layer functional state vector, the layer flow hidden state, and the set of allowed actions into the reinforcement learning strategy model, the target control action is simultaneously affected by the current operating state, the interlayer feeding offset state, and the layer functional constraints, so that the blown film parameter adjustment command corresponds to the functional layer feeding state of the three-layer co-extruded biodegradable film bag.
[0020] By setting offline pre-training, online updates, and safety control actions, when the layer functional state vectors obtained online exceed the coverage of offline training samples, the reinforcement learning policy model stops online updates and outputs safety control actions, or uses the current control action as the target control action, so that the control process in the uncovered production state enters a restricted execution state. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation on the scope of this application.
[0022] Figure 1 This is a flowchart of the reinforcement learning-based degradable membrane bag blown film parameter control method in the embodiments. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0024] All terms used in this application (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0025] To facilitate understanding, the relevant terms and concepts involved in the embodiments of this application will be introduced below.
[0026] Multi-source operational data refers to the data set collected within the same production batch from the outer extrusion unit, middle extrusion unit, inner extrusion unit, co-extrusion die, film bubble forming position, and traction winding position. Material type identification refers to the category information associated with each layer of degradable material batch, including at least the material system, filler type, drying treatment state, and heat sensitivity rating.
[0027] Layer functional constraints refer to the control requirements corresponding to the functions of the outer layer, middle layer, and inner layer. Among them, the outer layer continuity constraint is characterized by the continuous state of the outer layer surface, the local weak state, and the edge forming state. The middle layer material supply stability constraint is characterized by the continuous change state of the middle layer melt pressure, driving load, and filtration pressure difference. The inner layer heat sealing material supply constraint is characterized by the inner layer relative material supply state, heat sealing contact state, and heat sealing test results.
[0028] Specifically, the layer functional constraint satisfaction state refers to the set of satisfaction results of the outer layer continuity constraint, the middle layer material supply stability constraint, and the inner layer heat sealing material supply constraint within the same judgment period. The outer layer continuity constraint satisfaction state is jointly determined by the outer layer surface continuity state, local weak state, and edge forming state; the middle layer material supply stability constraint satisfaction state is jointly determined by the changes in the middle layer melt pressure, driving load, and filtration pressure difference within the continuous judgment window; and the inner layer heat sealing material supply constraint satisfaction state is jointly determined by the inner layer relative material supply state, heat sealing contact state, and heat sealing detection results.
[0029] The heat-sealing test results are linked to the corresponding production batch and judgment time after the heat-sealing test is completed. For judgment periods where no heat-sealing test results are obtained, the inner layer relative feeding state and heat-sealing contact state are used to determine the inner layer heat-sealing feeding constraint satisfaction status. The outer layer surface continuity state, local weak state, and edge forming state are determined based on at least one of the following: film bubble width, frost line position, total film thickness feedback data, winding tension, and offline layer thickness detection data; the heat-sealing contact state is determined based on the inner layer relative feeding state and the heat-sealing test results.
[0030] Example 1: like Figure 1 As shown, this embodiment employs a reinforcement learning-based method for controlling the blown film parameters of a three-layer co-extruded biodegradable film bag. This method is applied to the production process of a three-layer co-extruded biodegradable film bag, which includes an outer layer, a middle layer, and an inner layer. Layer functional constraints include outer layer continuity constraints, middle layer material supply stability constraints, and inner layer heat-sealing material supply constraints, including: Step S1: Obtain multi-source operating data and material type identification. The multi-source operating data is generated by at least the outer extrusion unit, the middle extrusion unit, the inner extrusion unit, the co-extrusion die, the film bubble forming process, and the traction winding process. Step S2: Generate a layer functional state vector based on multi-source operational data and material type identifiers; Step S3: Estimate the hidden state of layer flow based on the layer functional state vector. The hidden state of layer flow represents the material supply offset of the outer layer, middle layer and inner layer relative to the layer functional constraints. The estimation of the hidden state of layer flow includes: extracting the changes in rotational speed, load, melt pressure, melt temperature, and filtration differential pressure of each layer from the layer functional state vector, inputting the extraction results into the layer flow estimation model to obtain the feeding offset category of each layer; and making a consistency determination on the hidden inter-layer offset state and the overall forming deviation state based on the normalized deviation of total film thickness, normalized deviation of film bubble width, normalized change of traction speed, and feeding offset category within the continuous determination window, and writing the consistency determination result into the hidden state of layer flow. Step S4: Based on the hidden state of layer traffic and layer functional constraints, determine the candidate control actions, write the candidate control actions that weaken the layer functional constraints into the blocked candidate control actions, and output the set of allowed actions based on the unblocked candidate control actions. Step S5: Input the layer functional state vector, the layer flow hidden state, and the set of allowed actions into the reinforcement learning policy model to obtain the target control action; Step S6: Convert the target control action into a blown film parameter adjustment command and send it to the control system; Step S7: Update the reinforcement learning policy model based on the multi-source running data after execution.
[0031] A layer functional state vector refers to a set of state data organized according to source markers, acquisition times, and production stage states. Source markers are used to distinguish the extrusion position, co-extrusion position, bubble forming position, and traction winding position corresponding to the data. Acquisition times are aligned according to the transmission hysteresis relationship of material from the extrusion unit to the bubble forming position and thickness measurement position.
[0032] In one optional implementation, generating the layer functional state vector includes: binding the material type identifiers of the outer, middle, and inner layers to the screw speed, drive load, melt pressure, melt temperature state, and filtration pressure difference of the corresponding extrusion unit to form a layered operation sub-state, and recording the change direction of each layer at adjacent sampling times; writing the membrane bubble width, frost line position, total membrane thickness feedback data, traction speed, cooling air volume, and cooling air temperature into the shared molding sub-state; writing start-up, material change, steady-state production, and abnormal recovery into the production stage state; and combining the layered operation sub-state, the shared molding sub-state, and the production stage state to obtain the layer functional state vector.
[0033] The layer flow hidden state refers to the relative material supply state of each layer relative to the layer functional constraints, estimated based on the layer functional state vector. Material supply offset categories include outer layer material supply high category, outer layer material supply low category, middle layer material supply high category, middle layer material supply low category, inner layer material supply high category, inner layer material supply low category, and baseline material supply category. Non-baseline material supply categories refer to those that differ from the baseline material supply category, including at least the material supply high category and the material supply low category. The layer flow estimation model receives the layered operation sub-state and shared molding sub-state from the layer functional state vector, outputs the material supply offset category, and writes the material supply offset category into the layer flow hidden state.
[0034] In one optional implementation, the multi-source operating data includes screw speed, drive load, melt pressure, melt temperature, filtration differential pressure, die temperature, traction speed, cooling air volume, cooling air temperature, bubble width, frost line position, total film thickness feedback data, winding tension, ambient temperature, and ambient humidity for each extrusion unit. After acquiring the multi-source operating data, source markers are set according to the outer extrusion unit, middle extrusion unit, inner extrusion unit, co-extrusion die, bubble forming position, and traction winding position, and time alignment is performed according to the material transfer lag time from the extrusion unit to the thickness measurement position.
[0035] In one optional implementation, the layer flow rate estimation model is trained using historical blown film process data, which includes material type identification, extrusion operation data, film bubble formation data, total film thickness feedback data, offline layer thickness detection data, and heat sealing detection data. During training, the offline layer thickness detection data and heat sealing detection data are used as the annotation source for the material feed offset category, and material changeover transition data, filter differential pressure rise data, and long-term production drift data are written into the training samples, while retaining the corresponding production stage states. During training, the layer flow rate estimation model uses the layered operation sub-state and the shared forming sub-state as training inputs, and the material feed offset category as the training output target.
[0036] In one embodiment of Example 1, when determining the consistency of the correspondence between the material feeding offset category of each layer and the total film thickness feedback data, film bubble width, and traction speed changes, the total film thickness feedback data, film bubble width, and traction speed changes are converted to the same judgment scale. This ensures that data from different sources, after being aligned according to the material transfer delay time from the extrusion unit to the forming position and thickness measurement position, participate in the judgment at the same judgment time. The conversion method is as follows: , in, To determine the time Total film thickness normalization bias, To determine the time Total film thickness feedback data, For the target total film thickness, The allowable deviation for the target total film thickness; To determine the time Normalized bias of membrane bubble width To determine the time The width of the membrane bubble, For the target membrane bubble width, This refers to the allowable fluctuation range of the membrane bubble width. To determine the time The normalized change in traction speed To determine the time traction speed, To determine the time The traction speed at the previous judgment moment, The allowable range of variation in traction speed; To complete the determination time marker after aligning the transmission hysteresis time, the subsequent formulas... The meanings are the same. The target total film thickness, target bubble width, allowable deviation of target total film thickness, allowable fluctuation range of bubble width, and allowable variation range of traction speed are jointly determined by product specifications, material type identification, and production stage status, and are read from the process record of the corresponding production batch. When a process record matching the current material type identification is missing, the most recent valid process record under the same product specification is used; when this valid process record is missing, the writing of new hidden interlayer offset status and overall forming deviation status is stopped, and the hidden status of the previous valid layer flow is maintained; when there is no hidden status of the previous valid layer flow, the safety control action is invoked.
[0037] For example, the transmission hysteresis time refers to the time delay from the corresponding acquisition position of the extrusion unit to the film bubble forming position and the thickness measurement position for the same material segment. The transmission hysteresis time is determined based on the traction speed, equipment position spacing, and production stage status. When the traction speed is continuously changing, the data alignment relationship at the corresponding judgment time is determined using the traction speed records within the continuous judgment window. After time alignment is completed, the layered operation sub-state, shared forming sub-state, and production stage status corresponding to the same judgment time jointly participate in the generation of the layer functional state vector.
[0038] To avoid incorrect writing of the material offset category due to fluctuations in single-point sampling, the retention rate of the material offset category for each layer within the continuous decision window is statistically analyzed. The method is as follows: , in, For the first Layer at the time of determination The feed offset category remains proportional; This is the layer identifier, with values corresponding to the outer, middle, and inner layers; The number of sampling points contained in the continuous decision window; The offset sequence number of the sampling point within the continuous determination window; This is an indicator function that takes the value 1 when the condition is true and 0 when the condition is false. For the first Layer at the time of determination Feed offset category; For the first Layer at the time of determination The feed offset category has values of -1, 0, and 1. -1 indicates a feed offset category that is lower than the corresponding functional constraint, 0 indicates a base feed category that is within the corresponding functional constraint, and 1 indicates a feed offset category that is higher than the corresponding functional constraint.
[0039] When the total film thickness feedback data is within the target thickness range, the film bubble width is within the allowable fluctuation range, the traction speed does not change beyond the allowable amplitude, and the feed offset category of a certain layer remains a non-baseline feed category within the continuous determination window, the feed offset category corresponding to that layer is determined to be a hidden interlayer offset state. The corresponding determination method is as follows: , in, For the first Layer at the time of determination The hidden interlayer offset state; The threshold for determining traction speed stability is calibrated based on the normalized change distribution of traction speed corresponding to the material type identification and production stage status, and is limited to a range greater than 0 and less than or equal to 1. , , , and Refer to the aforementioned definition.
[0040] Furthermore, when writing the hidden inter-layer offset state into the layer flow hidden state, the corresponding layer identifier, feed offset category, and continuous decision window identifier are simultaneously recorded. The layer identifier is used to distinguish between outer, middle, and inner layers, and the continuous decision window identifier is used to determine the retention relationship of the feed offset category in continuous sampling points. When multiple layers meet the hidden inter-layer offset state writing conditions within the same decision period, the processing order used for subsequent candidate control action determination is generated according to the order of inner layer heat sealing feed constraint, middle layer feed stability constraint, and outer layer continuity constraint.
[0041] When the hidden layer offset state is 1, the corresponding layer's feed offset category is written into the layer flow hidden state. If different layers have feed offset categories with opposite directions within the same continuous decision window, then within the set of allowed actions, the candidate control action that reduces the relative feed of the higher layer and increases the relative feed of the lower layer is selected first. When the candidate control action conflicts with the inner layer heat sealing feed constraint, the middle layer feed stability constraint, or the abnormal recovery trigger condition, the safety control action is invoked, or the current control action is used as the target control action.
[0042] When the normalized deviation of total film thickness and the normalized deviation of film bubble width are in the same direction as the material feeding deviation categories of each layer involved in the offset, and the normalized change in traction speed does not offset the change direction corresponding to the normalized deviation of total film thickness, it is determined to be an overall forming deviation state. The corresponding determination method is as follows: like For all Established, For all Established, ,but , Other situations ; in, To determine the time The overall forming deviation state; The threshold for determining the consistency of total film thickness is calibrated based on the normalized deviation distribution of total film thickness corresponding to the material type identification and production stage status, and is limited to a range greater than 0 and less than or equal to 1. The threshold for determining the consistency of membrane bubble width is calibrated based on the normalized deviation distribution of membrane bubble width corresponding to the material type identification and production stage status, and is limited to a range greater than 0 and less than or equal to 1. To determine the time The number of layers involved in the offset is statistically significant; the remaining parameters are defined as described above.
[0043] When the overall forming deviation state is 1, the overall forming deviation state is written into the layer flow hidden state, and the subsequent control priority is shifted to the coordinated adjustment of traction speed, cooling air volume, cooling air temperature, and the common feeding direction of each layer; when the hidden inter-layer offset state is 1 and the overall forming deviation state is 0, the subsequent control priority is shifted to the screw speed, melt temperature state, or filter pressure difference related adjustment of the corresponding layer. This consistency determination is executed once in each determination cycle, and the continuous determination window is called according to the current production stage state; during the abnormal recovery stage, when the abnormal recovery conditions are triggered by melt pressure, filter pressure difference, membrane bubble width, or frost line position, the writing of new hidden inter-layer offset states is paused, and the previous effective layer flow hidden state is maintained or a safety control action is called.
[0044] Based on the above implementation, when the overall forming deviation state is written into the hidden state of layer flow, the normalized deviation direction of the total film thickness, the normalized deviation direction of the bubble width, the normalized change direction of the traction speed, and the statistical value of the number of layers involved in the offset are recorded. When both the hidden interlayer offset state and the overall forming deviation state meet the writing conditions within the same determination period, the hidden interlayer offset state is used as the input for layer function constraint determination, and the overall forming deviation state is used as the reference input for the coordinated adjustment of traction speed, cooling air volume, and cooling air temperature.
[0045] In one optional implementation, the reinforcement learning policy model uses the layer functional state vector and the layer flow hidden state as state inputs, selects actions from the set of allowed actions as actions, and uses hierarchical reward items as the basis for updating. The hierarchical reward items include layer functional constraint satisfaction items, membrane bubble stability items, total membrane thickness deviation items, thermal state constraint items, and action smoothing items. The update priority of layer functional constraint satisfaction items is higher than that of total membrane thickness deviation items. When a target regulation action reduces the total membrane thickness deviation and causes any layer functional constraint to be in an unsatisfied state, the target regulation action is marked as a low-priority action. When a target regulation action maintains the layer functional constraints and makes the membrane bubble width change continuously, the target regulation action is marked as a reusable action.
[0046] Specifically, the layer functional constraint satisfaction item in the tiered reward program is determined based on the layer functional constraint satisfaction status; the membrane bubble stability item is determined based on the continuous changes in membrane bubble width and frost line position; the total membrane thickness deviation item is determined based on the deviation between the total membrane thickness feedback data and the target total membrane thickness; the thermal state constraint item is determined based on the melt temperature status and die head temperature status; and the action smoothing item is determined based on the adjustment direction and adjustment amplitude category of the target control action within adjacent judgment periods. All of the above items are generated within the judgment period after the target control action is executed and participate in the reinforcement learning strategy model update after being bound to the corresponding target control action.
[0047] The offline pre-training samples for the reinforcement learning policy model are derived from historical layer functional state vectors, historical layer flow hidden states, historical regulation actions, historical layer functional constraint satisfaction states, and historical quality feedback data. Training objectives include selecting target regulation actions when layer functional constraints are satisfied, reducing the probability of selecting incorrect regulation actions when there is a deviation between the total membrane thickness feedback data and the layer flow hidden state, and maintaining smooth action output when the membrane bubble state changes continuously. During deployment, the reinforcement learning policy model receives the layer functional state vector, the layer flow hidden state, and the set of allowed actions, and outputs the target regulation action. After the target regulation action is executed, hierarchical reward items are generated based on the updated layer functional constraint satisfaction state, membrane bubble stable state, total membrane thickness deviation state, thermal state constraint state, and action smoothing state, and the reinforcement learning policy model is updated based on these hierarchical reward items.
[0048] In one optional implementation, the production stage includes a start-up stage, a material changeover stage, a steady-state production stage, and an anomaly recovery stage. During the start-up or material changeover stage, the set of actions to be executed is allowed to be output in the order of temperature, screw speed, traction speed, and cooling status. During the steady-state production stage, the set of actions to be executed includes screw speed adjustment actions and melt temperature adjustment actions for each layer. During the anomaly recovery stage, if the melt pressure, filtration differential pressure, membrane bubble width, or frost line position meets the anomaly recovery triggering conditions, the output of new target control actions is paused, and the safety control actions corresponding to the current production stage are invoked. The anomaly recovery triggering conditions are determined by the pressure threshold, differential pressure threshold, membrane bubble width fluctuation threshold, and frost line fluctuation threshold.
[0049] The production stage status is determined based on production instructions, changes in material type identification, continuous sampling data, and equipment alarm records. The start-up stage refers to the process of the extrusion unit heating up, material entering the co-extrusion die, and forming the initial film bubble. The material change stage refers to the process from when at least one layer's material type identification or material batch changes until the layer's functional state vector reaches continuous production conditions. The steady-state production stage refers to the process where the layer's functional state vector, film bubble width, frost line position, and total film thickness feedback data are all within continuous production conditions. The anomaly recovery stage refers to the process after melt pressure, filtration differential pressure, film bubble width, frost line position, or winding tension triggers an anomaly recovery condition. Safety control actions refer to control actions from the set of conservative control actions bound to the production stage status, including at least maintaining the current control action, reducing the control amplitude category, pausing online updates, limiting traction speed control actions, and limiting cooling state control actions. The current control action refers to the target control action that the control system executed in the previous judgment cycle and was not revoked by the anomaly recovery trigger condition.
[0050] In one optional implementation, updating the reinforcement learning policy model includes offline pre-training and online updating. Offline pre-training uses historical layer functional state vectors, historical layer traffic hidden states, historical regulation actions, and historical layer functional constraint satisfaction states to generate offline training samples. Online updating is initiated when the target regulation action belongs to the set of allowed actions and the layer functional constraint satisfaction state after execution is obtained. When the layer functional state vector obtained online exceeds the coverage of the offline training samples, online updating is stopped, and a safe regulation action is output, or the current regulation action is used as the target regulation action.
[0051] For example, the post-execution data used for online updates is multi-source operational data after the target control action has been completed and aligned with the transmission hysteresis time. If any key data among the total membrane thickness feedback data, membrane bubble width, traction speed, melt pressure, and filtration differential pressure is missing from the post-execution data, the current decision cycle generation model update state is paused online, while maintaining the parameter state of the reinforcement learning policy model in the previous effective update cycle.
[0052] The coverage range of offline training samples is jointly determined by the material type identifier, production stage status, the value range of the layered operation sub-state in the layer functional state vector, the value range of the shared forming sub-state, and the range of the layer flow hidden state category. If the layer functional state vector obtained online exhibits at least one of the following: material type identifier mismatch, production stage status mismatch, layered operation sub-state exceeding its corresponding historical range, or shared forming sub-state exceeding its corresponding historical range, it is determined to be outside the coverage range of offline training samples. This determination result is written into the model update state and used to trigger safety control actions, or to use the current control action as the target control action.
[0053] In an optional implementation, the method further includes generating a control traceability record, which includes the closed-loop control time, layer functional state vector, layer flow hidden state, shielded candidate control actions, set of allowed actions, target control action, blown film parameter adjustment command, layer functional constraint satisfaction state, and model update state. When abnormal heat sealing contact, edge wrinkles, winding abnormalities, or interlayer material feeding offset trends occur in the subsequent production process, the corresponding sample record is extracted according to the control traceability record, and the sample record is written into the training samples of the layer flow estimation model or reinforcement learning strategy model.
[0054] The control traceability record is generated sequentially according to the closed-loop control time and is bound to the production batch, material type identifier, and production stage status. The control traceability record uses the closed-loop control time as the index key and uses the production batch, material type identifier, production stage status, layer identifier, target control action, and model update status as retrieval fields. The model update status includes offline pre-training status, online update status, paused online update status, and safety control action invocation status. Sample records refer to data records extracted from the control traceability record and bound to subsequent quality feedback, including at least the layer functional state vector, layer flow latent state, target control action, layer functional constraint satisfaction status, and corresponding quality feedback. Before being used as training samples in the layer flow estimation model or reinforcement learning strategy model, sample records are bound to at least one of the following: offline layer thickness detection data, heat sealing detection data, and membrane bubble forming data.
[0055] Example 2: Based on Example 1, this example provides a specific way to shield candidate control actions that cause the intermediate screw speed or the intermediate melt temperature to change continuously in the same direction in the reinforcement learning-based degradable film bag blown film parameter control method. Candidate control actions refer to individual actions in the set of blown film parameter adjustment actions that the control system can execute. Each candidate control action includes the adjustment object, adjustment direction, adjustment amplitude category, and execution sequence. The adjustment objects include the outer layer extrusion unit, the middle layer extrusion unit, the inner layer extrusion unit, the co-extrusion die, the traction mechanism, and the cooling mechanism. The judgment of candidate control actions includes: before the target control action is output, eliminating candidate control actions that weaken the layer function constraints based on the layer flow hidden state and layer function constraints; when the layer flow hidden state indicates that the inner layer material supply is lower than the inner layer heat sealing material supply constraint, candidate control actions that reduce the inner layer screw speed, increase the traction speed, or enhance cooling and cause a relative decrease in the inner layer material supply are written into the masked candidate control actions.
[0056] In one optional implementation, the candidate control actions include the adjustment of the intermediate layer screw speed and the adjustment of the intermediate layer melt temperature. The determination of the candidate control actions further includes: when the layer flow rate latent state indicates that the intermediate layer feed fluctuation is transmitted to the bubble forming process, the presence of intermediate layer feed fluctuation transmission is determined based on the intermediate layer same-direction fluctuation retention degree determined by the intermediate layer melt pressure and the intermediate layer drive load, and the forming follow indicator of the bubble width or frost line position relative to the change direction of the intermediate layer feed side. When intermediate layer feed fluctuation transmission exists, and the intermediate layer screw speed increment or intermediate layer melt temperature increment is in the same direction as the corresponding increment in the previously executed control action, the corresponding candidate control action is written into the masked candidate control actions.
[0057] Furthermore, the set of allowed actions refers to the set of candidate control actions that are not included in the masked action list. When all candidate control actions are masked, the set of allowed actions consists of maintaining the current control action and a safe control action; when none of the candidate control actions are masked, the set of allowed actions includes all candidate control actions. After the set of allowed actions is output, it is used together with the layer functional state vector and the layer flow hidden state as the inference input of the reinforcement learning policy model.
[0058] Action masking refers to eliminating candidate control actions that weaken layer functional constraints based on the layer flow hidden state and layer functional constraints before the target control action is output. Continuous unidirectional change refers to the same control object maintaining the same control direction within consecutive decision periods.
[0059] In one embodiment of Example 2, when determining whether the fluctuation of the middle layer material supply is transmitted to the bubble forming process, the adjacent sampling changes of the middle layer melt pressure, the middle layer drive load, the bubble width and the position of the frost line are respectively converted into directional quantities with dead zones. Directional variables are used to eliminate minute fluctuations caused by sampling noise, and their conversion method is as follows: , in, It is a direction function with a dead zone; is the change in adjacent samples used for direction determination; is the dead zone threshold corresponding to the change in adjacent samples.
[0060] Take a positive value that is greater than the corresponding sensor sampling resolution and less than or equal to the corresponding allowable process fluctuation amplitude; when there are missing, out-of-bounds, or abnormal sensor markings in the data involved in the direction determination, the corresponding direction quantity is taken as 0, and this direction quantity is used as 0 in the calculation of the mid-layer same-direction fluctuation retention degree and molding follow-up mark.
[0061] The direction of change on the middle layer feeding side and the bubble forming side is obtained based on the direction function, in the following way: , Among them, superscript Indicates a middle-layer extrusion unit, superscript Indicates the location of membrane bubble formation, superscript Indicates the position of the frost line; subscript Indicates melt pressure, subscript Indicates the driving load, subscript Indicates the width of the membrane bubble, subscript Indicates the position of the frost line; To determine the time The direction of pressure change in the middle layer of the melt; To determine the time The pressure of the middle layer melt; To determine the time The pressure of the middle layer melt at the previous judgment moment; This is the threshold value for the dead zone of the middle layer melt pressure. To determine the time The direction of load change in the middle layer; To determine the time The middle layer drives the load; To determine the time The middle layer driver load at the previous decision moment; The dead-time threshold for the middle layer driver load; To determine the time The direction of change in the width of the membrane bubble; To determine the time The width of the membrane bubble; To determine the time The width of the membrane bubble at the previous determination time; The dead zone threshold is the width of the membrane vesicle. To determine the time The direction of change in the position of the frost line; To determine the time The location of the frost line; To determine the time The position of the frost line at the previous judgment moment; The dead zone threshold for the frost line location; Refer to the aforementioned definition.
[0062] Within a continuous decision window, when the mid-layer melt pressure or mid-layer drive load maintains the same direction of change, calculate the mid-layer unidirectional fluctuation retention rate: , in, To determine the time The degree of preservation of mid-level unidirectional fluctuations; This represents the number of sampling points contained in the continuous decision window of the middle layer; The offset number of the sampling point within the continuous decision window of the middle layer; This is an indicator function that takes the value 1 when the condition is true and 0 when the condition is false. To determine the time The direction of pressure change in the middle layer of the melt; To determine the time The direction of load change in the middle layer; and Refer to the aforementioned definition.
[0063] When the direction of change on the middle layer feeding side and the direction of change on the bubble forming side remain in the same direction within the middle layer continuous determination window, a forming follow-up indicator is generated: , in, To determine the time The molding follows the markings; The threshold for maintaining the mid-layer undulation forming is calibrated based on the proportional distribution of the change direction of the mid-layer material supply side and the change direction of the bubble forming side corresponding to the material type identification and production stage status, and is limited to a range greater than 0 and less than or equal to 1. To determine the time The direction of change in the width of the membrane bubble; To determine the time The direction of change in the position of the frost line; , , and Refer to the aforementioned definition.
[0064] When masking candidate control actions, the increments of the middle-layer screw speed and the middle-layer melt temperature in the candidate control action are compared in direction with the corresponding increments in the previous executed control action to obtain the middle-layer continuous unidirectional control masking indicator: , in, Candidate regulatory actions At the time of judgment The middle layer continuous unidirectional control shielding indicator, superscript This indicates that regulatory actions have been taken; Identify candidate regulatory actions; The threshold for the retention of unidirectional fluctuation in the middle layer is calibrated based on the distribution of unidirectional fluctuation retention of the middle layer melt pressure and the middle layer driving load corresponding to the material type identification and production stage status, and is limited to a range greater than 0 and less than or equal to 1. Candidate regulatory actions At the time of judgment The corresponding increase in the speed of the middle screw; To determine the time The speed increment of the middle screw corresponding to the control action executed at the previous judgment moment; Candidate regulatory actions At the time of judgment The corresponding increase in the temperature of the middle layer melt; To determine the time The temperature increment of the middle layer melt corresponding to the control action performed at the previous judgment time; and Refer to the aforementioned definition.
[0065] when When this happens, the corresponding candidate control action is removed from the set of allowed actions. This is used to prevent the continued output of the same direction of the intermediate layer screw speed increment or the intermediate layer melt temperature increment when the intermediate layer melt pressure or intermediate layer drive load has been continuously fluctuating in the same direction and the bubble width or frost line position has changed accordingly. When At that time, the corresponding candidate regulatory actions continue to participate in the action selection of the reinforcement learning strategy model.
[0066] In one example, the direction comparison of the intermediate screw speed increment and the intermediate melt temperature increment is classified into positive adjustment, negative adjustment, and hold adjustment. Hold adjustment does not constitute continuous same-direction adjustment; if the intermediate screw speed adjustment action and the intermediate melt temperature adjustment action were not executed in the previous judgment cycle, the corresponding candidate control action is retained in the set of allowed action actions, and the action selection continues to be performed by the reinforcement learning policy model.
[0067] The mid-layer continuous unidirectional control shielding flag is updated once in each judgment cycle, and the mid-layer continuous judgment window... The window is called according to the production stage status; a shorter window is used in the start-up stage and material change stage to adapt to the temperature and material supply establishment process, and a longer window is used in the steady-state production stage to suppress false shielding caused by single-point disturbances.
[0068] This embodiment , , , , , , , , , and All data are read from the threshold calibration record according to the material type identifier and production stage status. The threshold calibration record is only used to save the material type identifier, production stage status, threshold version identifier, calibration sample period, preset sample quantity lower limit and corresponding value.
[0069] and Take an integer greater than or equal to 2, and reread it from the threshold calibration record after the production stage status or material type identifier is switched, keeping it unchanged within the same continuous judgment window; when the number of calibration samples is lower than the preset lower limit of the number of samples, or when the threshold version identifier is inconsistent with the current version of the reinforcement learning strategy model, freeze the addition and writing of the hidden layer offset state, the overall forming deviation state, and the middle layer continuous unidirectional control shielding identifier, and use the previous valid set of allowed actions or call the safety control action; the previous valid set of allowed actions is the set of allowed actions that have completed the candidate control action judgment in the previous judgment period and have not triggered the abnormal recovery condition.
[0070] If the melt pressure, filtration differential pressure, membrane bubble width, or frost line position meet the abnormal recovery trigger conditions, then stop processing based on the current judgment cycle. The set of allowed actions is rewritten. If a valid set of allowed actions exists in the previous judgment period, actions that are identical to the safety control actions corresponding to the anomaly recovery phase are retained. If no actions are retained, the safety control actions are invoked. If no valid set of allowed actions exists, the safety control actions are invoked. When it is necessary to continue adjusting the total film thickness or the stable state of the film bubble, candidate control actions are selected from the following among the candidate control actions that are not masked by layer function constraints and anomaly recovery trigger conditions: traction speed adjustment, cooling airflow adjustment, cooling air temperature adjustment, outer layer screw speed compensation, and inner layer screw speed compensation.
[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0072] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of this application and form different embodiments. For example, all the embodiments above can be used in any combination. The information disclosed in this background section is intended only to enhance the understanding of the general background of this application and should not be construed as an admission or in any way implying that such information constitutes prior art known to those skilled in the art.
Claims
1. A method for controlling the blown film parameters of a co-extruded biodegradable film bag based on reinforcement learning, wherein the co-extruded biodegradable film bag comprises an outer layer, a middle layer, and an inner layer, and the layer functional constraints include an outer layer continuity constraint, a middle layer material feeding stability constraint, and an inner layer heat-sealing material feeding constraint, characterized in that, include: Acquire multi-source operational data and material type identifiers; A layer functional state vector is generated based on the multi-source operational data and material type identifier; Extract the rotation speed change, load change, melt pressure change, melt temperature change and filtration pressure difference change of each layer from the layer functional state vector, input them into the layer flow estimation model to obtain the feeding offset category of each layer, and combine the total film thickness normalized deviation, film bubble width normalized deviation, traction speed normalized change and feeding offset category in the continuous judgment window to make consistency judgment, and generate the layer flow hidden state including hidden inter-layer offset state and overall forming deviation state. Based on the hidden state of the layer traffic and the layer functional constraints, candidate control actions are determined, and candidate control actions that weaken the layer functional constraints are written into the blocked candidate control actions. The set of allowed actions is output based on the unblocked candidate control actions. The layer functional state vector, the layer traffic hidden state, and the set of allowed actions are input into the reinforcement learning policy model to obtain the target control action. The target control action is converted into blown film parameter adjustment commands and sent to the control system; The reinforcement learning strategy model is updated based on the multi-source running data after execution.
2. The method for adjusting the parameters of a degradable membrane bag blown film based on reinforcement learning according to claim 1, characterized in that, The multi-source operating data includes screw speed, drive load, melt pressure, melt temperature, filtration pressure difference, die temperature, traction speed, cooling air volume, cooling air temperature, film bubble width, frost line position, total film thickness feedback data, winding tension, ambient temperature, and ambient humidity for each extrusion unit. After acquiring the multi-source operating data, source markers are set according to the outer extrusion unit, middle extrusion unit, inner extrusion unit, co-extrusion die, film bubble forming position and traction winding position, and time alignment is performed according to the material transfer lag time from the extrusion unit to the thickness measurement position.
3. The method for adjusting the parameters of a degradable membrane bag blown film based on reinforcement learning according to claim 1, characterized in that, Generating the layer functional state vector includes: The material type identifiers of the outer, middle and inner layers are respectively bound to the screw speed, drive load, melt pressure, melt temperature state and filtration pressure difference of the corresponding extrusion unit to form a layered operation sub-state, and the change direction of each layer at adjacent sampling times is recorded. Write the membrane bubble width, frost line position, total membrane thickness feedback data, traction speed, cooling air volume and cooling air temperature into the shared molding sub-state; Write the startup, material change, steady-state production, and anomaly recovery status into the production stage status; The layered operation sub-state, the shared forming sub-state, and the production stage state are combined to obtain the layer functional state vector.
4. The method for adjusting the parameters of a degradable membrane bag blown film based on reinforcement learning according to claim 1, characterized in that, The layer flow estimation model is trained using historical blown film process data, which includes material type identification, extrusion operation data, film bubble formation data, total film thickness feedback data, offline layer thickness detection data, and heat sealing detection data. During training, offline layer thickness detection data and heat sealing detection data are used as the annotation source for material feeding offset categories. Material change transition data, filter pressure difference increase data and long-term production drift data are written into the training samples, and the corresponding production stage status is retained.
5. The method for adjusting the parameters of a degradable membrane bag blown film based on reinforcement learning according to claim 1, characterized in that, The candidate control action refers to a single action in the set of blown film parameter adjustment actions that the control system can execute. Each candidate control action includes the adjustment object, adjustment direction, adjustment amplitude category and execution order. The adjustment objects include the outer extrusion unit, the middle extrusion unit, the inner extrusion unit, the co-extrusion die, the traction mechanism, and the cooling mechanism; Determining the candidate control action includes: Before the target control action is output, candidate control actions that weaken the layer functional constraints are excluded based on the layer flow hidden state and layer functional constraints. When the layer flow rate latent state indicates that the inner layer feed is lower than the inner layer heat sealing feed constraint, candidate control actions that reduce the inner layer screw speed, increase the traction speed, or enhance cooling and cause the relative inner layer feed to decrease are written into the masked candidate control actions.
6. The method for adjusting the parameters of a degradable membrane bag blown film based on reinforcement learning according to claim 5, characterized in that, The candidate control actions include adjusting the speed of the middle layer screw and adjusting the temperature of the middle layer melt. The determination of the candidate control action also includes: When the layer flow rate latent state characterizes the transmission of the middle layer feed fluctuation to the bubble forming process, the presence of the middle layer feed fluctuation transmission is determined by the middle layer same-direction fluctuation retention degree determined by the middle layer melt pressure and the middle layer drive load, as well as the forming follow indicator of the bubble width or frost line position relative to the changing direction of the middle layer feed side. When the mid-layer feeding fluctuation is transmitted, and the mid-layer screw speed increment or mid-layer melt temperature increment is in the same direction as the corresponding increment in the previously executed control action, the corresponding candidate control action is written into the masked candidate control action.
7. The method for adjusting the parameters of a degradable membrane bag blown film based on reinforcement learning according to claim 1, characterized in that, The reinforcement learning strategy model uses the layer functional state vector and the layer flow hidden state as state inputs, actions from the set of allowed actions as optional actions, and layered reward items as the basis for updating. The tiered reward items include layer functional constraint satisfaction items, membrane bubble stability items, total membrane thickness deviation items, thermal state constraint items, and motion smoothing items. The update priority of the layer functional constraint satisfaction items is higher than that of the total membrane thickness deviation items. When a target control action reduces the total film thickness deviation and causes any layer's functional constraint to be unsatisfied, the target control action is marked as a low priority action. When a target control action maintains the functional constraints of the membrane layer and causes the membrane bubble width to change continuously, the target control action is marked as a reusable action.
8. The method for adjusting the parameters of a degradable membrane bag blown film based on reinforcement learning according to claim 3, characterized in that, The production stages include the start-up stage, material change stage, steady-state production stage, and anomaly recovery stage. During the start-up or material change phase, the set of actions can be output in the following order: temperature, screw speed, traction speed, and cooling status. During the steady-state production phase, the set of actions that can be performed includes screw speed adjustment actions and melt temperature adjustment actions for each layer. During the abnormal recovery phase, if the melt pressure, filtration differential pressure, membrane bubble width, or frost line position meet the abnormal recovery triggering conditions, the output of new target control actions will be paused, and the safety control actions corresponding to the current production stage will be invoked. The abnormal recovery triggering conditions are determined by the pressure threshold, differential pressure threshold, width fluctuation threshold, and frost line fluctuation threshold.
9. The method for adjusting the parameters of a degradable membrane bag blown film based on reinforcement learning according to claim 1, characterized in that, Updating the reinforcement learning policy model includes offline pre-training and online updating; Offline pre-training uses historical layer functional state vectors, historical layer traffic hidden states, historical regulation actions, and historical layer functional constraint satisfaction states to generate offline training samples. Online updates are initiated when the target control action belongs to the set of allowed actions and the state of the layer function constraints satisfied after execution is obtained. When the layer functional state vector obtained online exceeds the coverage of the offline training samples, online updates are stopped, and a safe adjustment action is output, or the current adjustment action is used as the target adjustment action.
10. The method for adjusting the parameters of a degradable membrane bag blown film based on reinforcement learning according to claim 1, characterized in that, It also includes generating control traceability records, which include closed-loop control time, layer function state vector, layer flow hidden state, blocked candidate control actions, set of allowed actions, target control action, blown film parameter adjustment instructions, layer function constraint satisfaction state and model update state; When abnormal heat sealing contact, edge wrinkles, abnormal winding, or interlayer feeding deviation occurs in the subsequent production process, the corresponding sample record is extracted according to the control traceability record, and the sample record is written into the training sample of the layer flow estimation model or reinforcement learning strategy model.