Industrialized production method of recycled plastic high-friction geogrid
By combining the whole-line colored Petri net model with multi-agent reinforcement learning, the method explicitly expresses the cross-station cycle coupling and constructs cycle safety check constraints, which solves the problems of uneven cycle and insufficient material adaptability in the production of recycled plastic high-friction geogrid, and realizes the improvement of production stability and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG HI SPEED CONSTRUCTION MANAGEMENT GROUP CO LTD
- Filing Date
- 2026-06-10
- Publication Date
- 2026-07-14
Smart Images

Figure CN122389655A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial automation technology, and in particular to an industrial production method for recycled plastic high-friction geogrid, which belongs to the field of industrial control software. Background Technology
[0002] Geogrids, as an important material in reinforced soil engineering, are widely used in highway subgrade reinforcement, railway embankments, surcharge preloading foundations, and ecological slope protection. The interfacial friction performance between the geogrid and the soil directly determines the load-bearing capacity and long-term deformation stability of the reinforced structure. In recent years, driven by both resource recycling and carbon emission reduction considerations, the preparation of welded geogrids using recycled plastic granules as raw materials has become an important development direction in the industry. However, recycled plastics suffer from significant batch-to-batch variations, large fluctuations in melt flow rate and impurity content, which can easily lead to cycle time mismatches between continuous production processes such as melt mixing, die extrusion, shaping and cooling, orientation stretching, ultrasonic welding of joints, and tension winding.
[0003] Several attempts have been made in the industry to address the aforementioned cycle time balancing problem. One approach is based on proportional-integral-derivative feedback control at the workstation level, suppressing speed fluctuations through local closed-loop control of the workstation servo speed ratio. However, relying solely on feedback from the workstation itself fails to represent cross-workstation coupling, resulting in generally large fluctuations in cycle time intervals. Another approach uses discrete event simulation to schedule the cycle time of the entire production line, driven by historical samples. This requires recalibration for situations where incoming material attributes change abruptly, limiting its online adaptability. Yet another approach introduces deep reinforcement learning for centralized control of the entire line. However, the policy network typically does not explicitly incorporate the material topology of the entire line, and it integrates multiple quality and capacity indicators in the training objective using a weighted sum method, easily falling into a degradation problem of sacrificing quality for output. None of these approaches provide a hard constraint on the physical feasibility of cycle time at the control strategy level. The policy output may require the entire line to output material at a rate that is physically impossible for the bottleneck loop, leading to material shortages in the buffer warehouse or insufficient processing in the processing warehouse.
[0004] It is evident that existing technologies in the industrial production of recycled plastic high-friction geogrids still suffer from problems such as insufficient cross-station cycle time coupling expression, lack of cycle time feasibility constraints, bias in strategy training objectives, and poor online adaptability to material attribute drift. There is an urgent need to propose a new production method that can analyze the lower bound of the entire line cycle time and construct cycle time safety check constraints accordingly. Summary of the Invention
[0005] The purpose of this invention is to provide an industrial production method for high-friction geogrids based on recycled plastics. This invention explicitly expresses cross-station cycle coupling, analyzes the lower bound hard constraint strategy output, and optimizes the process with a multi-class qualified signal necessary condition screening reward-guided strategy, which significantly improves the overall line cycle balance, effective capacity utilization, and the dimensional stability of non-uniform ribs and the reliability of welded node connections.
[0006] To solve the above-mentioned technical problems, the present invention provides an industrial production method for recycled plastic high-friction geogrid, the method comprising the following steps: Step 1: Deploy 6 processing stations sequentially along the material's forward direction, with variable buffer units between adjacent processing stations; each processing station is equipped with a local control unit, which collects the station's status vector and reports it to the edge server via a time-sensitive network; each batch of recycled plastic granules is assigned a batch identifier. Step 2: Establish a full-line colored Petri net model in the edge server, using processing stations as processing warehouses, variable buffer units as buffer warehouses, material transfer actions between processing warehouses and buffer warehouses as transitions, and batch identification as token colors; perform iterative exponentiation operations on the excitation times of transitions in the full-line colored Petri net model based on Max-plus algebra to obtain the analytical saturation beat lower bound of the entire line under the current configuration; Step 3: Construct a graph attention network in the edge server as the policy backbone for multi-agent reinforcement learning. Use the topology of the whole line colored Petri net model as the input graph and the lower bound of the analytic saturation beat as the beat safety check constraint. Use the QMIX algorithm to train the graph attention network. After training, the edge server runs the graph attention network according to the real-time status of the whole line in each control cycle. The output execution actions are sent to the local control unit of each workstation for execution via the time-sensitive network.
[0007] Furthermore, the six processing stations, arranged sequentially along the material's forward direction, are: melt compounding station, die extrusion station, shaping and cooling station, orientation stretching station, node ultrasonic welding station, and tension winding station. In the melt compounding station, recycled plastic granules are plasticized into a melt using a twin-screw extruder. The die extrusion station employs a dual-die structure, with the two dies extruding longitudinal and transverse ribs respectively. Both the longitudinal and transverse ribs have alternating raised friction sections and transition connecting sections along their length. The convex height and / or cross-sectional area of the rib surface of the raised friction section are greater than those of the transition connection section. The raised friction section is used to form a mechanical interlocking interface in the soil. The raised friction section and the transition connection section together constitute a continuous billet with non-uniform ribs. The shaping and cooling station performs water bath cooling and air cooling shaping on the longitudinal ribs and transverse ribs respectively. The orientation stretching station performs stretching and orientation treatment on the longitudinal ribs and transverse ribs respectively to obtain longitudinally oriented ribs and transversely oriented ribs. The tension winding station performs constant tension winding on the finished product.
[0008] Furthermore, the node ultrasonic welding station includes a cross-laying substation and an ultrasonic welding substation; the cross-laying substation includes a transverse rib fixed-length cutting module, a transverse transfer module, and a grid positioning module. The transverse rib fixed-length cutting module cuts the transverse oriented ribs into transverse ribs of a preset length. The transverse transfer module transfers the transverse ribs to the top of the longitudinal oriented ribs. The grid positioning module cross-overlays the transverse ribs with multiple longitudinal oriented ribs according to a preset grid spacing. The ultrasonic welding substation performs ultrasonic welding on the cross-overlay positions to form a welded node.
[0009] Furthermore, five variable buffer units are set between two adjacent processing stations; the variable buffer unit between the melt mixing station and the die extrusion station is a melt pressure stabilizing buffer unit, and its capacity parameter is controlled by the displacement of the servo regulating valve of the gear pump flow stabilizing chamber; the remaining four variable buffer units are tension floating material storage buffer units, and their capacity parameters are controlled by the servo displacement of the liftable floating idler; the capacity parameters of the melt pressure stabilizing buffer unit and the tension floating material storage buffer unit are uniformly recorded as the displacement of the servo regulating mechanism.
[0010] Furthermore, except for the tension winding station, the local control units of the other five processing stations periodically collect the servo torque, melt pressure or strip tension, process temperature, unit time output length calculated by the output encoder pulse, and the displacement of the servo adjustment mechanism of the downstream variable buffer unit of the station to form a station status vector; the local control unit of the tension winding station periodically collects the winding tension, winding roll diameter and roll length count to form a station status vector, and outputs a tension stability qualified signal; all station status vectors are reported to the edge server through a time-sensitive network at a fixed control cycle.
[0011] Furthermore, an online contour detection unit and a node welding qualification detection unit are set up before the tension winding station; the online contour detection unit collects the rib height, rib spacing and rib protrusion cross-sectional dimensions, and outputs a rib contour qualification signal; the node welding qualification detection unit collects the infrared thermograph of the welding node, ultrasonic welding energy, welding head downward displacement and welding pressure, and outputs a node welding qualification signal in combination with the node pull-out force calibration results obtained from periodic sampling inspections.
[0012] Furthermore, each batch of recycled plastic granules entering the melt blending station is assigned a batch identity identifier, which is determined by the incoming material spectrum identification result and the feeding time of this batch. The position of the batch identity identifier in the entire line is estimated in real time by the edge server based on the feeding time of this batch, the screw speed of the twin-screw extruder, the melt pressure, the melt flow rate, and the discharge length per unit time of each processing station, so as to obtain the current location of the batch identity identifier and the number of tokens it occupies.
[0013] Furthermore, in the whole-line colored Petri net model, one token represents one unit length of material. Max-plus algebra uses the maximum value operation instead of conventional addition and superposition operation instead of conventional multiplication. The whole-line colored Petri net model establishes longitudinal ribbed subnets and transverse ribbed subnets between the die extrusion station and the node ultrasonic welding station. Tokens in the longitudinal ribbed subnet represent one unit length of longitudinal material, and tokens in the transverse ribbed subnet represent one unit length of transverse material. The cross-laying sub-stations of the node ultrasonic welding station correspond to one merging transition. The merging transition is determined by the tokens in the buffer pool at the end of the longitudinal ribbed subnet and the buffer pool at the end of the transverse ribbed subnet. The tokens in the system serve as common inputs. The merging transition is only permitted to be activated when the number of tokens in both end buffer stations meets the preset token quantity condition corresponding to the grid spacing. After activation, one grid cell token is generated in the processing station corresponding to the ultrasonic welding station at the node. A minimum dwell time delay is defined for each processing station and each buffer station in the overall colored Petri net model. The minimum dwell time delay of a processing station is equal to the measured time required for the corresponding processing station to complete the processing of one unit length of material under calibrated conditions. The minimum dwell time delay of a buffer station is equal to the time required for the material to pass through the corresponding variable buffer unit at the maximum conveying rate under the current servo adjustment mechanism displacement. The measured time; for each buffer, a maximum token capacity and a minimum retention capacity are defined simultaneously. The maximum token capacity and minimum retention capacity are updated in linkage with the displacement of the servo adjustment mechanism of the corresponding variable buffer unit; upstream transitions of a buffer are only permitted to be activated when the number of tokens in the buffer is lower than the maximum token capacity, and downstream transitions of a buffer are only permitted to be activated when the number of tokens in the buffer is higher than the minimum retention capacity; the rule for determining the next activation time of a transition is as follows: for each input buffer of this transition, the activation time of the previous transition corresponding to this input buffer is superimposed with the minimum dwell time of this input buffer to obtain the superposition sum, and this sum is applied to all input buffers of this transition. The sum of the sums and the maximum value is taken as the next excitation time of this transition. For transitions that do not meet the permitted excitation conditions, the transition is marked as prohibited from excitation in the current iteration round, and the update of the excitation time of this transition is paused. The iterative exponentiation operation starts from an initial excitation time vector of zero vector and repeatedly updates the excitation times of all transitions in the entire line that meet the permitted excitation conditions according to the rule of taking the next excitation time value of the transition, so as to obtain the excitation time vector sequence. When the difference between the maximum and minimum values of the component differences of the excitation time vectors of two adjacent rounds is less than the preset convergence threshold, it is determined that the iteration has converged, and the average value of the component differences of the excitation time vectors of two adjacent rounds is taken as the lower bound of the analytical saturation beat of the entire line under the current configuration.Whenever the displacement of the servo adjustment mechanism of any variable buffer unit changes or the measured minimum dwell time of any processing station changes, the edge server re-executes an iterative exponentiation operation and refreshes the lower bound of the analytical saturation beat rate. The refreshed lower bound of the analytical saturation beat rate is broadcast to all local control units of all workstations via a time-sensitive network.
[0014] The industrial production method of recycled plastic high-friction geogrid of the present invention has the following beneficial effects: By modeling longitudinal and transverse ribbed subnets in parallel and rigorously characterizing the grid spacing conditions during cross-laying through confluence transitions, the coupling between the overall material topology and cross-station beat rate is explicitly expressed; based on Max-plus algebra, iterative exponentiation of the excitation times of transitions in the overall colored Petri net model is performed to obtain the lower bound of the overall analytical saturated beat rate, which is then directly applied as a beat rate safety check constraint to the multi-agent reinforcement learning training objective, ensuring that the execution actions output by the graph attention network do not require... The entire production line aims to output material at a rate that is physically impossible to achieve due to the bottleneck loop, fundamentally avoiding material shortages in the buffer warehouse and insufficient processing in the processing warehouse. The effective finished product length, selected based on node welding qualification signals, rib profile qualification signals, and tension stability qualification signals, serves as a single-step reward, avoiding the quality-for-output degradation strategy resulting from weighted combinations of multiple indicators. Using batch identification as token color and color histogram encoding as graph node input features, the system exhibits excellent online adaptability to drifting properties of recycled plastic materials, significantly improving the overall line cycle time balance, effective capacity utilization, and finished product quality consistency. Attached Figure Description
[0015] Figure 1 A schematic diagram of the overall process layout for the industrial production method of recycled plastic high-friction geogrid provided in this embodiment of the invention. Figure 2 The convergence process of Max-plus algebraic iterative exponentiation operation under three typical full-line configurations provided in the embodiments of the present invention; Figure 3 The graph attention network and QMIX algorithm control strategy used in this invention and the traditional PID control strategy are compared at the transition excitation time interval, as provided in the embodiments of this invention. Figure 4 The present invention provides a comparative curve showing the evolution of the effective finished product length at the tension winding station over production time during an 8-hour continuous production process. Detailed Implementation
[0016] The entire production line is arranged in a straight line along the material's direction of travel. Taking a demonstration line with an annual capacity of approximately 6 million square meters as an example, the total length is about 70 meters, and the equipment base is divided into independent foundations according to the workstations. The straight arrangement facilitates the constant movement of the melt and continuous billet, allowing each set of lifting and floating idlers to participate in tension balance by gravity in the vertical direction. In cases where the workshop length is limited, an L-shaped or U-shaped arrangement can be adopted, with deflection roller sets and temperature transition chambers added at the corners. The six processing stations arranged sequentially along the material's direction of travel are the melt mixing station, die extrusion station, shaping and cooling station, orientation stretching station, node ultrasonic welding station, and tension winding station.
[0017] The twin-screw extruder in the melt compounding station uses a co-rotating twin-screw configuration with a length-to-diameter ratio of 44:1. The screw diameter can be selected from 75mm to 120mm. The barrel is divided into 6 to 8 independent temperature control zones along the axial direction, and the typical feed rate is 80kg / h to 220kg / h. In the die extrusion station, the two dies are respectively machined with longitudinal and transverse ribs. Each die outlet surface has several micro-grooves with non-uniform ribbed cavities. The cavity depth of the raised friction section is 0.5mm to 1.2mm higher than that of the transition section cavity. A piezoelectric melt pressure sensor with a range of 0 to 50MPa is installed at the die inlet. The shaping and cooling station first passes through a constant temperature water bath section at 22℃, followed by an air-cooling section. The orientation stretching station feeds the cooled longitudinal and transverse ribs into the longitudinal stretching rollers and transverse clamping chains, respectively. The longitudinal orientation stretching ratio is typically 5:1 to 6.5:1, and the transverse orientation stretching ratio is typically 3.5:1 to 4.5:1. The ultrasonic welding station uses an ultrasonic frequency of 20kHz or 40kHz, an output power of 2.4kW to 4kW, a welding head pressing stroke of 0.6mm to 1.5mm, and a holding time of 80ms to 150ms. The tension winding station uses a direct-drive servo winding spindle with a roll diameter range of 1.5m to 2.5m and a winding tension closed-loop control range of 80N to 150N.
[0018] The process layout of the industrial production line for recycled plastic high-friction geogrid provided in this embodiment is as follows: Figure 1As shown, six processing stations are drawn from left to right along the material's forward direction: melt compounding station, die extrusion station, shaping and cooling station, orientation stretching station, node ultrasonic welding station, and tension winding station. The melt compounding station and tension winding station are located at opposite ends of the line: the former, represented by a rounded rectangle indicating the structural boundary of the twin-screw extruder, is responsible for melting and homogenizing recycled plastic granules and friction fillers; the latter, represented by concentric multiple rings indicating the finished roll being wound on the winding spindle, with the smallest solid dot inside indicating the center of the winding spindle. The die extrusion station is located downstream of the melt compounding station. At the die exit, two small triangles side by side indicate the dual-die structure, with the two dies marking the extrusion of longitudinal and transverse ribs respectively, ensuring that the longitudinal and transverse ribs are extruded synchronously and continuously at the same die exit. The shaping and cooling station, located after the die extrusion station, performs dimensional shaping and temperature gradient cooling on the extruded strip. The orientation and stretching station uses graded rollers to illustrate the bidirectional orientation effect of longitudinal and transverse stretching, ensuring that the longitudinal and transverse ribs achieve the required tensile modulus and frictional surface texture after orientation. The nodal ultrasonic welding station, located after the orientation and stretching station, uses multiple parallel ultrasonic welding heads to mark the nodal welding positions evenly distributed along the width direction. Each ultrasonic welding head performs nodal ultrasonic welding at the intersection of the longitudinal and transverse ribs to form a complete mesh-like geogrid. (The last sentence appears to be incomplete and possibly refers to a different process.) Figure 1 The overall process layout shown in this embodiment clearly defines the material movement direction, cycle time synchronization, and geometric correspondence among the six processing stations in one go, so that the subsequent overall cycle time optimization method based on time-colored Petri nets and the field monitoring method based on edge servers have clear physical references in the process logic.
[0019] Downstream from the extrusion station at the die head, the material feed path branches into two parallel paths, labeled as the longitudinal ribbed path and the transverse ribbed path, respectively. Each path connects one shaping and cooling station and one orientation stretching station in series. The connecting lines downstream of the two orientation stretching stations converge at the input end of the node ultrasonic welding station. The corresponding cross-laying sub-stations then lay the transverse ribs onto the longitudinal orientation ribs at a preset grid spacing.
[0020] Figure 1In this system, one variable buffer unit is set between every two adjacent processing stations, for a total of five. The melt pressure stabilizing buffer unit between the melt mixing station and the die extrusion station is illustrated by a horizontally placed ellipse with internal melt ripples. The remaining four variable buffer units are tension floating material storage buffer units, located between the shaping and cooling station and the orientation stretching station on the longitudinal ribbed path, between the orientation stretching station and the node ultrasonic welding station on the longitudinal ribbed path, at the corresponding positions on the transverse ribbed path, and between the node ultrasonic welding station and the tension winding station. Each tension floating material storage buffer unit is illustrated by a horizontal three-roller structure. The two solid small circles at the top represent the feed guide roller and the discharge guide roller, and the hollow circle at the bottom center represents the liftable floating idler roller. The three rollers are connected by a broken line to represent the strip winding path. The bidirectional vertical arrows on both sides of the floating idler roller indicate the servo displacement stroke of the idler roller in the vertical direction.
[0021] Figure 1 The edge server is drawn as a rounded rectangle at the top center, with a double horizontal line below representing the Time-Sensitive Network (TSN) backbone. Each processing station is connected vertically to the TSN backbone by a thin dashed line, with a solid dot near the connection point representing the station's local control unit's data acquisition point. This connection corresponds to the mechanism where the station's local control unit collects the station's state vector, reports it to the edge server via the TSN, and the edge server distributes the calculated execution actions to each station's local control unit via the TSN downlink.
[0022] Of the five variable buffer units, one located between the melt mixing station and the die extrusion station is a melt pressure stabilizing buffer unit. Its internal cavity volume is approximately 3L to 8L. The effective flow area of the cavity is adjusted by the displacement of a servo regulating valve at the inlet of the gear pump's flow stabilizing chamber. The stroke of the servo regulating valve is controllable within the range of 0mm to 25mm. This melt pressure stabilizing buffer unit is designed because the instantaneous melt output of the twin-screw extruder experiences low-frequency disturbances originating from feed fluctuations and screw pulsations, with typical pulsation amplitudes reaching 5% to 10%. The volumetric damping of the buffer chamber can suppress melt output pulsations to below 1.5%. The remaining four tension-floating storage buffer units consist of a transverse three-roll structure comprising a feed guide roller, a liftable floating idler roller, and a discharge guide roller. The liftable floating idler roller is driven vertically by a servo linear motor with a stroke of 0.2m to 0.6m. A magnetostrictive displacement sensor is used to obtain the vertical position of the idler roller, with a resolution of 0.05mm. The liftable floating idler allows for short-term differences in conveying speed between adjacent processing stations. This difference is absorbed by the vertical stroke of the floating idler, thus avoiding tensile strain impact caused by short-term speed mismatch. The displacement of the servo regulating valve of the melt pressure stabilizing buffer unit and the displacement of the liftable floating idler of the tension floating material storage buffer unit are uniformly recorded as the displacement of the servo regulating mechanism at the data acquisition level.
[0023] Each processing station is equipped with one local control unit. The hardware uses an industrial-grade programmable logic controller or an industrial control computer, and the real-time operating system uses an industrial real-time kernel that supports a time-sensitive network protocol stack. The local control unit is responsible for pulse planning of the servo drive and sensor data acquisition for its station. Servo torque, melt pressure, and strip tension are sampled at 1kHz, while process temperature, roll diameter, and roll length are sampled at 10Hz to 100Hz. Except for the tension winding station, the local control units of the other five processing stations combine the servo torque, melt pressure or strip tension, process temperature, the unit time output length converted from the output encoder pulse, and the displacement of the servo adjustment mechanism of the downstream variable buffer unit into a station state vector within each fixed control cycle. The servo torque is taken from the torque feedback register of the servo drive; the unit time output length is obtained by dividing the original encoder pulse count by the encoder line count (typically 2048 lines / revolution to 5000 lines / revolution), multiplying it by the circumference corresponding to one roll revolution, and finally dividing it by the sampling period. The local control unit at the tension winding station independently collects the winding tension, winding diameter, and winding length counts to form a station status vector. Based on whether the sliding variance of the winding tension is below a preset threshold, it outputs a tension stability and qualification signal. Sliding variance is used instead of instantaneous deviation because high-friction geogrids experience periodic tension disturbances during winding due to rib interlacing; instantaneous deviations can easily misinterpret normal rib disturbances as non-qualification signals. All station status vectors are reported to the edge server via a time-sensitive network at a fixed control cycle of 10ms to 20ms.
[0024] The time-sensitive network (TSN) employs the IEEE 802.1AS time synchronization protocol and the IEEE 802.1Qbv time-aware shaping protocol based on Ethernet, with a backbone bandwidth of 1Gbps and network jitter controlled within 1μs. This ensures that the status vectors of the workstations across six processing stations can be aligned to the same control cycle in the edge server. The edge server is an industrial server configuration, equipped with a 32-core general-purpose CPU and one to two general-purpose GPUs, 64GB to 128GB of memory, and a Linux operating system with the PREEMPT_RT patch.
[0025] Each batch of recycled plastic granules entering the melt blending station is scanned online by a near-infrared spectroscopy analyzer near the feed inlet of the loss-in-weight feeder, with a scanning wavelength range of 1000 nm to 2500 nm. Partial least squares discriminant analysis is used to match the obtained 256-channel reflectance spectrum with a pre-established calibration library of polyethylene, polypropylene, polyethylene terephthalate, and their mixtures to obtain the incoming material spectral identification result for this batch, including the main component resin type, typical melt flow rate range, and impurity level. The identification result, together with the batch's feed time, generates a 128-bit batch identifier. The first 64 bits are the hash of the incoming material spectral identification result, and the last 64 bits are the integer representation of the feed time in microseconds. Only by carrying both spectral and temporal identifiers can a specific batch be uniquely located during the overall production line simulation.
[0026] The batch identification identifier's position throughout the production line is advanced station by station along the material's direction of travel by the edge server, adhering to a mass conservation principle. Between the melt mixing station and the die extrusion station, due to axial mixing within the twin-screw extruder, the batch material does not advance downstream in an ideal step-like manner, but rather with a gradual change in residence time distribution. The edge server uses the batch's feed time as the initial zero point of its leading edge position, following... The leading edge position of this batch will be advanced downstream over time. For the first The leading edge of the batch identification mark is located at time [time] in the coordinate system along the extruder axis. The location is in millimeters; To control the cycle, a typical value is 10ms; For a moment The instantaneous volumetric flow rate is obtained by dividing the mass flow rate measured by the screw speed, melt pressure, and melt flow sensor of the twin-screw extruder by the melt density, with the unit being cubic millimeters per second; This represents the effective flow channel cross-sectional area of the screw at the current position of the leading edge in this batch, obtained from the screw geometry parameter table, in square millimeters, with common values between 800 and 1500 square millimeters. The equation is numerically derived periodically using an explicit Euler scheme. When... When crossing the actual geometric position of a certain processing station, the edge server subtracts the number of tokens per unit length in the corresponding storage space of the batch identity identifier in the whole line colored Petri net model from the upstream storage space and adds it accordingly in the downstream storage space, thereby completing the jump of tokens between storage spaces, that is, the true source of the storage space where the batch identity identifier is currently located and the number of tokens it occupies.
[0027] Below the die extrusion station, the melt has cooled into a strip-shaped billet, which can be regarded as an approximately rigid material movement process. Therefore, the batch advance along the material advance direction degenerates into a plug flow advance with a rate of the output length per unit time of this station. For the two independent material flows of longitudinal ribs and transverse ribs, the edge server maintains two sets of batch front queues respectively. When their respective fronts reach the cross-laying substation of the ultrasonic welding station, they are merged into the finished batch front.
[0028] In another feasible implementation, the batch front advancement between the melt mixing station and the die extrusion station can be replaced by an active calibration method based on fluorescent tracers, instead of the aforementioned mass conservation calculation. Each batch of feed carries approximately 0.05‰ of a fluorescent tracer by mass, which can be uniformly dispersed in the melt and has negligible impact on the mechanical properties of the finished product. A UV-excited laser fluorescence detector is placed on the die exit side to directly measure the actual time when the batch front reaches the die exit. Based on this, the actual parameters of the residence time distribution inside the extruder can be retrieved, which is suitable for applications requiring high batch traceability accuracy. In yet another implementation, the fixed control cycle of the station state vector can be shortened to 5ms or extended to 50ms. The former is suitable for applications with linear velocities greater than 20m / min and sensitive to tension fluctuations, while the latter is suitable for applications with linear velocities less than 10m / min and limited edge server computing power.
[0029] The full-line colored Petri net model resides in memory on the edge server as a composite data structure of adjacency list and attribute table. It is loaded once during the full-line startup phase, and the model instances remain in memory as the production line continues to run. The formal skeleton of the model can be denoted as... ,in For the collection of warehouses, For the set of changes, Let arcs represent the set of directed connections from places to transitions and from transitions to places. For the set of token colors, This is the initial identifier. Used to characterize the existing token distribution in each storage compartment at the moment of start-up of the entire line: the initial value of the token count in the processing compartment corresponding to the melt mixing station is 0, and the initial value of the token count in the buffer compartment corresponding to each tension floating material storage buffer unit is obtained by dividing the material storage length calculated by the liftable floating idler at the start position by the unit length.
[0030] Treasury Collection The system is divided into three disjoint subsets: the processing warehouse subset, the buffer warehouse subset, and the boundary warehouse subset. The processing warehouses correspond one-to-one with six processing stations, denoted as the melt mixing warehouse, die extrusion warehouse, shaping and cooling warehouse, orientation stretching warehouse, node ultrasonic welding warehouse, and tension winding warehouse. The buffer warehouses correspond one-to-one with five variable buffer units, the first of which corresponds to the melt stabilizing buffer unit, and the other four to the four material storage buffer units corresponding to the tension floating material storage buffer units. There are two boundary warehouses: the upstream material source warehouse and the downstream finished product collection warehouse. In the former, token generation follows a sequence of events assigned to batch identification, while in the latter, all tokens flowing out from the tension winding station are absorbed. Between the die extrusion station and the node ultrasonic welding station, the longitudinal ribs and the transverse ribs are two physically independent and parallel material flows. Therefore, separate longitudinal rib subnets and transverse rib subnets are established in this section. Each subnet contains one shaping and cooling chamber, one orientation and stretching chamber, and two material storage and buffer chambers. Parallel subnets are used instead of merging chambers to retain the key constraint that the two material flows must each meet the orientation and stretching conditions before cross-laying.
[0031] Set of Changes This is used to characterize two types of material transfer actions: the first type is the discharge action, where after processing a unit length of material in a certain processing warehouse, the corresponding token is transferred from that processing warehouse to the downstream buffer warehouse; the second type is the feeding action, where when a unit length of material is output from a certain buffer warehouse to the downstream processing station, the corresponding token is transferred from that buffer warehouse to the downstream processing station. Both types of actions are abstracted using transition excitation. A confluence transition is added at the cross-laying substation of the nodal ultrasonic welding station. The tokens in the buffer warehouse at the end of the longitudinal rib subnet and the buffer warehouse at the end of the transverse rib subnet are used as common inputs. After excitation, a grid cell token is generated in the nodal ultrasonic welding warehouse, representing the smallest welding grid cell that will be performed in the subsequent ultrasonic welding substation. The permissible activation condition for the merging transition is that the number of tokens in both end buffer pools simultaneously reaches the number corresponding to the preset grid spacing—for example, the preset grid spacing corresponds to accumulating 12 unit-length tokens in the longitudinal ribbed subnet end buffer pool and 8 unit-length tokens in the transverse ribbed subnet end buffer pool. After activation, 12 tokens are deducted from the longitudinal ribbed subnet end buffer pool, 8 tokens are deducted from the transverse ribbed subnet end buffer pool, and 1 grid cell token is generated in the node ultrasonic welding pool. This activation rule is equivalent to "welding can only begin when both materials are ready for the current welding," thus avoiding situations where one material arrives first but the other is not yet ready, yet welding is forcibly started.
[0032] Token color set Using batch identification as an element, each token is colored with a 128-digit batch identification assigned to it at the melt compounding station during the generation of self-recycled plastic granules. Color as a token attribute allows multiple tokens from different incoming batches to be accommodated simultaneously in the same storage area, accurately reflecting the physical reality of consecutive batches in the axial mixing zone of the twin-screw extruder. Tokens in the melt pressure stabilizing buffer storage area and the four material storage buffer storage areas are arranged in a first-in, first-out (FIFO) order, consistent with the plug flow of melt in the gear pump's flow stabilizing chamber and the segmented winding of the strip into and out of the tension floating material storage buffer unit.
[0033] Define a minimum dwell time for each processing location and each buffer location in the full-line colored Petri net model, denoted as processing location. The minimum dwell time is , buffer zone The minimum dwell time is All units are seconds. This equals the measured time required for the corresponding processing station to complete the processing of one unit length of material under calibrated working conditions. Taking the die head extrusion unit as an example, a typical... The value ranges from approximately 0.6s to 1.2s, and is specifically determined by the geometry of the die head flow channel and the target linear velocity. This is equal to the measured time required for the material to pass through the corresponding variable buffer unit at its maximum conveying rate under the current servo adjustment mechanism displacement. Taking a tension-floating material storage buffer unit as an example... The value typically ranges from 1s to 15s, changing linearly with the displacement of the liftable floating idler. and The actual measurement process is completed in advance during the production line inspection stage and written into the calibration table. During online operation, the table is directly consulted to avoid additional uncertainties introduced by online identification. Each buffer unit simultaneously defines a maximum token capacity and a minimum retention capacity. Taking a tension-floating material storage buffer unit as an example, the material storage length is short at the highest position and long at the lowest position of the lifting floating roller. Based on this, the maximum token capacity and minimum retention capacity are updated in conjunction with the current displacement of the servo adjustment mechanism. The minimum retention capacity reserves a certain number of tokens as tension buffer margin to prevent downstream processing stations from experiencing material shortages due to the buffer unit instantly dropping to zero. The maximum token capacity is a true constraint on the physical space of the variable buffer unit, preventing the lifting floating roller from reaching its physical upper limit or the internal pressure of the melt stabilizing buffer unit from exceeding the limit. Therefore, upstream changes in the buffer unit are only permitted when the number of tokens in the buffer unit is lower than the maximum token capacity, and downstream changes in the buffer unit are only permitted when the number of tokens in the buffer unit is higher than the minimum retention capacity.
[0034] In the fully colored Petri net model, define an excitation time variable for each transition, and denote the transition. The The next activation time is The unit is seconds. The rule for determining the next trigger time of the transition is expressed using Max-plus algebra. In Max-plus algebra, the maximum value operation replaces regular addition, and the superposition operation replaces regular multiplication, using symbols... The operation of finding the maximum value is represented by the symbol. This represents a superposition operation, therefore for two scalars... and , , The neutral element of the Max-plus algebra is and That is, any number and The largest number is itself, and any number plus... The superposition remains the same as the number itself. The most direct benefit of introducing Max-plus algebra is that it unifies the discrete dynamics, which originally required branching judgments to determine "when an event occurs," into linear dynamics. This allows cycle analysis, bottleneck identification, and stability analysis to be completed using the mature tool of matrix eigenvalues, without relying on event-by-event playback in event-driven simulations.
[0035] The iterative convergence process of the transition excitation time vector optimization method described in this embodiment on three different full-line configurations is as follows: Figure 2 As shown. Figure 2 The horizontal axis represents the iteration round. The scale is linear, ranging from 0 to 100, with the vertical axis representing the range of the vector component differences between the two excitation times. The unit is seconds, using a logarithmic scale with a base of 10. The logarithmic scale intervals are... to The selection of this interval makes The entire process, spanning more than three orders of magnitude, can be clearly demonstrated. Figure 2 The preset convergence threshold is drawn using a horizontal dashed line that runs through the entire graph. The three iterative convergence curves are distinguished by different symbols to correspond to full-line configurations A, B, and C. Hollow circles are marked at the intersections of the three curves with the horizontal dashed line, with the corresponding convergence cycle and the lower bound of the analytical saturation beat noted next to each circle: Full-line configuration A... , ; Configuration B of the entire line , The entire line is configured with C. , The three leader lines point to the intersection of the curve and the horizontal dashed line, and the annotation boxes use rounded corners to ensure that they do not overlap with the curve below. Figure 2 The upper center is given in the form of a square. Mathematical definition ,in The difference in the time vector between the two rounds of excitation. The Each component, as defined mathematically, directly corresponds to the engineering criterion for determining whether an iteration converges. (Using...) Figure 2 The three iterative convergence curves and the annotation of the analytical saturation beat lower bound shown in this embodiment are used to verify that the transition excitation time vector optimization method can converge to below the preset threshold within 100 iterations under different whole line configurations, and make the converged iterative solution directly approximate the analytical saturation beat lower bound, avoiding the problems of multiple start-stop tests and product scrapping inherent in the traditional whole line beat tuning method based on experience trial and error.
[0036] Figure 2 Three iterative convergence curves are plotted above, denoted as baseline condition A (full-line configuration), buffer capacity increase condition B (full-line configuration), and lower bound of the clock cycle decrease condition C (full-line configuration). All three curves approach from the upper left... to Starting nearby, that is, the range generated in the first iteration after the initial excitation vector starts from the zero vector is on the order of seconds, and then... The curves of configuration A exhibit a typical exponential decay combined with slow oscillations. (The curves decrease monotonically as they increase, tending towards their respective steady states.) to The interval has approached the convergence threshold; the curve of the whole-line configuration B has the steepest descent slope. to The interval has crossed the threshold, reflecting the overall line beat matrix after the buffer capacity is increased. The corresponding directed loop has a smaller time delay and looser clock constraints, thus the iterative exponentiation operation converges faster; the curve descent slope of the whole-line configuration C is the gentlest, until... It only approaches the threshold later because a decrease in the lower bound of the clock cycle means that the minimum dwell times on the bottleneck loops are closer together, thus... Maximum algebraic eigenvalues With a small difference from the second largest eigenvalue, the Max-plus iteration process exhibits a long "translational synchronization" prelude during the convergence phase.
[0037] change The rule for determining the next excitation time can be written as follows: ,Right now .in For change The collection consisting of all input libraries; For input library The previous transition was the token injection. That change; For the first change Secondary trigger moment; for The minimum dwell time. The intuitive meaning of this equation is: the next firing of this transition can only be carried out after all tokens in the input library have been dwelled for the minimum dwell time, and the latest one to arrive ready determines the actual firing time of this transition.
[0038] In the fully colored Petri net model, the excitation times of all transitions satisfying the permitted excitation conditions are arranged in a uniform column vector, denoted as . ,in The total number of transitions in the fully colored Petri net model. This represents the vector transpose. Using matrix operations in Max-plus algebra, the above transition value rules can be combined into... .in for A Max-plus matrix of order 1 is called the full-line beat matrix; The Line number Column elements Defined as: if the transition It is a change in a certain input library The previous transition, then If changes Not a change The previous transition of any input library, then The Max-plus rule for multiplying a matrix and a column vector is as follows: This approach replaces addition in the typical matrix-vector multiplication with taking the maximum value and multiplication with superposition. This replacement is feasible because the causal sequence between transitions in the entire beat matrix precisely satisfies the non-negative equivalent algebraic structure of "taking the maximum value and then superimposing," thus ensuring that the matrix iteration process is strictly monotonically convergent in the Max-plus sense. For transitions that do not meet the permitted activation conditions, in the current iteration round, the transition is marked as prohibited from activation, and the update of the activation time of the transition is paused—specifically, all elements in the entire beat matrix whose row is the transition are temporarily set to zero. It will be restored to its original value only when the permission activation condition is met again.
[0039] Iterative exponentiation operation from the initial excitation time vector Depart, repeatedly follow Update to obtain the sequence of excitation time vectors. A core property of Max-plus algebra guarantees that when the whole-line beat matrix... When the corresponding directed graph has at least one directed cycle and there are no forbidden excitation state transitions on the cycle, a constant exists after a finite number of iterations. Make The components tend to be the same. The constant Equal to the whole line beat matrix The largest ratio among the sum of the cycle delays of all directed cycles in a directed graph and the number of transitions on those cycles is... .in for The set of all directed cycles in a corresponding directed graph; For directed circuits The sum of the minimum dwell times corresponding to each of the above transitions; For directed circuits The number of transitions. The largest of these ratios is the maximum cyclic mean in graph theory, referred to as the maximum cyclic mean in Max-plus algebra. The largest algebraic eigenvalue, and The corresponding directed loop is the bottleneck loop of the entire line. (Constant) As the lower bound of the analytical saturation cycle time for the entire production line under the current configuration, its engineering significance lies in the fact that, regardless of how the control system adjusts the buffer capacity and speed ratio commands, the actual time interval between two adjacent discharge events of the entire production line cannot be less than [a certain value]. This is because the interval is determined by the physical dwell time delay of the bottleneck loop. Utilizing As a lower bound, it provides a hard value for subsequent cycle safety checks, while also preventing the controller from pursuing the unrealistic goal of "infinite speed".
[0040] In engineering implementation, the convergence criterion for iterative exponentiation does not rely on a strict limit definition, but instead uses the following numerical criterion: calculate the difference between the vectors at the two excitation times. Take the difference between the maximum and minimum values among all components of this difference. ,when Less than the preset convergence threshold When determining if the iteration converges, Typical values range from 0.001s to 0.01s. The round number for determining convergence is denoted as... At this time, arithmetic mean of each component As the lower bound of the analytical saturation beat rate of the entire line under the current configuration, and assigned to The difference between the maximum and minimum values is used as the criterion instead of comparing each component individually because the growth of each component during the Max-plus iteration is "translational synchronization" during the convergence phase—that is, each component advances simultaneously at the same pace. It is the most direct indicator for detecting overall synchronicity; using the average value of multiple components instead of arbitrarily selecting one component is to obtain a more robust estimate in the slight oscillations that still exist during the convergence phase. At the numerical implementation level, iterative exponentiation operations can be implemented using dense matrix multiplication, with n typically between 50 and 200. A single iteration takes about tens of microseconds on a general-purpose CPU, and in typical cases, 20 to 80 iterations are sufficient to meet the convergence threshold. The overall computation time is much shorter than the fixed control cycle of 10ms to 20ms, thus ensuring that the rolling derivation of the whole-line colored Petri net model can be carried out in real time.
[0041] Whenever the displacement of the servo adjustment mechanism of any variable buffer unit changes, or the measured value of the minimum dwell time of any processing station changes, the edge server will not indiscriminately re-execute the iterative exponentiation operation, but will first compare... or The edge server will update the data only if the relative change between the old and new values exceeds a preset refresh threshold, typically set at 1% to 5% of the relative change. or Write the whole line beat matrix The corresponding element position, obtained from the convergence of the previous iteration. The initial trigger vector for this iteration of exponentiation is used to restart the iteration until the convergence threshold is met again to obtain a new result. The refresh threshold is introduced to avoid frequent refreshes caused by measurement noise consuming edge server computing resources; using the result of the previous iteration as a new initial excitation time vector is equivalent to a warm start, which can reduce the number of iterations required for reconvergence to about 1 / 3 to 1 / 2 of the original. The newly obtained... The message is broadcast to all local control units at each workstation via a time-sensitive network using a dedicated scheduling time slot. The broadcast message uses a fixed-length data packet of 1 frame, which consumes very little bandwidth and will not compete with the normal workstation status vector reporting message.
[0042] After completing the real-time computation of the full-line colored Petri net model and the analytical saturation beat lower bound, the edge server continues to load the graph attention network as the policy backbone for multi-agent reinforcement learning within the same software process. The graph attention network is implemented using the PyTorch deep learning framework during the training phase and serialized into a static computation graph via TorchScript during the deployment phase. The inference time per iteration is controlled within 2ms, ensuring that even when the control period tightens to 5ms, the output of the graph attention network can still be delivered to the time-sensitive network's sending queue before the control period expires.
[0043] The input graph of the graph attention network is a graph structure obtained by isomorphic transformation of the full-line colored Petri net model, denoted as . .in This is a set of graph nodes, where each node corresponds to either a processing warehouse or a buffer warehouse. The two boundary warehouses (incoming material warehouse and finished product warehouse) are excluded because they do not participate in control decisions. This equals the sum of 6 processing warehouses and 5 buffer warehouses, totaling 11. However, the processing warehouses and buffer warehouses in the longitudinal and transverse ribbed subnets are each mapped to independent graph nodes. Therefore, the actual... The value can be extended to between 13 and 15, depending on whether the longitudinal and transverse ribbed subnets are further subdivided into intermediate states. For the set of directed connections between graph nodes, the generation rule is: In the Petri net model with full line coloring, if the library... Through changes To the warehouse Output token, then in Add 1 line by Corresponding graph node points Directed edges corresponding to the graph nodes are added. Simultaneously, the input libraries on both sides of the merging transition add directed edges to the graph nodes corresponding to the ultrasonic welding library, thus explicitly representing the merging relationship between longitudinally oriented ribs and transversely oriented ribs during the cross-laying process in the graph structure. The consideration for setting reverse directed edges is that the state of the downstream processing station (e.g., whether the ultrasonic welding station is waiting for materials) also affects the speed ratio decision of the upstream processing station. Therefore, in addition to each directed edge from upstream to downstream, an extra directed edge from downstream to upstream is added so that information can be aggregated in both directions simultaneously. However, the semantics of the reverse edge are explicitly labeled as a "feedback edge," and in the attention calculation of the graph attention network, it uses an independent linear transformation from the forward edge to avoid semantic confusion between the two types.
[0044] The input feature for each graph node is a feature of length [length missing]. real vectors, Typical values range from 32 to 64. For the graph node corresponding to the processing library, this vector is formed by concatenating the workstation status vector of the corresponding processing station with the token color code currently being processed at this processing station. The specific content of the workstation status vector has been given above; the token color code adopts a fixed-dimensional embedding vector of the batch identity identifier, that is, the 128-bit batch identity identifier is mapped to a length of [missing value] through an embedding lookup table. real vectors, A typical value is 16. The embedding lookup table is initialized with a random orthogonal matrix during the initial operation of the production line and is updated synchronously with other parameters of the graph attention network during the multi-agent reinforcement learning training process. The embedding vector is used to ensure that the batch identifications corresponding to similar incoming material spectrum recognition results are in close positions in the embedding space, which is beneficial for the graph attention network to give coherent decision outputs for incoming material batches with similar components. For the graph node corresponding to the buffer, the input feature vector is composed of the number of tokens in this buffer, the displacement of the servo adjustment mechanism, and the color histogram encoding of all tokens queued in this buffer. The color histogram encoding is generated as follows: all tokens queued in the buffer are binned according to the principal component category of the incoming material spectrum to which their batch identification belongs (usually 4 to 8 bins), the token count in each bin is divided by the current total number of tokens in this buffer to obtain the normalized frequency, and then concatenated into a vector with a length equal to the number of bins. The reason for using normalized frequency instead of the original count is that the original count will drift significantly with the total number of tokens in the buffer, making it difficult for graph attention networks to learn stable batch composition features; while normalized frequency decouples "batch composition" from "total number", with the former represented by color histogram encoding and the latter by token count scalar representation. This decoupling can significantly improve training stability.
[0045] The graph attention network consists of 3 layers. Three layers were chosen because the graph attenuation network corresponds to the fully line-colored Petri net model. The longest and shortest paths between any two graph nodes typically do not exceed 6 when the number of graph nodes is between 13 and 15. Three layers of graph attention computation are sufficient to ensure that information between any pair of graph nodes travels back and forth more than once. However, excessively deep layers can easily lead to an oversmoothing phenomenon where the features of each graph node tend to become uniform. Each layer of the graph attention network performs attention on each graph node... Perform the following calculations: First, set the graph nodes... Input feature vector The input feature vector of all its first-order neighboring graph nodes Each undergoes a linear transformation shared by this layer. , , Mapped to query vector Key vector AND value vector .in For layer number, For graph nodes The set of first-order adjacent graph nodes, , , All Real matrix of shape For the first The output feature dimension of the layer. Secondly, graph nodes. The aggregation coefficient is obtained by normalizing the inner product of the query vector and the key vector of each first-order adjacent graph node using softmax. .in express Transpose of; in the denominator It is a scaling factor. The reason for introducing this scaling is that when Large inner product The variance also increases accordingly. Directly feeding it into the softmax function will cause the output distribution to tend to extremes, and the gradient will approximately disappear during backpropagation. This precisely restores the variance of the inner product to... The magnitude is adjusted to keep the gradient stable. Next, the value vectors of first-order neighboring graph nodes are weighted and summed using the aggregation coefficients to obtain the neighbor aggregation feature. Finally, the adjacent aggregated features are linked to the graph nodes. The eigenvector obtained by self-connection linear transformation The output features of this layer are obtained by adding elements one by one and then processing them with the ReLU activation function. .in This is the self-connected linear transformation matrix of this layer, and... , , The ReLU activation function does not share parameters; its specific form is as follows: This means taking the maximum value between each element and 0 in the input vector. The consideration behind introducing a self-connected linear transformation is that if there are only adjacent aggregated features without a direct path to the features of the node itself, the initial features of the node will be diluted by the features of adjacent nodes after passing through multiple layers, losing its own workstation identification. Fusing self-connected features and adjacent aggregated features through summation rather than concatenation maintains a constant feature dimension, facilitating multi-layer stacking. In engineering practice, the typical output feature dimension values for the first two layers in a three-layer structure are... The third layer outputs feature dimensions. The progressive narrowing of the graph attention network results in a more compact graph node representation output at the final layer.
[0046] The servo speed ratio command for each machining station is physically a continuous scalar, typically ranging from 0.5 to 2.0, representing the multiple of the servo speed ratio of this station relative to the nominal speed ratio of the production line. To match the discrete motion space of the QMIX algorithm, this continuous interval is uniformly discretized into 7 preset speed ratio levels, i.e. The displacement command of the servo adjustment mechanism of each variable buffer unit is also a continuous scalar. The displacement range of the servo adjustment valve of the melt voltage regulating buffer unit from 0mm to 25mm is uniformly discretized into 5 preset displacement levels, i.e. The displacement range of the lifting floating idler roller in the tension floating material storage buffer unit, from 0.2m to 0.6m, is uniformly distributed into 5 preset displacement levels, i.e. In practical engineering, if there is a greater sensitivity to beat balance within certain intervals, these intervals can be subdivided. A typical approach is to refine the speed ratio increments around 1.00 from 0.25 to 0.10 to achieve finer speed ratio adjustment. The output features of the graph attention network at the last layer are mapped through a fully connected layer to obtain the motion scores of each machining station at the preset speed ratio and the motion scores of each variable buffer unit at the preset displacement; the specific form is as follows: ,in For graph nodes The motion score vector of the corresponding workstation or variable buffer unit at all preset positions, with a length equal to the number of preset positions for that workstation or variable buffer unit; and Graph nodes The corresponding output layer weight matrix and bias vector; To determine the number of layers in the graph attention network, in this embodiment... Each workstation or variable buffer unit maintains its own independent output layer because the preset gear semantics of different workstations and variable buffer units are not consistent. For example, the speed ratio gear of the orientation stretching workstation reflects the adjustment of the stretching orientation ratio, while the speed ratio gear of the node ultrasonic welding workstation reflects the adjustment of the welding cycle. Finally, the gear corresponding to the maximum motion score is taken as the current execution action of this workstation or variable buffer unit. ,in for The Each component.
[0047] Before the execution action is sent to the local control unit at the workstation, the edge server adds a transitional step on the servo side to prevent sudden changes in gear level: if the difference between the gear level corresponding to the execution action in the current cycle and the execution action in the previous cycle exceeds two adjacent gear levels, the execution action in the current cycle is forcibly truncated to within two adjacent gear levels of the previous cycle. This transitional step controls the biased exploration of the multi-agent reinforcement learning algorithm towards long-term rewards within two adjacent gear levels, preserving adjustment bandwidth while avoiding mechanical jitter and resonance of the servo actuator caused by sudden changes in instructions.
[0048] The QMIX algorithm treats each processing station and each variable buffer unit as an agent, resulting in 11 agents in total (6 processing stations and 5 variable buffer units). The value of each agent's local action in the current control cycle is... The output features of the corresponding graph nodes of this agent in the last layer are then passed through an independent fully connected layer, i.e. ,in For the local observation-action history of this intelligent agent, This refers to the action taken by the agent in the current control cycle. and These are the linear head weight vector and bias scalar of the local action values of the agent, respectively. The local action values of all 11 agents are synthesized into the global action value through a monotonically mixed network containing only positive coefficients. .in For joint observation of action history by all intelligent agents, For the joint execution of actions by all intelligent agents, The global state vector for the entire line, in this embodiment, is formed by concatenating the input feature vectors of all graph nodes in order of their graph node numbers, and its length is equal to... ; This is a functional representation of a monotonic hybrid network, typically implemented as a multilayer perceptron with two layers and 32 hidden layers. All elements of its weight matrix are evaluated using absolute value operations before being fed into the forward computation to ensure... For all This monotonicity guarantee holds true. This monotonicity is given by the core theorem of the QMIX algorithm: when each agent selects, under its local observations, ... When executing the action that takes the maximum value, the combined execution action just makes The maximum value is taken, allowing the strategy obtained through centralized training to be deployed in a distributed manner. Forcing positive weights involves generating the weight matrix of the hybrid network using a single supernetwork, which in turn... Let the input be weights and the output be weights activated by an absolute value function—let the output of the first fully connected layer of a supernetwork be denoted as . The weights of the first layer of the hybrid network are then set to... ,in It represents the absolute value of each element.
[0049] Training samples are generated by forward extrapolation on the whole-line colored Petri net model: each time a batch of training samples is generated, the edge server uses the current real-time state of the whole line as the starting point for extrapolation, and completely copies the current token count of each place, the batch identity of each token, and the most recent excitation time of each transition in the whole-line colored Petri net model to the extrapolation copy. Then, the execution action currently output by the graph attention network drives the extrapolation copy to extrapolate forward according to the next excitation time value rule of the transition, the upstream and downstream permitted excitation conditions, and the grid spacing conditions of the merging transition for a set number of control cycles. In this embodiment, the typical extrapolation length is 50 to 200 control cycles, which is equivalent to 0.5s to 4s of material advance time. During the extrapolation process, a clock safety check is performed synchronously within each control cycle: if the interval between two adjacent excitation times of any transition is less than the value corresponding to the lower bound of the analytical saturation clock, a preset clock violation penalty value is added to the single-step reward of this extrapolation sample. The preset penalty value for violations of the beat is a negative number, and the absolute value is 5 to 10 times the upper limit of the effective finished product length calculated based on the rated maximum capacity of the entire line equipment within a single simulation cycle, so as to ensure that the violation samples are always in the worst position in the QMIX training target.
[0050] The transition triggering time intervals collected by the edge server during the continuous operation phase of the entire line in this embodiment are as follows: Figure 3 As shown. Figure 3 The horizontal axis represents the control cycle number, ranging from 0 to 200, and uses a linear scale to characterize the sequence of transition excitation time intervals collected by the edge server in each control cycle during the continuous operation of the entire line. Figure 3 The vertical axis represents the transition excitation time interval, in seconds, ranging from 0.7 seconds to 1.7 seconds, using a linear scale. This range is defined by the lower bound of the analytical saturation beat. Centered on the clock, it covers the allowable redundancy range of clock cycles in normal production upwards and the possible irregularity range of clock cycles downwards. Figure 3 The measured values of the excitation time intervals for each control cycle are shown in scatter plot form, and the temporal trend is displayed by a broken line connecting the scatter plots; a horizontal dashed line running through the entire graph is placed on the vertical axis. The position indicates the reference position for the lower bound of the analytical saturation beat, serving as a benchmark for determining whether a beat has entered a saturation or violation state. (Using...) Figure 3 The timing sequence shown in the continuous 200 control cycles allows this embodiment to directly visualize and compare the actual fluctuation of the entire line's operating rhythm with the lower bound of the analytical saturation rhythm. This enables the edge server to identify the rhythm offset trend and trigger corresponding feedback control at the control cycle level, avoiding the problem of delayed rhythm anomaly detection caused by data backhaul delay in traditional centralized control schemes.
[0051] Figure 3Three curves and a set of markers are plotted above. The first curve corresponds to the traditional PID control strategy, exhibiting obvious characteristics of low-frequency sinusoidal disturbance superimposed with random noise. The overall curve is... The above fluctuation is approximately 0.4s, with a peak-to-valley difference of about 0.4s. This fluctuation reflects that traditional PID control, when faced with the drift in the properties of recycled plastic materials and the constraint of buffer capacity, can only rely on local feedback at each workstation for individual adjustment. Cross-workstation coupling cannot be resolved, and therefore the cycle interval inevitably fluctuates with the material's movement. The second curve corresponds to the graph attention network and QMIX algorithm control strategy of this invention, exhibiting a very small disturbance, with the curve closely following the material's movement. The curve moves along a horizontal line with a disturbance peak-to-valley difference of only about 0.05s. This curve is significantly closer to the first curve. Furthermore, the fact that it doesn't penetrate downwards directly reflects that the multi-agent reinforcement learning employed in this invention can compress the entire line's beat interval to a position close to the lower bound of the analytical saturation beat while ensuring beat safety. The third curve, the horizontal dashed line running through the entire graph, corresponds to the lower bound of the analytical saturation beat. s, the horizontal dashed line is in Figure 3 The middle part serves as both a visual reference for cycle safety checks and an objective baseline for comparing the performance of the two control strategies.
[0052] Figure 3 The traditional PID control curve is marked with a cross-shaped point above. The position of the horizontal dashed line shows that approximately 12 downward crossing events occurred within the entire 200 control cycles. Each downward crossing corresponds to one cycle violation event, meaning that the actual cycle interval is less than the analytical saturation cycle lower bound. This implies that the entire line attempted to output material at a rate physically impossible for the bottleneck loop during that control cycle, directly resulting in material shortages in the buffer or insufficient processing in the processing warehouse. The graph attention network and QMIX algorithm control curve of this invention did not experience any downward crossings throughout the entire cycle. The horizontal dashed line reflects that the beat violation penalty value, through the QMIX training objective function, has pushed the policy toward a stable operating point that "approaches but does not break through".
[0053] The specific definition of single-step reward is: the effective finished product length at the tension winding station at the end of the simulation cycle when the downstream of the entire line simultaneously passes the node welding qualification signal, rib contour qualification signal, and tension stability qualification signal. The three types of qualification signals are screened using "necessary conditions" rather than a weighted combination; that is, if any one of the qualification signals is unqualified, the length of that finished product segment is not included in the effective finished product length. Necessary condition screening avoids the problem of the strategy falling into "trading quantity for defective products" degenerates due to the lack of objective basis for weight selection in the weighted summation method, ensuring that the optimization direction of the graph attention network is strictly aligned to maximize output while ensuring quality.
[0054] The training objective of the QMIX algorithm is to minimize the temporal difference error between the global action value and the single-step reward, specifically in the form of... ,in .in The single-step reward is the final reward after being adjusted for a preset beat violation penalty value; For the joint observation-action history of the next control cycle; Candidate joint execution actions for the next control cycle; This is the discount factor, typically 0.95. To achieve the overall value of the target line's actions, it consists of 1 set and A target network with the same structure but delayed parameter updates is provided. The target network parameters are hard-copied from the main network every 200 training steps to mitigate the non-steady state of the temporal difference objective. The optimizer is Adam, with an initial learning rate of 0.0005, which decays to 0.5 every 50,000 training steps until the temporal difference error stabilizes. The criterion for "stability" is that the relative change of the moving average of the temporal difference error over 2000 consecutive training steps is less than 2%. During training, the simulation copy continuously generates new samples from the latest real-time state of the entire production line, thus ensuring that the training samples are homologous to the latest operating conditions of the production line and eliminating the need to rely on historical data replay.
[0055] In this embodiment, the effective finished product length sequence of the tension winding station after continuous operation of the method for one standard production shift is as follows: Figure 4 As shown. Figure 4 The horizontal axis represents production time, in minutes, ranging from 0 to 480, corresponding to an 8-hour continuous operation of a standard production shift; the vertical axis represents the effective finished product length at the tension winding station, in meters per minute, ranging from 40 to 160, using a linear scale. The effective finished product length is the portion of the material wound at the tension winding station that simultaneously passes three screening tests: node welding qualification signal, rib profile qualification signal, and tension stability qualification signal. Figure 4 The curve illustrates the evolution of the effective finished product length over production time. The curve rapidly climbs from a lower level to the rated range within the first few minutes after production starts, then remains stable within the rated range, maintaining a relatively stable continuous output in the latter part of the shift. (Using...) Figure 4 The 8-hour continuous operation sequence shown in this embodiment is used to verify that the method can maintain a stable output of effective finished product length throughout the entire standard production shift, thereby supporting the effective production capacity in industrial continuous production scenarios and matching the cycle time of downstream handling, inspection, and warehousing.
[0056] Figure 4The diagram shows two main curves and one vertical dashed line. The first curve corresponds to the traditional PID control strategy, exhibiting low-frequency sinusoidal fluctuations with an average speed of approximately 100 meters per minute and a peak-to-valley difference of approximately 24 meters per minute, superimposed with random noise. The second curve corresponds to the graph attention network and QMIX algorithm control strategy of this invention. It coincides with the first curve during the production time interval from 0 to 60 minutes. At the 60-minute mark, the vertical dashed line marks the control strategy switching point. From this point onward, the second curve exhibits an approximately linear upward transition phase, lasting approximately 60 minutes before entering a steady state. The steady-state average speed is approximately 138 meters per minute, and the steady-state peak-to-valley difference is approximately 12 meters per minute. The transition phase reflects that the graph attention network requires several control cycles to adapt to the current incoming batch attributes during the initial switching phase. The steady-state phase reflects the full release of the overall line's cycle time balancing capability under the constraint of cycle time safety checks.
[0057] This invention improves the effective production capacity by approximately 31.5% compared to traditional PID control in the steady-state phase. This figure is obtained by dividing the difference between the average effective finished product length in the steady-state phase and the average effective finished product length in the baseline phase of traditional control by the average value of the baseline phase. This improvement comes from cycle time equalization, which brings the nominal production capacity of the entire line closer to the physical upper limit corresponding to the bottleneck loop, and the necessary condition screening single-step reward-guided multi-agent strategy, which maximizes the winding length while ensuring that all three types of qualified signals are true simultaneously.
[0058] The present invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the invention. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of the present invention. For those skilled in the art, various improvements and modifications can be made to the present invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. An industrial production method for high-friction geogrid made from recycled plastics, characterized in that, The method includes the following steps: Step 1: Deploy 6 processing stations sequentially along the material's forward direction, with variable buffer units between adjacent processing stations; each processing station is equipped with a local control unit, which collects the station's status vector and reports it to the edge server via a time-sensitive network; each batch of recycled plastic granules is assigned a batch identifier. Step 2: Establish a full-line colored Petri net model in the edge server, using processing stations as processing warehouses, variable buffer units as buffer warehouses, material transfer actions between processing warehouses and buffer warehouses as transitions, and batch identification as token colors; perform iterative exponentiation operations on the excitation times of transitions in the full-line colored Petri net model based on Max-plus algebra to obtain the analytical saturation beat lower bound of the entire line under the current configuration; Step 3: Construct a graph attention network in the edge server as the policy backbone for multi-agent reinforcement learning. Use the topology of the whole line colored Petri net model as the input graph and the lower bound of the analytic saturation beat as the beat safety check constraint. Use the QMIX algorithm to train the graph attention network. After training, the edge server runs the graph attention network according to the real-time status of the whole line in each control cycle. The output execution actions are sent to the local control unit of each workstation for execution via the time-sensitive network.
2. The industrial production method of recycled plastic high-friction geogrid according to claim 1, characterized in that... The six processing stations, arranged sequentially along the material's direction of travel, are: melt compounding station, die extrusion station, shaping and cooling station, orientation stretching station, ultrasonic welding station for joints, and tension winding station. In the melt compounding station, recycled plastic granules are plasticized into a melt using a twin-screw extruder. The die extrusion station employs a dual-die structure, with two dies extruding longitudinal and transverse ribs respectively. Both the longitudinal and transverse ribs have alternating raised friction sections and transition connecting sections along their length. The convex height and / or cross-sectional area of the rib surface of the friction section is greater than that of the transition connection section. The convex friction section is used to form a mechanical interlocking interface in the soil. The convex friction section and the transition connection section together constitute a continuous billet with non-uniform ribs. The shaping and cooling station performs water bath cooling and air cooling shaping on the longitudinal ribs and transverse ribs respectively. The orientation stretching station performs stretching and orientation treatment on the longitudinal ribs and transverse ribs respectively to obtain longitudinally oriented ribs and transversely oriented ribs. The tension winding station performs constant tension winding on the finished product.
3. The industrial production method of recycled plastic high-friction geogrid according to claim 2, characterized in that... The ultrasonic welding station for nodes includes a cross-laying substation and an ultrasonic welding substation. The cross-laying substation includes a transverse rib fixed-length cutting module, a transverse transfer module, and a grid positioning module. The transverse rib fixed-length cutting module cuts the transverse oriented ribs into transverse ribs of a preset length. The transverse transfer module transfers the transverse ribs to the top of the longitudinal oriented ribs. The grid positioning module cross-overlays the transverse ribs with multiple longitudinal oriented ribs according to a preset grid spacing. The ultrasonic welding substation performs ultrasonic welding on the cross-overlay positions to form a welded node.
4. The industrial production method of recycled plastic high-friction geogrid according to claim 2, characterized in that... Five variable buffer units are set between two adjacent processing stations. The variable buffer unit between the melt mixing station and the die extrusion station is a melt pressure stabilizing buffer unit, and the capacity parameter of the melt pressure stabilizing buffer unit is controlled by the displacement of the servo regulating valve of the gear pump flow stabilizing chamber. The remaining four variable buffer units are tension floating material storage buffer units, and the capacity parameter of the tension floating material storage buffer units is controlled by the servo displacement of the liftable floating idler. The capacity parameters of the melt pressure stabilizing buffer unit and the tension floating material storage buffer unit are collectively referred to as the displacement of the servo regulating mechanism.
5. The industrial production method of recycled plastic high-friction geogrid according to claim 4, characterized in that... Except for the tension winding station, the local control units of the other 5 processing stations periodically collect the servo torque, melt pressure or strip tension, process temperature, output length per unit time calculated by the output encoder pulse, and the displacement of the servo adjustment mechanism of the downstream variable buffer unit of the station to form a station status vector; the local control unit of the tension winding station periodically collects the winding tension, winding diameter and winding length count to form a station status vector and outputs a tension stability qualified signal; all station status vectors are reported to the edge server through a time-sensitive network at a fixed control cycle.
6. The industrial production method of recycled plastic high-friction geogrid according to claim 2, characterized in that... An online contour detection unit and a node welding qualification detection unit are set up before the tension winding station. The online contour detection unit collects the rib height, rib spacing and rib protrusion cross-sectional dimensions, and outputs a rib contour qualification signal. The node welding qualification detection unit collects the infrared thermograph of the welding node, ultrasonic welding energy, welding head downward displacement and welding pressure, and outputs a node welding qualification signal in combination with the node pull-out force calibration results obtained from periodic sampling inspections.
7. The industrial production method of recycled plastic high-friction geogrid according to claim 2, characterized in that... Each batch of recycled plastic pellets entering the melt mixing station is assigned a batch identity identifier. The batch identity identifier is determined by the incoming material spectrum identification result and the feeding time of this batch. The position of the batch identity identifier in the entire line is estimated in real time by the edge server based on the feeding time of this batch, the screw speed of the twin-screw extruder, the melt pressure, the melt flow rate, and the discharge length per unit time of each processing station, so as to obtain the current location of the batch identity identifier and the number of tokens it occupies.
8. The industrial production method of recycled plastic high-friction geogrid according to claim 3, characterized in that... In the full-line colored Petri net model, one token represents one unit length of material. Max-plus algebra uses the maximum value operation instead of conventional addition and superposition operation instead of conventional multiplication. The full-line colored Petri net model establishes longitudinal ribbed subnets and transverse ribbed subnets between the die extrusion station and the node ultrasonic welding station. Tokens in the longitudinal ribbed subnet represent one unit length of longitudinal material, and tokens in the transverse ribbed subnet represent one unit length of transverse material. The cross-laying sub-stations of the node ultrasonic welding station correspond to one merging transition. The merging transition is determined by the tokens in the buffer pool at the end of the longitudinal ribbed subnet and the tokens in the buffer pool at the end of the transverse ribbed subnet. As a common input, the merging transition is only permitted to be activated when the number of tokens in both end buffer depots meets the token quantity condition corresponding to the preset grid spacing. After the merging transition is activated, one grid cell token is generated in the processing depot corresponding to the ultrasonic welding station at the node. A minimum dwell time delay is defined for each processing depot and each buffer depot in the whole-line colored Petri net model. The minimum dwell time delay of the processing depot is equal to the measured time required for the corresponding processing station to complete the processing of one unit length of material under the calibrated working conditions. The minimum dwell time delay of the buffer depot is equal to the measured time required for the material to pass through the corresponding variable buffer unit at the maximum conveying rate under the current servo adjustment mechanism displacement. Time; For each buffer, a maximum token capacity and a minimum retention capacity are defined simultaneously. The maximum token capacity and minimum retention capacity are updated in conjunction with the displacement of the servo adjustment mechanism of the corresponding variable buffer unit. Upstream transitions of a buffer are only permitted to be triggered when the number of tokens in the buffer is lower than the maximum token capacity, and downstream transitions of a buffer are only permitted to be triggered when the number of tokens in the buffer is higher than the minimum retention capacity. The rule for determining the next trigger time of a transition is as follows: For each input buffer of this transition, the trigger time of the previous transition corresponding to this input buffer is superimposed with the minimum dwell time of this input buffer to obtain the superposition sum. This superposition sum is applied to all input buffers of this transition. The maximum sum is taken as the next excitation time of this transition; for transitions that do not meet the permitted excitation conditions, the transition is marked as prohibited from excitation in the current iteration round and the update of the excitation time of this transition is paused; the iterative exponentiation operation starts from an initial excitation time vector of zero vector and repeatedly updates the excitation times of all transitions in the entire line that meet the permitted excitation conditions according to the rule of taking the next excitation time value of the transition, so as to obtain the excitation time vector sequence; when the difference between the maximum and minimum values of the component differences of the excitation time vectors of two adjacent rounds is less than the preset convergence threshold, it is determined that the iteration has converged, and the average value of the component differences of the excitation time vectors of two adjacent rounds is taken as the lower bound of the analytical saturation beat of the entire line under the current configuration;Whenever the displacement of the servo adjustment mechanism of any variable buffer unit changes or the measured minimum dwell time of any processing station changes, the edge server re-executes an iterative exponentiation operation and refreshes the lower bound of the analytical saturation beat rate. The refreshed lower bound of the analytical saturation beat rate is broadcast to all local control units of all workstations via a time-sensitive network.