Reinforcement learning-based multi-parameter collaborative optimization control method for impregnated paper production process
By adopting a multi-parameter collaborative optimization control method based on reinforcement learning, the problem of parameter optimization relying on experience in the impregnated paper production process was solved. This method achieves full-process data transparency and real-time optimization, improves the stability and efficiency of the production process, and meets the dynamic adjustment of market demands.
Patent Information
- Application Number
- CN202511580744.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-10-31
AI Technical Summary
In the traditional impregnated paper production process, the optimization of process parameters relies heavily on experience-based adjustments, lacks scientific quantitative decision-making basis, makes multi-parameter coordinated control difficult, and has insufficient real-time monitoring and dynamic adjustment capabilities for the production process. It is difficult to adapt to fluctuations in raw materials and changes in market demand, and there is a lack of effective means to balance and optimize energy consumption control and product quality, resulting in unstable product quality and low production efficiency.
A multi-parameter collaborative optimization control method based on reinforcement learning is adopted to achieve panoramic perception and real-time optimization of the production process by establishing a comprehensive data acquisition network, data fusion technology, knowledge graph, and digital twin model. Specific steps include: deploying a sensor network, constructing a data acquisition network, performing data preprocessing and fusion, constructing a knowledge graph, establishing a digital twin model, designing a multi-objective optimization decision framework, and using reinforcement learning algorithms for adaptive control.
It achieves full-process data transparency and standardization in the impregnated paper production process, improves data quality and information integrity, provides a high-fidelity simulation verification environment, and can simultaneously handle multiple constraints such as product quality, production efficiency, energy consumption costs and equipment safety, thus solving the problem of insufficient adaptability of traditional methods in complex industrial environments.
Smart Images

Figure CN121028731B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent manufacturing, and in particular to a multi-parameter collaborative optimization control method for impregnated paper production process based on reinforcement learning. BACKGROUND
[0002] As an important component of surface decoration materials for man-made boards, the production process of impregnated paper involves multiple complex process links such as raw paper pretreatment, impregnating liquid preparation, impregnating drying, hot pressing forming, etc. In the traditional production process of impregnated paper, there is a strong coupling and nonlinear relationship between multiple key parameters such as temperature, pressure, concentration, and speed. Small changes in process parameters may cause significant fluctuations in product quality. The main technical challenges currently faced by the industry include: process parameter optimization relies heavily on experience adjustment, lacking scientific quantitative decision-making basis; multi-parameter collaborative control is difficult, lacking effective coordination mechanism between process links; real-time monitoring and dynamic adjustment capability of production process is insufficient, difficult to adapt to raw material fluctuations and market demand changes; there is a lack of effective means for balancing optimization between energy consumption control and product quality. These technical bottlenecks seriously restrict the stability of impregnated paper product quality and the improvement of production efficiency, and it is urgent to introduce advanced intelligent control technology to break through the limitations of traditional production mode. SUMMARY
[0003] In order to solve the above problems, the purpose of the present application is to provide a multi-parameter collaborative optimization control method for impregnated paper production process based on reinforcement learning, which solves the problem of multi-parameter collaborative optimization control in impregnated paper production process.
[0004] To achieve the above purpose, the present application adopts the following technical scheme:
[0005] The multi-parameter collaborative optimization control method for impregnated paper production process based on reinforcement learning comprises the following steps:
[0006] S1: Establish a comprehensive data acquisition network to realize panoramic perception of the production process and obtain real-time data stream;
[0007] S2: Based on data fusion technology, intelligently integrate data of different sources and different types to obtain a fused data set;
[0008] S3: According to the fused data set and based on the process mechanism, a knowledge graph of the impregnated paper production process is constructed;
[0009] S4: Combine the fused data set and the knowledge graph to construct a digital twin model of the impregnated paper production line;
[0010] S5: According to the digital twin model, production plan and constraint conditions, design a multi-objective optimization decision framework to realize collaborative optimization at different time scales.
[0011] Further, a comprehensive data collection network is established to realize panoramic perception of the production process and obtain real-time data flow, as follows:
[0012] A group of sensors is deployed at key process nodes of impregnated paper production, including concentration sensors, temperature sensors, liquid level sensors, and pH value sensors installed in the impregnation tank area; thickness detectors, speed sensors, and tension sensors configured in the coating area; temperature field sensor arrays, humidity sensors, and hot air flow meters arranged in the drying and curing area; and roll diameter measurement and visual sensors set in the winding area.
[0013] For key equipment, including pumps, fans, heaters, and pressure rollers, vibration sensors, current sensors, and bearing temperature sensors are installed to realize real-time monitoring of the health status of the equipment and form an image of the equipment operation status through the collection of equipment operation parameters.
[0014] Near-infrared spectrometers are also deployed to detect paper gum content, and laser scanning is used to detect surface flatness. Temperature and humidity sensors and air pressure sensors are used to monitor the temperature, humidity, and air pressure environment parameters in the workshop.
[0015] A comprehensive data collection network is established to obtain the sensor data mentioned above and collect real-time data flow.
[0016] Further, the comprehensive data collection network adopts a hierarchical data collection network, including a field device layer, a data aggregation layer, and a data center layer, as follows:
[0017] The field device layer connects all sensors to field I / O modules through field buses according to the principle of proximity, and the field bus uses the Profinet protocol.
[0018] The data aggregation layer sets up data aggregation servers in each section, configures industrial computers, runs real-time database software, and connects to the central data center through gigabit Ethernet.
[0019] The data center layer establishes a central database server cluster, uses a time series database design to optimize the storage and query efficiency of time series data, and is equipped with a data backup and recovery system to ensure data security.
[0020] Further, based on data fusion technology, different sources and types of data are intelligently integrated to obtain a fused data set, as follows:
[0021] A unified data preprocessing pipeline is established to perform standardized processing on data from different sources. First, data cleaning is performed to identify and handle missing values, outliers, and duplicates. Then, data format unification is performed to convert all timestamps to the standard UTC format, retain significant digits for numerical data and unify units, establish a standard encoding mapping table for text data, and unify resolution and format for image data. Next, time alignment processing is performed to establish a time offset compensation mechanism and map all data to a unified time reference. For data with different sampling frequencies, resampling techniques are used to ensure the consistency of time series. Finally, data standardization is performed using Z-score normalization for numerical data and one-hot encoding or label encoding for categorical data to ensure the comparability of data with different dimensions and ranges. A standardized multi-source data set is obtained after preprocessing.
[0022] Based on the standardized multi-source data set, a hierarchical data fusion strategy is developed in combination with the production process characteristics and business requirements of impregnated paper to obtain a fused data set.
[0023] Further, the hierarchical data fusion strategy includes data-level fusion and feature-level fusion, each level using different fusion algorithms and weight distribution mechanisms, as follows: the first layer of data-level fusion uses a weighted average method to fuse the measurements of nine thermocouples for impregnation tank temperature; for equipment vibration data, an adaptive Kalman filter is used to dynamically adjust filter parameters based on vibration frequency spectrum characteristics to suppress noise interference and improve signal quality; the second layer of feature-level fusion constructs a comprehensive feature vector by combining measurements of different physical quantities into a comprehensive index reflecting system state, uses principal component analysis dimension reduction technology to project high-dimensional original features into a low-dimensional space, uses independent component analysis to separate mixed signals and extract independent source signal components; the nonlinear relationship between features is identified through mutual information and correlation analysis to construct a feature correlation atlas; finally, after two layers of fusion processing, the original multi-source data is converted into a fused data set.
[0024] Further, based on the fused data set and the process mechanism, a knowledge graph of the impregnated paper production process is constructed as follows:
[0025] The semantic structure of the field concept is defined, the classification relationship, composition relationship, and dependency relationship between concepts are established, and a knowledge framework for the impregnated paper production field is formed.
[0026] Based on mechanism knowledge modeling and equation integration, the penetration of resin into paper fibers during the impregnation process is a key physical process that follows an extended form of Darcy's law, with penetration velocity related to pressure gradient, liquid viscosity, and porous medium permeability:
[0027] ;
[0028] where: v is the penetration velocity, k is the effective permeability, μ is the dynamic viscosity, P is the pressure gradient, ρ is the liquid density, g is the gravitational acceleration, σ is the surface tension, θ is the contact angle, r is the pore radius;
[0029] In the knowledge graph, the mechanism equation is represented by the MechanisticEquation entity, and each physical quantity in the equation is mapped to the corresponding Parameter entity, connected through the hasVariable relationship; the calibration results of the equation parameters are stored in the EquationParameter entity, including parameter values, calibration accuracy, and applicable range information;
[0030] The drying kinetics equation is:
[0031] ;
[0032] where X is the moisture content on a wet basis, t is the time, h m is the mass transfer coefficient; A is the heat transfer area, m s is the dry basis mass, Y s is the surface vapor concentration, Y ∞ is the ambient vapor concentration, k(T, RH) is the drying constant, is the equilibrium moisture content;
[0033] The effect of temperature on the drying constant follows the Arrhenius relationship:
[0034] ;
[0035] where: k0 is the pre-exponential factor, E a is the activation energy, R is the gas constant, and T is the absolute temperature;
[0036] The curing process of thermosetting resin follows the law of chemical reaction kinetics, and the change of the curing degree with time can be described by the Kamal model:
[0037] ;
[0038] where α is the curing degree, k1 and k2 are the reaction rate constants, and m and n are the reaction orders;
[0039] A mechanism knowledge network is formed in the knowledge graph, connected to the related process through the hasEquation relationship, connected to the affected process parameters through the affectsParameter relationship, and connected to the physical and chemical principles through the basedOnPrinciple relationship;
[0040] The expert experience is converted into standard fuzzy reasoning rules, and a Mamdani reasoning method is adopted;
[0041] Based on the above-mentioned knowledge framework in the field of impregnated paper production, mechanism knowledge modeling and equation integration and fuzzy reasoning rules, a knowledge graph of the impregnated paper production process is constructed.
[0042] Further, combined with the fused data set and knowledge graph, a digital twin model of the impregnated paper production line is constructed, specifically as follows:
[0043] Based on the fused data set and knowledge graph, the digital twin model of the impregnated paper production line adopts a five-dimensional architecture design, including a physical space, a virtual space, a connection layer, a data layer and a service layer; the physical space includes actual production equipment, a sensor network, a control system and material flow; the virtual space includes digital representation of the physical system, integrating geometric models, physical models, behavior models and knowledge models; the connection layer realizes real-time bidirectional interaction between the physical space and the virtual space, including data acquisition, state synchronization and control instruction issuing functions; the data layer manages structured knowledge from the fused data set and knowledge graph; the service layer provides various intelligent services based on the digital twin model, including state monitoring, fault diagnosis and performance prediction;
[0044] The digital twin model has two operating modes, real-time synchronization and predictive simulation. In real-time synchronization mode, the digital twin model is time-synchronized with the physical system, receives sensor data in real time, updates the state of the digital twin model, and accurately maps the physical system. In predictive simulation mode, based on the current state and boundary conditions, the future evolution trajectory of the system is predicted.
[0045] Further, the virtual space includes digital representation of the physical system, integrating geometric models, physical models, behavior models and knowledge models, specifically:
[0046] The geometric model establishes a three-dimensional digital representation of the production line, including equipment geometric models, plant space models and material flow path models, and obtains accurate geometric information through CAD data import. The physical model integrates mechanism equations in the knowledge graph to establish a mathematical model describing physical phenomena;
[0047] The behavior model describes the dynamic response characteristics of the system, including control system behavior, equipment motion behavior and material flow behavior. The control system behavior adopts a state space model:
[0048] ;
[0049] wherein x is a state vector, u is an input vector, y is an output vector, w and v are process noise and measurement noise, A, B, C and D are system matrices; is the time derivative of the state vector;
[0050] The dynamics of the device is governed by the equations of multibody dynamics:
[0051] ;
[0052] where q is the generalized coordinate, M is the mass matrix, C is the damping matrix, K is the stiffness matrix, Q is the generalized force, Q c is the constraint force; is the generalized acceleration vector; is the generalized velocity vector; t represents the time t;
[0053] The knowledge model adopts a hybrid modeling method, combining mechanism models and data-driven models:
[0054] ;
[0055] where f physics is the physical mechanism model, f ML is the machine learning model, and a is the fusion weight, which is dynamically adjusted by an adaptive algorithm: is the predicted output value of the knowledge model; is the physical model parameter vector; is the machine learning model parameter vector;
[0056] ;
[0057] where e physics and e ML are the historical errors of the physical model and the machine learning model, respectively, s is the activation function, w α is the weight parameter vector; x(t) is the input feature at the current time; e physics (t) is the historical prediction error of the physical model at time t; e ML (t) is the historical prediction error of the machine learning model at time t; b α is the bias parameter, used to adjust the baseline value of the weight.
[0058] Further, in the prediction simulation mode, forward simulation is performed based on the current state and boundary conditions to predict the future evolution trajectory of the system, as follows:
[0059] Based on time series analysis and recurrent neural networks for short-term prediction, a long short-term memory network is used to capture the time series dependence:
[0060] ;
[0061] ;
[0062] ;
[0063] ;
[0064] ;
[0065] ;
[0066] wherein, is the hidden state of time step t; is the forget gate; is the input gate; is the candidate value, i.e., the candidate memory cell at the current time instant; is the cell state, representing the memory cell state at the current time instant; is the output gate; is the hidden state at the current time instant; , , , are the weight matrices of the forget gate, the input gate, the candidate value and the output gate, respectively; , , , are the bias vectors of the forget gate, the input gate, the candidate value and the output gate, respectively; σ is the sigmoid function; ⊙ is the Hadamard product; X t represents the input at time step t;
[0067] The mid-term prediction is performed using the state-space model:
[0068] ;
[0069] wherein, x t is the state vector at time t; x t+1 is the state vector at the next time instant (t+1); u t is the control input vector at time t; y t is the observation output vector at time t; A is the state transition matrix; B is the control input matrix; C is the output matrix; D is the feedforward matrix; w t is the process noise vector; v t is the observation noise vector;
[0070] The parameter matrix is obtained through system identification:
[0071] ;
[0072] wherein, θ is the parameter vector to be estimated; is the optimal estimation value of the parameter; N is the total length of the observation data; y t is the actual observation value at time t; Model prediction value based on parameter θ; Square norm of prediction error; λ is regularization coefficient.
[0073] Further, the multi-objective optimization decision framework realizes the collaborative optimization of different time scales, as follows:
[0074] Production process optimization involves multiple conflicting objectives such as quality, cost, efficiency, and energy consumption. Multi-objective optimization method is used to solve the Pareto optimal solution set:
[0075] minF(x)=[f1(x),f2(x),f3(x),f4(x)] T ;
[0076] subject to: g(x)≤0,h(x)=0;
[0077] Where f1(x) is the quality index, f2(x) is the production cost, f3(x) is the energy consumption index, and f4(x) is the production efficiency. g(x) represents the maximum load of equipment; h(x) represents the production quota;
[0078] Reinforcement learning adaptive control uses deep deterministic policy gradient algorithm to realize adaptive control policy learning, where the Actor network updates:
[0079] ;
[0080] Where, is the parameter vector of the Actor network; is the deterministic policy function; s t is the state at time t; a is the action variable; is the state distribution of the behavior policy; β represents the behavior policy; is the Critic network; s is the state; is the gradient of the Q function to the action; is the parameter of the Critic network; represents the integral of the state under the behavior policy distribution; J is the objective function;
[0081] Critic network update:
[0082] ;
[0083] Where, t is the TD target value; is the experience replay buffer; r t is the obtained immediate reward; s t+1 is the next state;
[0084] Target network soft update:
[0085] ;
[0086] wherein, are target Critic network parameters; are target Actor network parameters; tau is a soft update coefficient;
[0087] Based on the prediction ability of the digital twin model, a nonlinear model predictive control is adopted to realize rolling optimization:
[0088] ;
[0089] ;
[0090] wherein, L(x,u) is a stage cost function, V f (x) is a terminal cost function, N is a prediction horizon length; k is a time step index in the prediction horizon; f(x k ,u k ) is a system state transition function; is a control input constraint set; is a state constraint set; is a terminal constraint set; x k ,u k respectively represent the system state vector and the control input vector at the kth step;
[0091] The cost function is designed in a weighted multi-objective form:
[0092] ;
[0093] wherein, w1, w2, w3 are weight coefficients, Q, R, S are weight matrices, Delta u is a control increment, u ref is a control reference value vector, and x ref is a state reference value vector.
[0094] The present application has the following beneficial effects:
[0095] 1. The present application constructs a comprehensive data acquisition network combined with multi-source heterogeneous data fusion technology, breaks through the limitations of traditional single data source, realizes the data transparency and standardization of the whole process from raw material input to finished product output, significantly improves the data quality and information integrity, secondly, based on the process mechanism, a knowledge graph is constructed, the experience knowledge, theoretical mechanism and real-time data are organically combined to form a structured knowledge representation system, and through the deep integration of digital twin model and reinforcement learning algorithm, the real-time synchronous mapping of physical system and virtual system is realized, which provides a high-fidelity simulation verification environment for multi-parameter collaborative optimization;
[0096] 2. The present application can simultaneously handle multiple constraint targets of product quality, production efficiency, energy consumption cost and equipment safety through the target optimization decision framework, solving the problem of insufficient adaptability of traditional single-target optimization methods in complex industrial environments. BRIEF DESCRIPTION OF DRAWINGS
[0097] Figure 1 The flowchart of the method of the present application. DETAILED DESCRIPTION
[0098] The present application will be further described in detail below in combination with the drawings and specific examples:
[0099] REFERENCE Figure 1 In this embodiment, a multi-parameter collaborative optimization control method for impregnated paper production process based on reinforcement learning is provided, including the following steps:
[0100] S1: Establish a comprehensive data acquisition network to realize panoramic perception of the production process and obtain real-time data streams;
[0101] S2: Based on data fusion technology, intelligently integrate data of different sources and types to obtain a fused data set;
[0102] S3: According to the fused data set and based on the process mechanism, construct a knowledge graph of the impregnated paper production process;
[0103] S4: Combine the fused data set and the knowledge graph to construct a digital twin model of the impregnated paper production line;
[0104] S5: According to the digital twin model, production plan and constraint conditions, design a multi-objective optimization decision framework to realize collaborative optimization at different time scales.
[0105] In this embodiment, a comprehensive data acquisition network is established to realize panoramic perception of the production process and obtain real-time data streams, which are as follows:
[0106] Deploy sensor groups at key process nodes of impregnated paper production, install concentration sensors (conductivity type, refractive type), temperature sensors (thermocouple, thermal resistance), liquid level sensors (ultrasonic, pressure type), pH value sensors in the impregnated tank area; configure thickness detectors (laser, eddy current), speed sensors (encoder, radar), tension sensors (pressure type) in the coating area; arrange temperature field sensor arrays, humidity sensors, hot air flow meters in the drying and curing area; set up roll diameter measurement and visual sensors in the winding area;
[0107] For key equipment, including pumps, fans, heaters, and pressure rollers, vibration sensors, current sensors, and bearing temperature sensors are installed to monitor the health status of the equipment in real time. Through the collection of equipment operating parameters, an image of the equipment operating status is formed.
[0108] A near-infrared spectrometer is deployed to detect the glue content of paper, a laser scanning device is set up to detect surface flatness, and temperature and humidity sensors and air pressure sensors are used to monitor the temperature, humidity, and air pressure environment parameters in the workshop.
[0109] A comprehensive data collection network is constructed to obtain the sensor data and collect real-time data streams.
[0110] In this embodiment, the comprehensive data collection network adopts a hierarchical data collection network, including a field device layer, a data aggregation layer, and a data center layer, as follows:
[0111] The field device layer connects all sensors to the field I / O module through the field bus according to the nearest principle, adopts modular design, and supports multiple signal types (analog input AI, digital input DI, pulse input PI, etc.). The field bus uses the Profinet protocol, with a transmission speed of 100 Mbps, supporting hot plug and redundant configuration.
[0112] The data aggregation layer sets up data aggregation servers in each section, configures industrial-grade computers (Intel i7 processor, 16GB memory, 1TB SSD storage), runs real-time database software, and connects to the central data center through a gigabit Ethernet network.
[0113] The data center layer establishes a central database server cluster, uses a time series database design to optimize the storage and query efficiency of time series data, and is equipped with a data backup and recovery system to ensure data security.
[0114] In this embodiment, based on data fusion technology, different sources and types of data are intelligently integrated to obtain a fused data set, as follows:
[0115] A unified data preprocessing pipeline is established to perform standardized processing on data from different sources. First, data cleaning operations are performed to identify and process missing values (using forward filling, backward filling, linear interpolation, mean filling, etc.), outliers (based on the 3σ criterion, quartile method, and isolation forest algorithm), and duplicate values (based on timestamps and data source identifiers). Then, data format unification is performed to convert all timestamps to the standard UTC format, retain the number of significant digits for numerical data and unify the unit system (temperature unified to Celsius, pressure unified to MPa, speed unified to m / min, etc.), establish a standard encoding mapping table for text data, and unify the resolution and format of image data. Next, time alignment processing is performed to establish a time offset compensation mechanism and map all data to a unified time reference. Resampling techniques are used for data with different sampling frequencies (interpolation for upsampling and filtering for anti-aliasing for downsampling) to ensure the consistency of time series. Finally, data standardization is performed using Z-score normalization (mean of 0 and standard deviation of 1) for numerical data and one-hot encoding or label encoding for categorical data to ensure the comparability of data with different dimensions and ranges. The preprocessed standardized multi-source data set is obtained.
[0116] Based on the standardized multi-source data set, a hierarchical data fusion strategy is developed based on the production process characteristics and business requirements of impregnated paper to obtain a fused data set.
[0117] In this embodiment, the hierarchical data fusion strategy includes data-level fusion (redundant measurement fusion of similar sensors) and feature-level fusion (feature combination of different types of data). Different fusion algorithms and weight distribution mechanisms are used at each level, as follows: the first layer of data-level fusion uses a weighted average method to fuse the measurements of 9 thermocouples for impregnation tank temperature. For device vibration data, an adaptive Kalman filter is used to dynamically adjust the filter parameters based on the vibration frequency spectrum characteristics to suppress noise interference and improve signal quality. The second layer of feature-level fusion constructs a comprehensive feature vector by combining measurements of different physical quantities into a comprehensive index reflecting the system state. Principal component analysis is used to reduce the dimensionality of high-dimensional original features to retain more than 95% of the information. Independent component analysis is used to separate mixed signals and extract independent source signal components. Mutual information and correlation analysis are used to identify nonlinear relationships between features and construct a feature correlation atlas. After two layers of fusion processing, the original multi-source data is transformed into a fused data set.
[0118] In this embodiment, based on the fused data set and the process mechanism, a knowledge graph of the impregnated paper production process is constructed as follows:
[0119] The semantic structure of the field concept is defined, the classification relationship, composition relationship, and dependency relationship between concepts are established, and the knowledge framework of the impregnated paper production field is formed.
[0120] The ProcessStep class is defined as an abstract concept of a process step, which specifically includes the Preparation, Impregnation, Coating, Drying, Curing, and Winding subclasses. Each process step is connected to related process parameters through the hasParameter relationship, connected to the equipment used through the usesEquipment relationship, and connected to the input and output materials through the hasInput and hasOutput relationships. The sequence dependency between process steps is established through the precedes and follows relationships, forming a complete process flow chain. The time ontology is introduced to model the dynamic characteristics of the process, including the duration, start time, and end time attributes.
[0121] Equipment is an important component of the production process, and is classified by function into the Impregnation Tank, Coating Head, Drying Chamber, Pump, and Sensor subclasses. The equipment's composition structure is described through the hasComponent relationship, its spatial location is described through the locatedIn relationship, and its operating state is described through the hasState relationship. The material Material includes the Base Paper, Resin, Solvent, and Additive subclasses, each with specific physical and chemical properties, which are connected through the hasProperty relationship. There are mixing, dissolution, and reaction chemical relationships between materials, which are important components of process mechanism knowledge.
[0122] Process parameters and quality indicators are key numerical concepts in the knowledge graph, and their numerical characteristics and constraint relationships are modeled. Parameters include Temperature, Pressure, Concentration, Viscosity, Speed, and Tension subcategories. Quality indicators include Resin Content, Thickness, Uniformity Index, and Surface Quality subcategories. The relationship between parameters and indicators is established through the affects relationship to establish the impact relationship and the correlatedWith relationship to establish the correlation relationship. The constraint concept Constraint is introduced to describe the value constraints of parameters, including minimum value constraints (minValue), maximum value constraints (maxValue), and target value constraints (targetValue).
[0123] Causal relationship and impact relationship modeling: Causal relationship is the most important relationship type in the knowledge graph, used to model the impact mechanism of parameter changes on quality results. The causes relationship is defined to represent direct causal relationships, including impact strength (strength), time delay (delay), and nonlinearity (nonlinearity) attributes. The affects relationship is used to represent a relatively broad impact relationship, without emphasizing strict causality. The directionality of the relationship is naturally expressed through the directed edges of RDF, while the positiveAffects and negativeAffects properties are introduced to refine the nature of the impact. For complex multi-factor impact relationships, the complexAffects relationship is defined, which is associated with an InfluenceFunction. The function contains specific mathematical expressions and parameter settings. Time delay is represented by the hasDelay data property, supporting the modeling of lag effects and dynamic response characteristics.
[0124] The hierarchical relationship is used to express the classification relationship and the inclusion relationship between concepts, which is realized by the standard OWL class hierarchy and object property hierarchy. In addition to the simple is-a relationship, the part-of relationship is defined to represent the composition relationship, the component-of relationship to represent the component relationship, and the belongs-to relationship to represent the affiliation relationship. These relationships support the hierarchical organization of knowledge and multi-level reasoning. Spatial relationships are expressed through locatedIn, adjacentTo, upstream, downstream, etc. to support spatial reasoning and propagation analysis. Time relationships are expressed through before, after, during, overlaps, etc. to support temporal reasoning and event correlation analysis. Quantitative relationships are expressed through hasValue, hasThreshold, hasTarget, etc. data properties to connect concept entities and specific numerical data.
[0125] The dynamic characteristics of the production process need to be modeled through dynamic relationships to support the expression of state changes and process evolution. The StateTransition class is defined to represent state transitions, which are connected through the fromState and toState relationships before and after the transition, and the triggeredBy relationship to connect the conditions or events that trigger the transition. The event concept Event is introduced as the driving factor of state changes, which has attributes such as occurrence time (occurTime), duration (duration), severity (severity), etc. Process evolution is modeled through the ProcessEvolution class, which describes the trend of parameter changes over time, including trend type (trendType), change rate (changeRate), prediction interval (predictionInterval), etc. A versioning relationship management mechanism is established to mark the valid period of the relationship through validFrom and validTo timestamps, supporting dynamic updating and historical tracing of knowledge;
[0126] Based on the modeling of mechanism knowledge and equation integration, the penetration of resin in paper fibers during impregnation is a key physical process that follows the extended form of Darcy's law, with penetration speed related to pressure gradient, liquid viscosity, and porous medium permeability factors:
[0127] ;
[0128] where: v is the penetration speed, k is the effective permeability, μ is the dynamic viscosity, P is the pressure gradient, ρ is the liquid density, g is the acceleration of gravity, σ is the surface tension, θ is the contact angle, and r is the pore radius;
[0129] In the knowledge graph, the mechanism equation is represented by the MechanisticEquation entity, and each physical quantity in the equation is mapped to the corresponding Parameter entity through the hasVariable relationship. The calibration results of the equation parameters are stored in the EquationParameter entity, including parameter values, calibration accuracy, and applicable range information.
[0130] The drying kinetics equation is:
[0131] ;
[0132] where X is the wet basis moisture content, t is the time, h m is the mass transfer coefficient; A is the heat transfer area, m s is the dry basis mass, Y s is the surface vapor concentration, Y ∞ is the ambient vapor concentration, k(T, RH) is the drying constant (a function of temperature and relative humidity), is the equilibrium moisture content;
[0133] The effect of temperature on the drying constant follows the Arrhenius relationship:
[0134] ;
[0135] where: k0 is the pre-exponential factor, E a is the activation energy, R is the gas constant, and T is the absolute temperature;
[0136] The curing process of thermosetting resin follows the law of chemical reaction kinetics, and the change of curing degree with time can be described by the Kamal model:
[0137] ;
[0138] where α is the curing degree, k1 and k2 are the reaction rate constants, and m and n are the reaction orders;
[0139] In the knowledge graph, a mechanism knowledge network is formed, which is connected to the related process through the hasEquation relationship, connected to the affected process parameters through the affectsParameter relationship, and connected to the physical and chemical principles through the basedOnPrinciple relationship.
[0140] Expert experience is converted into standard fuzzy inference rules, and the Mamdani inference method is used;
[0141] The typical quality control rules are as follows:
[0142] Rule R1:
[0143] ;
[0144] wherein concentration is concentration; temperature is temperature; speed is speed; resin content uniformity is resin content uniformity;
[0145] Rule R2:
[0146] ;
[0147] wherein viscosity is viscosity; pressure is pressure; penetration effect is penetration effect;
[0148] Rule R3:
[0149] ;
[0150] wherein drying temperature is drying temperature; humidity is humidity; foaming risk is foaming risk;
[0151] The triggering strength of each rule is calculated by norm calculation:
[0152] ;
[0153] wherein, is the matching degree (activation degree, membership) of rule i, that is, the activation strength of the fuzzy rule under the current input; represents the multi norm operation, the membership degrees of each input variable are integrated to obtain the overall matching degree; represents the current value x n of the nth input variable. n The membership degree of fuzzy set A
[0154] Based on the above knowledge framework, mechanism knowledge modeling and equation integration and fuzzy reasoning rules in the field of impregnated paper production, a knowledge graph of the impregnated paper production process is constructed.
[0155] In this embodiment, combined with the fused data set and knowledge graph, a digital twin model of the impregnated paper production line is constructed, specifically as follows:
[0156] Based on the fused dataset and knowledge graph, the digital twin model of the impregnated paper production line adopts a five-dimensional architecture, including physical space, virtual space, connection layer, data layer, and service layer. The physical space includes the actual production equipment, sensor network, control system, and material flow. The virtual space contains the digital representation of the physical system, integrating geometric model, physical model, behavioral model, and knowledge model. The connection layer enables real-time bidirectional interaction between the physical and virtual spaces, including data acquisition, status synchronization, and control command issuance. The data layer manages structured knowledge from the fused dataset and knowledge graph. The service layer provides various intelligent services based on the digital twin model, including status monitoring, fault diagnosis, and performance prediction.
[0157] The digital twin model operates in two modes: real-time synchronization and predictive simulation. In real-time synchronization mode, the digital twin model maintains time synchronization with the physical system, receiving sensor data in real time to update its state and achieve accurate mapping of the physical system. In predictive simulation mode, forward simulation is performed based on the current state and boundary conditions to predict the future evolution trajectory of the system, providing forward-looking support for decision-making. The fusion of these two modes is achieved through a parallel computing architecture. The real-time synchronization thread ensures the current accuracy of the model, while the predictive simulation thread provides future insights. A mode switching mechanism is established to dynamically select the appropriate operating mode based on application requirements and computing resources.
[0158] In this embodiment, the virtual space includes a digital representation of the physical system, integrating geometric models, physical models, behavioral models, and knowledge models. Specifically:
[0159] The geometric model establishes a three-dimensional digital representation of the production line, including equipment geometric models, plant space models, and material flow path models. Accurate geometric information is obtained by importing CAD data. The physical model integrates the mechanistic equations in the knowledge graph to establish a mathematical model describing the physical phenomena.
[0160] The behavioral model describes the dynamic response characteristics of the system, including control system behavior, equipment motion behavior, and material flow behavior; the control system behavior adopts a state-space model.
[0161] ;
[0162] Where x is the state vector, u is the input vector, y is the output vector, w and v are the process noise and measurement noise, and A, B, C, and D are the system matrices; The time derivative of the state vector;
[0163] The dynamic behavior of the equipment is governed by multibody dynamics equations:
[0164] ;
[0165] where q is the generalized coordinate, M is the mass matrix, C is the damping matrix, K is the stiffness matrix, Q is the generalized force, Q c is the constraint force; is the generalized acceleration vector; is the generalized velocity vector; t represents the time t;
[0166] The knowledge model adopts a hybrid modeling method combining mechanism models and data-driven models:
[0167] ;
[0168] where f physics is the physical mechanism model, f ML is the machine learning model, and a is the fusion weight, which is dynamically adjusted by an adaptive algorithm: is the predicted output value of the knowledge model; is the physical model parameter vector; is the machine learning model parameter vector;
[0169] ;
[0170] where e physics and e ML are the historical errors of the physical model and the machine learning model, respectively, s is the activation function, w α is the weight parameter vector; x(t) is the input feature (process parameters, etc.) at the current time; e physics (t) is the historical prediction error of the physical model at time t; e ML (t) is the historical prediction error of the machine learning model at time t; b α is the bias parameter, which is used to adjust the reference value of the weight.
[0171] In this embodiment, in the prediction simulation mode, forward simulation is performed based on the current state and boundary conditions to predict the future evolution trajectory of the system, as follows:
[0172] Based on time series analysis and recurrent neural networks, short-term prediction is performed, and a long short-term memory network is used to capture time sequence dependencies:
[0173] ;
[0174] ;
[0175] ;
[0176] ;
[0177] ;
[0178] ;
[0179] where, is the hidden state at time step t; is the forget gate; is the input gate; is the candidate value, i.e., the candidate memory cell at the current time step; is the cell state, representing the memory cell state at the current time step; is the output gate; is the hidden state at the current time step; 、 、 、 are the weight matrices of the forget gate, input gate, candidate value, and output gate, respectively; 、 、 、 are the bias vectors of the forget gate, input gate, candidate value, and output gate, respectively; σ is the sigmoid function; ⊙ is the Hadamard product; X t represents the input at time step t;
[0180] The mid-term prediction is performed using a state-space model:
[0181] ;
[0182] where, x t is the state vector at time t; x t+1 is the state vector at the next time step (t+1); u t is the control input vector at time t; y t is the observed output vector at time t; A is the state transition matrix; B is the control input matrix; C is the output matrix; D is the feedforward matrix; w t is the process noise vector; v t is the observation noise vector;
[0183] The parameter matrices are obtained through system identification:
[0184] ;
[0185] where, θ is the parameter vector to be estimated; is the optimal estimate of the parameters; N is the total length of the observation data; y t is the actual observation value at time t; is the model prediction value based on the parameters θ; represents the squared norm of the prediction error; λ is the regularization coefficient.
[0186] In this embodiment, a multi-objective optimization decision framework is implemented to achieve collaborative optimization of different time scales, as follows:
[0187] Production process optimization involves multiple conflicting objectives such as quality, cost, efficiency, and energy consumption. A multi-objective optimization method is used to solve the Pareto optimal solution set:
[0188] minF(x)=[f1(x),f2(x),f3(x),f4(x)] T ;
[0189] subject to: g(x)≤0,h(x)=0;
[0190] where f1(x) is the quality index (maximization), f2(x) is the production cost (minimization), f3(x) is the energy consumption index (minimization), and f4(x) is the production efficiency (maximization); g(x) represents the maximum load of equipment; h(x) represents the production quota;
[0191] Reinforcement learning adaptive control uses the deep deterministic policy gradient algorithm to achieve adaptive control policy learning, where the Actor network is updated as follows:
[0192] ;
[0193] where is the parameter vector of the Actor network; is the deterministic policy function; s t is the state at time t; a is the action variable; is the state distribution of the behavior policy; β represents the behavior policy; is the Critic network; s is the state; is the gradient of the Q function with respect to the action; is the parameter of the Critic network; represents the integral of the state under the behavior policy distribution; J is the objective function;
[0194] Critic network update:
[0195] ;
[0196] where y t is the TD target value; is the experience replay buffer; r t is the immediate reward obtained; s t+1 is the next state;
[0197] Target network soft update:
[0198] ;
[0199] wherein, are target Critic network parameters; are target Actor network parameters; τ is a soft update coefficient;
[0200] Based on the prediction ability of the digital twin model, a nonlinear model predictive control is used to realize rolling optimization:
[0201] ;
[0202] ;
[0203] wherein, L(x, u) is a stage cost function, V f (x) is a terminal cost function, N is a prediction horizon length; k is a time step index within the prediction horizon; f(x k , u k ) is a system state transition function; is a control input constraint set; is a state constraint set; is a terminal constraint set; x k , u k respectively represent a system state vector and a control input vector at the kth step;
[0204] The cost function is designed as a weighted multi-objective form:
[0205] ;
[0206] wherein, w1, w2, w3 are weight coefficients, Q, R, S are weight matrices, Δu is a control increment, u ref is a control reference value vector, x ref is a state reference value vector.
[0207] Those skilled in the art will appreciate that embodiments of the present application can be provided as methods, systems, or computer program products. Accordingly, the present application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage media, etc.) having computer-usable program code embodied therein.
[0208] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks.
[0209] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks.
[0210] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks.
[0211] The above description is only preferred embodiments of the present application, and is not intended to limit the present application to other forms described above. Any person skilled in the art can make modifications or improvements to the above-mentioned disclosed technical content without departing from the technical solutions of the present application. However, any simple modification, equivalent change and modification of the above-mentioned embodiments without departing from the technical solutions of the present application, according to the technical essence of the present application, still belongs to the protection scope of the present application.
Claims
1. A multi-parameter collaborative optimization control method for impregnated paper production process based on reinforcement learning, characterized in that, The method comprises the following steps: S1: Establish a comprehensive data collection network to realize panoramic perception of the production process and obtain real-time data flow; S2: Based on data fusion technology, intelligently integrate data of different sources and types to obtain a fused data set; S3: According to the fused data set and based on the process mechanism, a knowledge graph of the impregnated paper production process is constructed; S4: A digital twin model of the impregnated paper production line is constructed in combination with the fused data set and the knowledge graph; S5: According to the digital twin model, the production plan and the constraint conditions, a multi-objective optimization decision framework is designed to realize collaborative optimization at different time scales; According to the fused data set and based on the process mechanism, the knowledge graph of the impregnated paper production process is constructed as follows: The semantic structure of the field concept is defined, the classification relationship, composition relationship and dependency relationship between the concepts are established, and the knowledge framework of the impregnated paper production field is formed; Based on mechanism knowledge modeling and equation integration, the penetration of resin in paper fibers in the impregnation process is a key physical process, which follows the extended form of Darcy's law, and the penetration speed is related to the pressure gradient, liquid viscosity and porous medium permeability factors: ; where: v is the permeation velocity, k is the effective permeability, μ is the dynamic viscosity, is the pressure gradient, p is the liquid density, g is the gravitational acceleration, σ is the surface tension, θ is the contact angle, and r is the pore radius. In the knowledge graph, the mechanism equation is represented by a MechanisticEquation entity, and each physical quantity in the equation is mapped to a corresponding Parameter entity, and the connection is established through a hasVariable relationship; the calibration results of the equation parameters are stored in an EquationParameter entity, including parameter values, calibration accuracy and applicable range information; The drying kinetics equation is: ; where X is the wet basis moisture content, t is time, h m is the mass transfer coefficient; A is the heat transfer area, m s is the dry basis mass, Y s is the surface vapor concentration, Y ∞ is the ambient vapor concentration, k(T, RH) is the drying constant, is the equilibrium moisture content; The influence of temperature on the drying constant follows the Arrhenius relationship: ; where: k0 is the pre-exponential factor, E a is the activation energy, R is the gas constant, and T is the absolute temperature; The curing process of thermosetting resin follows the law of chemical reaction kinetics, and the change of curing degree with time can be described by the Kamal model: ; Wherein, a is the curing degree, k1 and k2 are the reaction rate constants, and m and n are the reaction orders; In the knowledge graph, a mechanism knowledge network is formed, which is connected to the related process through a hasEquation relationship, connected to the affected process parameters through an affectsParameter relationship, and connected to the physical and chemical principles through a basedOnPrinciple relationship; Expert experience is converted into standard fuzzy reasoning rules, and the Mamdani reasoning method is adopted; Based on the above knowledge framework of the impregnated paper production field, mechanism knowledge modeling and equation integration and fuzzy reasoning rules, the knowledge graph of the impregnated paper production process is constructed.
2. The method according to claim 1, wherein The establishment of a comprehensive data collection network to realize panoramic perception of the production process and obtain real-time data flow is as follows: A group of sensors is deployed at the process nodes of impregnated paper production, concentration sensors, temperature sensors, liquid level sensors and pH value sensors are installed in the impregnation tank area; thickness detectors, speed sensors and tension sensors are configured in the coating area; temperature field sensor arrays, humidity sensors and hot air flow meters are arranged in the drying and curing area; roll diameter measurement and visual sensors are set in the winding area; For key equipment, including pumps, fans, heaters, and pressure rollers, vibration sensors, current sensors, and bearing temperature sensors are installed to monitor the health status of the equipment in real time. Through the collection of equipment operating parameters, an image of the equipment operating status is formed. A near-infrared spectrometer is deployed to detect the glue content of paper, a laser scanning device is set up to detect surface flatness, and a temperature and humidity sensor and a barometric pressure sensor are used to monitor the temperature, humidity, and barometric pressure environment parameters in the workshop. A data acquisition network is constructed to obtain the sensor data and collect real-time data streams.
3. The method of claim 2, wherein the method is characterized by, The data acquisition network adopts a hierarchical data acquisition network, including a field device layer, a data aggregation layer, and a data center layer, as follows: The field device layer connects all sensors to the field I / O module through the field bus according to the nearest principle, and the field bus uses the Profinet protocol. The data aggregation layer sets up data aggregation servers in each section and connects them to the central data center through a gigabit Ethernet network. The data center layer establishes a central database server cluster, uses a time series database design to optimize the storage and query efficiency of time series data, and is equipped with a data backup and recovery system to ensure data security.
4. The method of claim 1, wherein the method is characterized by, The data fusion technology is used to intelligently integrate data from different sources and types to obtain a fused data set, as follows: a unified data preprocessing pipeline is established to perform standardized processing on data from different sources. First, data cleaning operations are performed to identify and handle missing values, outliers, and duplicates. Then, data format unification is performed to convert all timestamps to the standard UTC format, retain significant digits for numerical data and unify units, establish a standard encoding mapping table for text data, and unify resolution and format for image data. Next, time alignment processing is performed to establish a time offset compensation mechanism and map all data to a unified time reference. For data with different sampling frequencies, resampling techniques are used to ensure the consistency of time series. Finally, data standardization is performed to normalize numerical data using Z-score standardization and categorical data using one-hot encoding or label encoding to ensure the comparability of data with different dimensions and ranges. A standardized multi-source data set is obtained. Based on the standardized multi-source data set, combined with the characteristics of the impregnated paper production process and business requirements, a hierarchical data fusion strategy is developed to obtain a fused data set.
5. The method of claim 4, wherein the method is characterized by, The hierarchical data fusion strategy includes data level fusion and feature level fusion, each level uses different fusion algorithm and weight distribution mechanism, as follows: the first layer data level fusion uses weighted average method, for the temperature of the impregnation tank, the measurement values of 9 thermocouples are fused by weighted average; for the equipment vibration data, an adaptive Kalman filter is used, the filter parameters are dynamically adjusted according to the vibration spectrum characteristics, the noise interference is suppressed, and the signal quality is improved; the second layer feature level fusion constructs a comprehensive feature vector, combines the measurement values of different physical quantities into a comprehensive index reflecting the system state, uses principal component analysis dimension reduction technology to project high-dimensional original features into a low-dimensional space, uses independent component analysis to separate mixed signals and extract independent source signal components; the nonlinear relationship between features is identified through mutual information and correlation analysis, and a feature correlation atlas is constructed; after two layers of fusion processing, the original multi-source data is converted into a fused data set.
6. The method of claim 1, wherein, The combined fused data set and knowledge graph are used to construct a digital twin model of the impregnated paper production line, as follows: based on the fused data set and the knowledge graph, the digital twin model of the impregnated paper production line adopts a five-dimensional architecture design, including a physical space, a virtual space, a connection layer, a data layer and a service layer; the physical space includes actual production equipment, a sensor network, a control system and material flow; the virtual space includes digital representation of the physical system, integrating geometric models, physical models, behavior models and knowledge models; the connection layer realizes real-time bidirectional interaction between the physical space and the virtual space, including data acquisition, state synchronization and control instruction issuing functions; the data layer manages structured knowledge from the fused data set and the knowledge graph; the service layer provides various intelligent services based on the digital twin model, including state monitoring, fault diagnosis and performance prediction; the digital twin model has two operating modes, real-time synchronization and predictive simulation, in the real-time synchronization mode, the digital twin model is time-synchronized with the physical system, real-time sensor data is received, the digital twin model state is updated, and accurate mapping of the physical system is realized; in the predictive simulation mode, forward simulation is performed based on the current state and boundary conditions to predict the future evolution trajectory of the system.
7. The method of claim 6, wherein the method is characterized by, The virtual space includes digital representation of the physical system, integrating geometric models, physical models, behavior models and knowledge models, as follows: the geometric model establishes a three-dimensional digital representation of the production line, including equipment geometric models, plant space models and material flow path models, accurate geometric information is obtained through CAD data import, the physical model integrates mechanism equations in the knowledge graph to establish mathematical models describing physical phenomena; the behavior model describes the dynamic response characteristics of the system, including control system behavior, equipment motion behavior and material flow behavior; the control system behavior uses a state space model: ; where x is a state vector, u is an input vector, y is an output vector, w and v are process noise and measurement noise, A, B, C, and D are system matrices; is a time derivative of the state vector; the equipment dynamics behavior uses a multi-body dynamics equation: ; where q is the generalized coordinate, M is the mass matrix, C is the damping matrix, K is the stiffness matrix, Q is the generalized force, Q c is the constraint force; is the generalized acceleration vector; is the generalized velocity vector; t denotes the time instant t; the knowledge model uses a hybrid modeling method combining mechanism models and data-driven models: ; Wherein, f physics is a physical mechanism model, f ML is a machine learning model, and a is a fusion weight dynamically adjusted by an adaptive algorithm: is a prediction output value of the knowledge model. is a physical model parameter vector. is a machine learning model parameter vector. ; where e physics and e ML are the historical errors of the physical model and the machine learning model, respectively, σ is an activation function, w α is a weight parameter vector; x(t) is the input feature at the current time; e physics (t) is the historical prediction error of the physical model at time t; e ML (t) is the historical prediction error of the machine learning model at time t; and b α is a bias parameter used to adjust the baseline value of the weight.
8. The method of claim 7, wherein the method is characterized by, in the predictive simulation mode, forward simulation is performed based on the current state and boundary conditions to predict the future evolution trajectory of the system, as follows: Short-term prediction is based on time series analysis and recurrent neural network, and long short-term memory network is used to capture time series dependence: ; ; ; ; ; ; wherein, is the hidden state at time step t; is the forget gate; is the input gate; is the candidate value, i.e., the candidate memory cell at the current time instant; is the cell state, representing the memory cell state at the current time instant; is the output gate; is the hidden state at the current time instant; , , , are weight matrices of the forget gate, the input gate, the candidate value, and the output gate, respectively; , , , are bias vectors of the forget gate, the input gate, the candidate value, and the output gate, respectively; σ is the sigmoid function; ⊙ is the Hadamard product; X t represents the input at time step t; Medium-term prediction is based on state space model: ; where x t is the state vector at time t; x t+1 is the state vector at the next time (t+1); u t is the control input vector at time t; y t is the observation output vector at time t; A is the state transition matrix; B is the control input matrix; C is the output matrix; D is the feedforward matrix; w t is the process noise vector; v t is the observation noise vector; Parameter matrix is obtained through system identification: ; where θ is the parameter vector to be estimated; is the optimal estimate of the parameter; N is the total length of the observation data; y t is the actual observation value at time t; is the model prediction value based on the parameter θ; denotes the squared norm of the prediction error; λ is the regularization coefficient.
9. The method of claim 8, wherein the method is characterized by, The multi-objective optimization decision framework realizes the collaborative optimization of different time scales, which is as follows: Production process optimization involves multiple conflicting objectives such as quality, cost, efficiency, and energy consumption, and multi-objective optimization method is used to solve the Pareto optimal solution set: minF(x) = [fl(x), f2(x), f3(x), f4(x)] T ; Subject to: g(x)≤0,h(x)=0; Where f1(x) is the quality index, f2(x) is the production cost, f3(x) is the energy consumption index, and f4(x) is the production efficiency; g(x) represents the maximum load of equipment; h(x) represents the production quota; Reinforcement learning adaptive control, using deep deterministic policy gradient algorithm to realize adaptive control strategy learning, where Actor network update: ; where, is the parameter vector of the Actor network; is the deterministic policy function; s t is the state at time t; a is the action variable; is the state distribution of the behavior policy; β represents the behavior policy; is the Critic network; s is the state; is the gradient of the Q function with respect to the action; is the parameter of the Critic network; represents the integral over the state under the behavior policy distribution; J is the objective function; Critic network update: ; where y t is the TD target value; is the experience replay buffer; r t is the obtained immediate reward; s t+1 is the next state; Target network soft update: ; wherein, are target Critic network parameters; are target Actor network parameters; τ is a soft update coefficient; Based on the prediction ability of digital twin model, nonlinear model predictive control is used to realize rolling optimization: ; ; where L(x, u) is a stage cost function, V f (x) is a terminal cost function, N is the prediction horizon length; k is the time step index within the prediction horizon; f(x k ,u k ) is a system state transition function; is a control input constraint set; is a state constraint set; is a terminal constraint set; x k ,u k denote the system state vector and control input vector at the k-th step, respectively. Cost function is designed in a weighted multi-objective form: ; wherein w1, w2, w3 are weight coefficients, Q, R, S are weight matrices, Δu is a control increment, u ref is a control reference value vector, x ref is a state reference value vector.
Citation Information
Patent Citations
Quality inspection early warning system for impregnated paper production
CN116609354A
Impregnated paper raw material management system fusing knowledge graph and neural network
CN116720819A