Ink path optimization control method, system, device and medium based on reinforcement learning
By employing a multi-sensor fusion and reinforcement learning-based ink path optimization control method, the problems of flow control accuracy and bubble suppression in OLED inkjet printing have been solved, achieving high precision and stability of the ink path system and supporting the manufacturing of high-resolution OLED displays.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIHUA LAB
- Filing Date
- 2026-02-14
- Publication Date
- 2026-05-26
Smart Images

Figure CN121697353B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of ink path control, and in particular to an ink path optimization control method, system, device and medium based on reinforcement learning. Background Technology
[0002] As an emerging manufacturing method to replace traditional vapor deposition processes, OLED inkjet printing technology has significant advantages in manufacturing large-size, flexible OLED displays, including high material utilization and low equipment manufacturing and maintenance costs. However, its industrialization process is severely hampered by core bottlenecks such as insufficient flow control accuracy, poor stability, and inadequate bubble suppression in the ink path system. Although existing technologies attempt to improve pulsation and accuracy by employing parallel dual-diaphragm pump structures or modular ink path designs, they largely rely on feedback from a single sensor, resulting in insufficient overall accuracy, stability, and adaptive adjustment capabilities. Summary of the Invention
[0003] This application aims to improve at least one technical problem in the background art.
[0004] This application provides a reinforcement learning-based ink path optimization control method, comprising: an application to an ink path control system, the ink path control system including a control device, an adjustment mechanism electrically connected to the control device, and multiple sensors, the ink path optimization control method comprising:
[0005] Acquire the status information of each sensor to obtain the status information of multiple sensors;
[0006] Multiple sensor state information is fused to generate a system state vector;
[0007] Based on the system state vector and a pre-defined multi-objective reinforcement learning decision model, a set of control instructions is generated.
[0008] According to the control instruction set, the driving adjustment mechanism executes the corresponding working mode.
[0009] According to some technical solutions of this application, the step of obtaining the state information of each sensor to obtain multiple sensor state information specifically includes:
[0010] The plurality of sensors include a flow sensor, a pressure sensor, an ultrasonic sensor, a viscosity sensor, and a temperature sensor;
[0011] Flow data in the ink path is collected using a flow sensor;
[0012] Pressure data in the pipeline is collected using pressure sensors;
[0013] Data on bubble concentration in ink is collected using an ultrasonic sensor.
[0014] Ink viscosity data is collected using a viscosity sensor;
[0015] Ink temperature data is collected using a temperature sensor;
[0016] Based on flow rate data, pressure data, bubble concentration data, viscosity data, and temperature data, multiple sensor status information is generated.
[0017] According to some technical solutions of this application, the process of fusing multiple sensor state information to generate a system state vector specifically includes:
[0018] Adaptive weighted averaging fusion is performed on flow rate data, pressure data, viscosity data, and temperature data from multiple sensor status information to generate static fusion components;
[0019] Kalman filtering is applied to fuse the bubble concentration data from the status information of the multiple sensors to generate a dynamic fusion component.
[0020] The static and dynamic fusion components are integrated to generate the system state vector.
[0021] According to some technical solutions of this application, after driving the adjustment mechanism to execute the corresponding working mode according to the control instruction set, it further includes:
[0022] After the regulating mechanism executes the corresponding working mode, it returns to obtain the status information of each sensor in order to obtain new status information of multiple sensors.
[0023] A new system state vector is generated based on the new state information from multiple sensors;
[0024] Based on the new system state vector, optimize and update the multi-objective reinforcement learning decision model.
[0025] According to some technical solutions of this application, the generation of a control instruction set based on the system state vector and a preset multi-objective reinforcement learning decision model specifically includes:
[0026] The regulating mechanism includes a diaphragm pump, a proportional valve, and a bubble suppression device;
[0027] Based on the system state vector, the diaphragm pump speed adjustment amount is determined by a multi-objective reinforcement learning decision model.
[0028] Based on the system state vector, the proportional valve opening adjustment amount is determined through a multi-objective reinforcement learning decision model.
[0029] Based on the bubble concentration index in the system state vector, the working mode of the bubble suppression device is determined by a multi-objective reinforcement learning decision model.
[0030] A set of control commands is generated based on the adjustment amount of the diaphragm pump speed, the adjustment amount of the proportional valve opening, and the working mode of the bubble suppression device.
[0031] According to some technical solutions of this application, the method of determining the working mode of the bubble suppression device based on the bubble concentration index in the system state vector and using a multi-objective reinforcement learning decision model specifically includes:
[0032] The operating modes include mild inhibition mode, moderate inhibition mode, and strong inhibition mode;
[0033] If the bubble concentration index is lower than the preset first threshold, the bubble suppression device is determined to be in mild suppression mode by a multi-objective reinforcement learning decision model.
[0034] If the bubble concentration index is between the preset first threshold and the preset second threshold, the bubble suppression device is determined to be in a moderate suppression mode by a multi-objective reinforcement learning decision model.
[0035] If the bubble concentration index reaches or exceeds the second threshold, the bubble suppression device is determined to be in strong suppression mode by a multi-objective reinforcement learning decision model.
[0036] This application also provides a reinforcement learning-based ink path optimization control system, which includes a control device, an adjustment mechanism electrically connected to the control device, and a plurality of sensors; the control device is used to execute the reinforcement learning-based ink path optimization control method described in the above technical solution.
[0037] This application also provides a reinforcement learning-based ink path optimization control device, the ink path optimization control device comprising: a memory and at least one processor, wherein the memory stores instructions;
[0038] At least one of the processors invokes the instructions in the memory to cause the ink path optimization control device to execute the various steps of the reinforcement learning-based ink path optimization control method as described above.
[0039] This application also provides a computer-readable storage medium storing instructions that, when executed by a processor, implement the various steps of the reinforcement learning-based ink path optimization control method described above.
[0040] The ink path optimization control method provided in this application has at least the following beneficial effects: by fusing information from multiple sensors and achieving multi-objective adaptive optimization through reinforcement learning, it effectively suppresses the generation and accumulation of bubbles, and can dynamically adjust the strategy according to ink characteristics, environmental changes, etc., thereby improving the accuracy, stability and robustness of the ink path system, and thus providing key technical support for the industrialization breakthrough of high-resolution OLED manufacturing. Attached Figure Description
[0041] Figure 1 A flowchart illustrating the ink path optimization control method provided in this application embodiment;
[0042] Figure 2 This is a schematic diagram of the structure of the ink path optimization control system provided in the embodiments of this application;
[0043] Figure 3 This is a schematic diagram of the ink path optimization control device provided in an embodiment of this application. Detailed Implementation
[0044] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0045] In the description of this application, it should be understood that the orientation descriptions, such as up, down, front, back, left, right, etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation or be constructed or operated in a specific orientation. Therefore, they should not be construed as limiting the present invention.
[0046] In the description of this application, unless otherwise expressly defined, terms such as "setup," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this application in conjunction with the specific content of the technical solution.
[0047] The following is combined Figures 1 to 3 Embodiments of the present invention will be described.
[0048] Current mainstream control methods, such as pneumatic control, are greatly affected by ambient air pressure and have slow response times. Single-diaphragm pump speed control inherently suffers from flow pulsation, which leads to droplet volume errors and fails to meet the stringent requirement of ±5% for ultra-high resolution OLED displays with a pixel density of 300ppi and above. Furthermore, air bubbles in the ink path system are generated due to temperature changes, pipe sealing, or mechanical movement. Traditional optical sensors struggle to effectively detect micron-sized bubbles, and their accumulation can cause ejection failures or ink droplet distortion. Simultaneously, the system must balance multiple objectives, including flow accuracy, bubble suppression, and response speed. These objectives involve inherent trade-offs; for example, improving flow accuracy requires reducing pump speed, but this may exacerbate bubble formation. Traditional PID control or rule-based control strategies, with their fixed parameters and lack of predictability, struggle to achieve adaptive balance. While existing technologies attempt to improve pulsation and accuracy using parallel dual-diaphragm pump structures or modular ink path designs, they largely rely on single sensor feedback and fail to fundamentally address the intelligent decision-making problems arising from multi-parameter coupling, nonlinearity, and time-varying characteristics.
[0049] Based on the above, this application provides a reinforcement learning-based ink path optimization control method, which is applied to an ink path control system. The ink path control system includes a control device, an adjustment mechanism electrically connected to the control device, and multiple sensors. The ink path optimization control method includes:
[0050] S100, acquire the status information of each sensor to obtain multiple sensor status information; this method first senses the physical state of the ink path through heterogeneous sensors, such as the state of flow rate, pressure, and bubbles, to form multi-dimensional raw data.
[0051] S200 fuses the state information from multiple sensors to generate a system state vector. Then, a specific fusion algorithm integrates these heterogeneous, potentially noisy, data into a system state vector that can comprehensively and accurately characterize the current operating status of the system.
[0052] The S300 generates a set of control commands based on the system state vector and a pre-set multi-objective reinforcement learning decision model. This vector is fed into a pre-trained multi-objective reinforcement learning decision model. During the training phase, the multi-objective reinforcement learning decision model has learned how to make optimal decisions under complex multi-objective constraints, such as high precision, low bubble, and high stability. Therefore, it can output a set of optimal control commands based on the real-time state vector.
[0053] S400, based on the control instruction set, drives the adjustment mechanism to execute the corresponding working mode. Finally, the instruction is transmitted to the specific adjustment mechanism, translating the control instructions into actions of the specific actuators, thereby achieving precise and adaptive control of the ink path system. For example, the adjustment mechanism comprises a diaphragm pump, a proportional valve, and a bubble suppression device. Based on the instruction set, the diaphragm pump speed, the proportional valve opening, and the activation state of the bubble suppression device are adjusted respectively. For instance, by adjusting the bubble suppression device to a light mode, the flow path is adjusted, the flow rate is reduced, and slight vibration promotes bubble buoyancy; adjusting the bubble suppression device to a medium mode allows for localized pressure to promote bubble dissolution; adjusting the bubble suppression device to a strong mode temporarily interrupts printing and flushes the ink path system.
[0054] In this way, by introducing an intelligent control architecture that combines multi-sensor fusion perception with multi-objective reinforcement learning decision-making, the Molu control system has the ability to fully perceive the state, intelligently optimize decisions, and accurately execute actions. When facing nonlinear, time-varying, and multi-objective coupled Molu systems, it can improve control accuracy and adaptability, thereby achieving a dynamic balance between multiple competing objectives such as flow accuracy and bubble suppression.
[0055] Therefore, by fusing information from multiple sensors and achieving multi-objective adaptive optimization through reinforcement learning, the generation and accumulation of bubbles can be effectively suppressed. Furthermore, the strategy can be dynamically adjusted according to ink characteristics and environmental changes, thereby improving the accuracy, stability, and robustness of the ink path system. This can provide key technical support for the industrialization breakthrough of high-resolution OLED manufacturing.
[0056] To obtain data reflecting a specific aspect of the ink path's state, in some embodiments, step S100 involves acquiring the state information of each sensor to obtain multiple sensor state information, including:
[0057] Based on the key state parameters during ink path operation, the multiple sensors include a flow sensor, a pressure sensor, an ultrasonic sensor, a viscosity sensor, and a temperature sensor.
[0058] S110 collects flow data in the ink path through a flow sensor for real-time monitoring of ink flow in the ink supply line, with a sampling frequency ≥1kHz;
[0059] S120 uses a pressure sensor to collect pressure data in the pipeline, which is used to detect pressure changes at key points in the pipeline and identify air bubbles and blockages;
[0060] S130 uses an ultrasonic sensor to collect bubble concentration data in ink, which is used to detect the content and distribution of tiny bubbles in ink.
[0061] S140 collects ink viscosity data through a viscosity sensor to monitor changes in ink viscosity and compensate for the effects of temperature.
[0062] The S150 uses a temperature sensor to collect ink temperature data for monitoring ink temperature and for viscosity compensation.
[0063] S160 generates multiple sensor status information based on flow rate data, pressure data, bubble concentration data, viscosity data, and temperature data.
[0064] By deploying five types of sensors, a multi-sensor status information system is formed that comprehensively reflects the current operating status of the ink system. This system constructs a comprehensive sensing system covering the fluid dynamics (flow rate, pressure, etc.), fluid physical properties (viscosity, temperature, etc.), and key defect indicators such as bubble concentration. Thus, by configuring a highly targeted and comprehensive sensor combination, a raw data foundation is provided for subsequent information fusion and intelligent decision-making.
[0065] In some embodiments, step S200 involves fusing multiple sensor state information to generate a system state vector, specifically including:
[0066] S210, adaptive weighted average fusion is performed on the flow rate, pressure, viscosity, and temperature data from multiple sensor status information to generate a static fusion component. For data reflecting the static or slowly changing properties of the system, such as flow rate, pressure, viscosity, and temperature, whose changes are relatively stable, adaptive weighted average fusion is used. Specifically, adaptive weighted average fusion is performed using the following formula:
[0067]
[0068] as well as
[0069]
[0070] in, This is the final output value after fusion. The number of similar sensors participating in the fusion. For the first The raw measurement values of each sensor, For the first The adaptive weights of each sensor satisfy the constraint that the sum of the weights is 1.
[0071] S220: Kalman filtering is applied to the bubble concentration data from multiple sensor status information to generate a dynamic fused component. For data reflecting bubble concentration, which has strong temporal correlation and dynamic change characteristics, Kalman filtering is used for fusion. Kalman filtering is an optimal recursive estimation algorithm that fuses data from multiple ultrasonic bubble sensors through recursive prediction-update steps to achieve the optimal estimate of bubble concentration. Specifically, it fuses the current measurement value with the optimal estimate from the previous moment, effectively filtering out measurement noise and predicting the future trend of bubble concentration, thus obtaining the optimal estimate of bubble state and improving the accuracy of bubble concentration estimation. This reduces the shortcomings of traditional methods in tracking transient and changing signals like bubbles, such as tracking lag and high noise interference, thereby improving the accuracy, reliability, and information content of the system state vector.
[0072] S230 integrates the static and dynamic fusion components to generate a system state vector, thereby comprehensively reflecting the static state and dynamic changes of the ink path system.
[0073] Further, in step S230, the static fusion component and the dynamic fusion component are integrated to generate a system state vector. Specifically, this includes:
[0074] S231, extract real-time flow rate, real-time pressure, ink viscosity and ink temperature from the static fusion components, and extract four types of core parameters that can directly reflect the basic operating status of the ink system.
[0075] S232, Extract the bubble concentration index from the dynamic fusion component; this index is the core parameter for evaluating the bubble suppression effect and determining the bubble suppression mode.
[0076] S233, based on historical sequences of real-time flow and pressure values, calculates the flow and pressure change trends; for example, the historical sequence includes real-time flow and pipeline pressure values from the past n moments. The historical state sequence helps the model adapt to the time-varying characteristics of the pipeline system, further enhancing the adaptive capability of the control strategy. Based on the historical sequences of real-time flow and pressure values, time-series data analysis algorithms, such as the finite difference method, are used to calculate the flow and pressure change trends. The sensor-collected flow or pressure time-series data is first preprocessed by moving average filtering and a unified sampling interval, then the rate of change is calculated using first-order difference to determine the direction and magnitude of change, second-order difference is used to determine trend stability, and finally, it is quantified into discrete level identifiers, forming flow or pressure change trend data that can be integrated into the system state vector, adapting to the needs of reinforcement learning decision-making.
[0077] Therefore, by capturing the dynamic evolution patterns of the two types of core parameters, reinforcement learning models can predict the direction of system state evolution, adjust control strategies in advance, reduce control lag, improve system response speed, and compensate for the inadequacy of relying solely on real-time values to reflect the changing state of the system.
[0078] S234 integrates real-time flow rate, real-time pressure, bubble concentration, ink viscosity, ink temperature, flow rate variation trends, pressure variation trends, and historical state sequences to obtain a system state vector. This results in a comprehensive and dimensionally complete system state vector, significantly improving the comprehensiveness and depth of state representation.
[0079] In some embodiments, after driving the adjustment mechanism to execute the corresponding working mode according to the control instruction set, step S400 further includes starting the next cycle immediately after each control action is executed:
[0080] S510, after the regulating mechanism executes the corresponding working mode, returns to obtain the status information of each sensor in order to obtain new status information of multiple sensors, that is, to re-collect new status information of multiple sensors through multiple types of sensors.
[0081] S520 generates a new system state vector based on new sensor state information; the new system state vector is used to obtain the ink path state after the control command is executed.
[0082] S530 optimizes and updates the multi-objective reinforcement learning decision model based on the new system state vector. By collecting new sensor state information, it evaluates the results of completing control commands, calculating the reward value using the model's internal reward function. Based on this reward value, the system uses it as a feedback signal, combined with key risk indicators from the new system state vector, to fine-tune the decision model's internal parameters using reinforcement learning algorithms, such as existing gradient descent algorithms. For example, when a significant increase in bubble risk is detected, the weight of the bubble suppression reward may be temporarily increased. In this way, based on the calculated reward value, the model's parameters are updated through reinforcement learning algorithms, adjusting the decision strategy so that the model can output better control commands in subsequent control processes.
[0083] For example, the weighting coefficients in a multi-objective reward function Automatic updates are achieved through a dynamic preference adjustment mechanism, namely:
[0084]
[0085] in, For learning rate, For the first The optimization objective is... Weighting coefficients at time points For the first The optimization objective is... The weight coefficients after the update at each time step.
[0086] In this way, based on the preset dynamic weight adjustment strategy, the weight coefficients of the multi-objective reward function in the multi-objective reinforcement learning decision model are updated according to the key risk indicators in the new system state vector; by adjusting the weights, multiple competitive objectives can be dynamically balanced under different working conditions.
[0087] Therefore, through online updates, the decision-making model can continuously adapt to specific production environments, optimize control strategies for specific ink characteristics, and even compensate for slow changes in equipment performance. This significantly improves the system's long-term robustness, adaptability, and intelligence.
[0088] In some embodiments, before generating the control instruction set based on the system state vector and a preset multi-objective reinforcement learning decision model, step S300 further includes:
[0089] S310, construct a Markov decision process including a state space and an action space; its state space is the system state vector, and its action space is the control action space composed of control commands from the regulating mechanism; specifically, the state space is designed as follows:
[0090]
[0091] in, This represents the real-time traffic value. Indicates the pipeline pressure value. Indicates the bubble concentration index. Indicates the viscosity value of the ink. Indicates the ink temperature value. Indicates the trend of traffic changes. Indicates the trend of pressure changes. This represents the sequence of inertial historical states used to characterize the system.
[0092] The motion space is designed as follows:
[0093]
[0094] in, This is the adjustment amount for the diaphragm pump speed. This is the adjustment amount of the proportional valve opening. The bubble suppression module has three operating modes, for example, 0 indicates off, 1 indicates mild, and 2 indicates strong.
[0095] S320 constructs a multi-objective reward function that includes a flow accuracy reward, a bubble suppression reward, a system stability reward, and a control efficiency reward. The multi-objective reward function transforms the four optimization objectives of flow accuracy, bubble suppression, system stability, and control efficiency into quantifiable scalar reward values. Each reward item corresponds to a different optimization direction, and a penalty mechanism is triggered when constraints are violated.
[0096] The reward function is the core indicator for evaluating the overall performance of the entire ink path control system. It integrates multiple optimization objectives, such as flow accuracy, bubble suppression, system stability, and control efficiency, into a quantifiable scalar value to guide the decision-making of the reinforcement learning agent. Specifically, it is designed using a multi-objective weighted approach, and its specific form is defined as follows:
[0097]
[0098] in, The total reward function value. For traffic accuracy rewards, For bubble suppression rewards, To control efficiency rewards, As a reward for system stability, To control efficiency rewards, For flow accuracy weighting coefficient, This is the bubble suppression weighting coefficient. For system stability weighting coefficients, To control the efficiency weighting coefficient, To constrain the penalty weighting coefficient for violations.
[0099] The reward functions for each item in the formula are defined as follows:
[0100] The flow accuracy bonus is:
[0101]
[0102] in, This represents the actual flow rate. The target flow rate;
[0103] The bubble suppression reward is:
[0104]
[0105] in, This is the bubble weighting coefficient. This represents the current bubble concentration.
[0106] The system stability reward is:
[0107]
[0108] in, For stability weighting coefficients, The rate of change of flow rate;
[0109] The control efficiency bonus is:
[0110]
[0111] in, Efficiency weighting coefficient;
[0112] The penalty for violating the restrictions is:
[0113]
[0114] in, The indicator functions for each constraint violation;
[0115] Specifically, the constraint violation indication function mainly includes two items, one of which is a safety pressure penalty. ,in , The maximum safe pressure is the point at which a significant, fixed penalty is applied to prevent pipe rupture when the system pressure exceeds the maximum safe pressure. One type is the critical bubble penalty. ,in , This indicates the danger threshold for bubble concentration. When the detected bubble concentration exceeds the critical value, it may immediately cause nozzle blockage or spray failure, triggering a severe penalty.
[0116] The S330 uses a reinforcement learning algorithm for iterative training based on the state space, action space, and multi-objective reward function until convergence, resulting in a well-trained multi-objective reinforcement learning decision model. During training, the model continuously explores the impact of different actions on the system state, adjusting its decision strategy based on the reward values fed back from the reward function, until iterative training reaches convergence, forming a well-trained decision model capable of achieving multi-objective balanced optimization. Through systematic training, a highly optimized initial decision model that balances multiple objectives and performance indicators, such as precise flow rate accuracy and bubble suppression, can be obtained.
[0117] For example, the ink flow control problem is formalized as a Markov Decision Process (MDP): the system state is the system state vector; the optional actions are all possible combinations of control commands, such as different combinations of pump speed and valve opening. Then, a multi-objective reward function is designed, which integrates multiple optimization objectives, such as high flow accuracy, low bubble concentration, small system fluctuations, and efficient adjustment actions, into a scalar reward signal through weighted summation. Finally, on simulation environments or historical data, reinforcement learning algorithms are used, such as existing PPO (Proximal Policy Optimization) and DDPG (Deep Deterministic Policy Gradient), to allow the model to repeatedly try different approaches, learning the optimal control policy by maximizing the cumulative reward.
[0118] In some embodiments, the regulating mechanism includes a diaphragm pump, a proportional valve, and a bubble suppression device; specifically, the diaphragm pump suppresses flow pulsation by optimizing the start-up time difference between the two diaphragm pumps; the proportional valve precisely adjusts its opening according to the control quantity; and the bubble suppression device includes an ultrasonic defoaming device and a vacuum defoaming device.
[0119] S300 generates a control instruction set based on the system state vector and a preset multi-objective reinforcement learning decision model, specifically including:
[0120] S310, based on the system state vector, determines the diaphragm pump speed adjustment amount through a multi-objective reinforcement learning decision model; for example, based on the flow data, flow change trend and other information in the system state vector, the speed adjustment amount of the diaphragm pump is determined to achieve precise flow control;
[0121] S320, based on the system state vector, determines the proportional valve opening adjustment amount through a multi-objective reinforcement learning decision model; for example, it determines the proportional valve opening adjustment amount based on pipeline pressure data, flow demand and other information to optimize pipeline pressure distribution.
[0122] S330, based on the bubble concentration index in the system state vector, determines the working mode of the bubble suppression device through a multi-objective reinforcement learning decision model; thus, based on the bubble concentration index in the system state vector, the working mode of the bubble suppression device is determined to suppress bubbles in a targeted manner.
[0123] S340 generates a control command set based on the diaphragm pump speed adjustment, proportional valve opening adjustment, and the operating mode of the bubble suppression device. When faced with a specific system state vector, the model's internal policy network simultaneously outputs three decisions: a speed adjustment command for the diaphragm pump, a proportional valve opening adjustment command for fine-tuning flow and pressure, and a mode selection command for the bubble suppression device. Thus, the model comprehensively considers the entire system's state and implements coordinated control of the flow generation mechanism, flow regulation mechanism, and bubble treatment mechanism to ensure synchronized optimization of the ink path system's flow, pressure, and bubble state.
[0124] In some embodiments, S340, based on the bubble concentration index in the system state vector, determines the working mode of the bubble suppression device through a multi-objective reinforcement learning decision model, specifically including: based on the degree of influence of bubble concentration in the ink path system on printing quality, the system compares the real-time monitored bubble concentration index with a preset safety threshold, that is, a preset first threshold and a preset second threshold, triggers different levels of countermeasures, and divides the working mode into mild suppression mode, moderate suppression mode and strong suppression mode;
[0125] S341, if the bubble concentration index is lower than the preset first threshold, the bubble suppression device is determined to be in mild suppression mode by the multi-objective reinforcement learning decision model. In mild suppression mode, the bubble concentration index is less than the first threshold. The bubble is promoted to float by gentle means such as adjusting the flow path, reducing the flow rate, and slight vibration, with prevention as the main focus, so as to avoid excessive suppression from affecting the stability of the system and to minimize the overall energy consumption and impact.
[0126] S342, if the bubble concentration index is between a preset first threshold and a preset second threshold, the bubble suppression device is determined to be in a moderate suppression mode by a multi-objective reinforcement learning decision model; in the moderate suppression mode, i.e., the first threshold ≤ bubble concentration index < second threshold, the ultrasonic defoaming device is activated to promote bubble dissolution through local pressure. The existing bubbles are locally treated through active suppression methods such as ultrasonic defoaming.
[0127] S343, if the bubble concentration index reaches or exceeds the second threshold, the bubble suppression device is determined to be in strong suppression mode by a multi-objective reinforcement learning decision model. Strong suppression mode, i.e., bubble concentration index ≥ the second threshold, is considered an emergency and the highest level of processing is initiated, such as vacuum degassing, or even pausing printing for rinsing, to thoroughly remove high-concentration bubbles, avoid ejection failure or ink droplet distortion, and ensure system safety.
[0128] Therefore, through a tiered strategy, the system can take the most appropriate suppression measures according to the severity of the bubble threat, which can significantly improve the bubble suppression effect and effectively avoid printing defects caused by bubbles. This enables efficient and targeted suppression of bubbles of different concentrations, which in turn helps to minimize interference with the normal operation of the ink system while ensuring the bubble suppression effect, thus ensuring the quality stability of high-resolution OLED printing.
[0129] This application also provides a reinforcement learning-based ink path optimization control system, which includes a control device, an adjustment mechanism electrically connected to the control device, and multiple sensors; the control device is used to execute the reinforcement learning-based ink path optimization control method described in the above embodiment. The ink path optimization control system includes an acquisition module 100, a fusion module 200, a generation module 300, and a drive module 400. The acquisition module is used to acquire multi-sensor state information through multiple sensors; the fusion module is used to fuse the multi-sensor state information to generate a system state vector; the generation module is used to generate a control command set based on the system state vector and a preset multi-objective reinforcement learning decision model; the drive module 400 drives the adjustment mechanism to execute the corresponding working mode according to the control command set.
[0130] In some optional embodiments, the overall architecture of the Molu optimization control system includes a multi-sensor fusion perception layer, an intelligent decision-making layer, and a precise execution layer. The multi-sensor fusion perception layer is responsible for collecting multi-dimensional state information of the Molu system. The intelligent decision-making layer adopts the multi-objective reinforcement learning (MORL) framework to model the Molu control problem as a Markov decision process. The precise execution layer transforms the control commands of the intelligent decision-making layer into the actions of specific actuators.
[0131] This application also provides a reinforcement learning-based ink path optimization control device, the ink path optimization control device comprising: a memory and at least one processor, the memory storing instructions; at least one processor calling the instructions in the memory to cause the ink path optimization control device to execute the various steps of the reinforcement learning-based ink path optimization control method as described in the above embodiments.
[0132] Figure 3This is a schematic diagram of the structure of an ink path optimization control device 600 provided in an embodiment of the present invention. The ink path optimization control device 600 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 610 (e.g., one or more processors) and a memory 620, and one or more storage media 630 (e.g., one or more mass storage devices) storing application programs 633 or data 632. The memory 620 and storage media 630 can be temporary or persistent storage. The program stored in the storage media 630 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the ink path optimization control device 600. Furthermore, the processor 610 may be configured to communicate with the storage media 630 and execute the series of instruction operations in the storage media 630 on the ink path optimization control device 600 to implement the steps of the methods provided in the above-described method embodiments.
[0133] The ink path optimization control device 600 may also include one or more power supplies 640, one or more wired or wireless network interfaces 650, one or more input / output interfaces 660, and / or one or more operating systems 631, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 3 The illustrated electronic device structure does not constitute a limitation on the electronic device and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0134] This application also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of the above-described reinforcement learning-based ink path optimization control method.
[0135] The preferred embodiments of the present invention have been described in detail above, but the present disclosure is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention, and these equivalent modifications or substitutions are all included within the scope defined by the claims of the present disclosure.
Claims
1. A method for ink path optimization control based on reinforcement learning, characterized in that: An ink path optimization control method is applied to an ink path control system, which includes a control device, an adjustment mechanism electrically connected to the control device, and multiple sensors. The status information of each sensor is acquired to obtain multiple sensor status information; the sensor status information includes flow rate data, bubble concentration data, pressure data, viscosity data, and temperature data. Multiple sensor state information is fused to generate a system state vector; Based on the system state vector and a pre-defined multi-objective reinforcement learning decision model, a set of control instructions is generated. According to the control instruction set, the driving adjustment mechanism executes the corresponding working mode; The process of fusing multiple sensor state information to generate a system state vector specifically includes: performing adaptive weighted averaging fusion on flow rate data, pressure data, viscosity data, and temperature data from multiple sensor state information to generate a static fusion component; performing Kalman filtering fusion on bubble concentration data from multiple sensor state information to generate a dynamic fusion component; and integrating the static and dynamic fusion components to generate a system state vector. The process of generating a control command set based on the system state vector and a preset multi-objective reinforcement learning decision model specifically includes: the regulating mechanism comprising a diaphragm pump, a proportional valve, and a bubble suppression device; determining the diaphragm pump speed adjustment amount based on the system state vector using the multi-objective reinforcement learning decision model; determining the proportional valve opening adjustment amount based on the system state vector using the multi-objective reinforcement learning decision model; determining the bubble suppression device's operating mode based on the bubble concentration index in the system state vector using the multi-objective reinforcement learning decision model; and generating a control command set based on the diaphragm pump speed adjustment amount, the proportional valve opening adjustment amount, and the bubble suppression device's operating mode. Before generating the control instruction set, the process based on the system state vector and a pre-defined multi-objective reinforcement learning decision model further includes: constructing a Markov decision process containing a state space and an action space; constructing a multi-objective reward function including a flow accuracy reward term, a bubble suppression reward term, a system stability reward term, and a control efficiency reward term; and iteratively training the multi-objective reinforcement learning model using a reinforcement learning algorithm based on the state space, action space, and multi-objective reward function until the model converges to obtain the trained multi-objective reinforcement learning decision model. After driving the adjustment mechanism to execute the corresponding working mode according to the control instruction set, the method further includes: after the adjustment mechanism executes the corresponding working mode, returning to obtain the state information of each sensor to obtain new multiple sensor state information; generating a new system state vector based on the new multiple sensor state information; and optimizing and updating the multi-objective reinforcement learning decision model according to the new system state vector.
2. The reinforcement learning-based ink path optimization control method according to claim 1, characterized by: The acquisition of the status information of each sensor to obtain multiple sensor status information specifically includes: The plurality of sensors include a flow sensor, a pressure sensor, an ultrasonic sensor, a viscosity sensor, and a temperature sensor; Flow data in the ink path is collected using a flow sensor; Pressure data in the pipeline is collected using pressure sensors; Data on bubble concentration in ink is collected using an ultrasonic sensor. Ink viscosity data is collected using a viscosity sensor; Ink temperature data is collected using a temperature sensor; Based on flow rate data, pressure data, bubble concentration data, viscosity data, and temperature data, multiple sensor status information is generated.
3. The ink path optimization control method based on reinforcement learning according to claim 1, characterized in that: The method for determining the operating mode of the bubble suppression device based on the bubble concentration index in the system state vector using a multi-objective reinforcement learning decision model specifically includes: The operating modes include mild inhibition mode, moderate inhibition mode, and strong inhibition mode; If the bubble concentration index is lower than the preset first threshold, the bubble suppression device is determined to be in mild suppression mode by a multi-objective reinforcement learning decision model. If the bubble concentration index is between the preset first threshold and the preset second threshold, the bubble suppression device is determined to be in a moderate suppression mode by a multi-objective reinforcement learning decision model. If the bubble concentration index reaches or exceeds the second threshold, the bubble suppression device is determined to be in strong suppression mode by a multi-objective reinforcement learning decision model.
4. A reinforcement learning-based ink path optimization control system, characterized in that: It includes a control device, an adjustment mechanism electrically connected to the control device, and a plurality of sensors; the control device is used to execute the ink path optimization control method based on reinforcement learning as described in any one of claims 1-3.
5. A reinforcement learning-based ink path optimization control device, characterized in that, The ink path optimization control device includes: a memory and at least one processor, wherein the memory stores instructions; At least one of the processors invokes the instructions in the memory to cause the ink path optimization control device to perform the steps of the reinforcement learning-based ink path optimization control method as described in any one of claims 1-3.
6. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instructions are executed by the processor, they implement the steps of the reinforcement learning-based ink path optimization control method as described in any one of claims 1-3.