Coating production line control method, device, equipment, storage medium and program product
By real-time monitoring and dynamically adjusting the production environment and operating parameters of the coating production line equipment, and using the deep reinforcement learning model to optimize the parameters, the problems of uneven film quality and light transmittance fluctuations caused by the fixation of equipment parameters are solved, the production efficiency and quality are improved, and energy consumption and defective rate are reduced.
Patent Information
- Application Number
- CN202510548928.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-08
AI Technical Summary
The fixed parameters of existing coating production line equipment lead to problems such as uneven film quality and fluctuations in light transmittance, which require frequent shutdown and adjustment, affecting production efficiency and quality consistency.
By monitoring the production environment and operating parameters of coating production line equipment, the pre-trained deep reinforcement learning model is used to adjust the parameters, the equipment operation parameters are optimized in real time, and the equipment parameters are dynamically adjusted in combination with the deep reinforcement learning algorithm.
Significantly improve the production efficiency and quality of coatings, reduce energy consumption and defective rates, and improve equipment utilization and production stability.
Smart Images

Figure CN120447490A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of production scheduling technology, and in particular to a coating production line control method, device, equipment, storage medium and program product. Background Art
[0002] With the growing demand for high-performance coated glass in the architectural, automotive, and photovoltaic industries, optimizing coating production line efficiency and quality stability has become a key industry challenge. Existing glass coating production lines often use fixed parameter control or manual intervention based on empirical rules. For example, these lines drive automated equipment based on preset process parameters (such as sputtering power and annealing temperature), relying on regular manual inspections to adjust the production process.
[0003] However, in the existing production process, equipment parameters are fixed. Environmental fluctuations (such as temperature drift, load changes) or differences in material batches can easily lead to problems such as uneven film thickness and transmittance fluctuations. Frequent shutdowns are required to adjust equipment parameters, seriously restricting production line efficiency and quality consistency.
[0004] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of this application is to provide a coating production line control method, device, equipment, storage medium and program product, aiming to solve the technical problem that the fixed parameters of existing coating production line equipment affect the coating production quality.
[0006] To achieve the above objectives, the present application proposes a coating production line control method, the method comprising: Monitor the production environment parameters and equipment operation parameters of coating production line equipment; By using a pre-trained deep reinforcement learning model, parameter adjustment is performed according to the production environment parameters and equipment operating parameters to obtain an equipment parameter adjustment amount, wherein the pre-trained deep reinforcement learning model is obtained by training a preset initial deep reinforcement learning model based on pre-acquired historical coating production data; The equipment operating parameters of the coating production line equipment are adjusted according to the equipment parameter adjustment amount, and the coating production line equipment after the parameter adjustment is controlled to perform coating production.
[0007] In one embodiment, before the step of adjusting parameters of the pre-trained deep reinforcement learning model according to the production environment parameters and the equipment operating parameters, the method further includes: Acquiring historical coating production data, wherein the historical coating production data includes historical production environment data, historical coating control data, and historical coating detection data; Based on the historical production environment data, historical coating control data and historical coating detection data, a preset initial deep reinforcement learning model is trained to obtain a pre-trained deep reinforcement learning model.
[0008] In one embodiment, before the step of training a preset initial deep reinforcement learning model based on the historical production environment data, the historical coating control data, and the historical coating detection data to obtain a pre-trained deep reinforcement learning model, the step further includes: Defining a state space according to a production environment parameter range of the coating production line equipment; Defining an action space according to a coating control parameter range of the coating production line equipment; Define the reward function based on the preset film layer standard parameters and production target parameters; A model is constructed based on the state space, action space, and reward function to obtain an initial deep reinforcement learning model.
[0009] In one embodiment, the step of training a preset initial deep reinforcement learning model based on the historical production environment data, historical coating control data, and historical coating detection data to obtain a pre-trained deep reinforcement learning model includes: Performing time series alignment processing on the historical production environment data, the historical coating control data, and the historical coating detection data to generate initial training samples; Performing data enhancement on the initial training samples by a sliding window method to obtain data-enhanced training samples; The data-augmented training samples are input into the initial deep reinforcement learning model for offline iterative training to obtain a pre-trained deep reinforcement learning model.
[0010] In one embodiment, the production environment parameters include one or more of equipment load, substrate specifications, temperature and pressure, and the equipment operation parameters include one or more of magnetron sputtering power, reaction gas flow and annealing temperature.
[0011] In one embodiment, after the step of adjusting the equipment operating parameters of the coating production line equipment according to the equipment parameter adjustment amount and controlling the coating production line equipment after the parameter adjustment to perform coating production, the method further includes: Performing film layer detection on the glass produced by the coating production line equipment to obtain film layer detection parameters; Calculating a deviation between the film layer detection parameter and a preset film layer standard parameter, and comparing the deviation value with a preset deviation threshold; If the deviation value exceeds a preset deviation threshold, the reward function of the pre-trained deep reinforcement learning model is updated based on the deviation value, and the pre-trained deep reinforcement learning model is retrained.
[0012] In addition, to achieve the above-mentioned purpose, the present application also proposes a coating production line control device, which includes: Monitoring module, used to monitor the production environment parameters and equipment operating parameters of the coating production line equipment; an adjustment module, configured to adjust parameters according to the production environment parameters and equipment operating parameters using a pre-trained deep reinforcement learning model to obtain an equipment parameter adjustment amount, wherein the pre-trained deep reinforcement learning model is obtained by training a preset initial deep reinforcement learning model based on pre-acquired historical coating production data; A control module is used to adjust the equipment operating parameters of the coating production line equipment according to the equipment parameter adjustment amount, and control the coating production line equipment after the parameter adjustment to perform coating production.
[0013] In addition, to achieve the above-mentioned purpose, the present application also proposes a coating production line control device, which includes: a memory, a processor, and a computer program stored on the memory and runnable on the processor, and the computer program is configured to implement the steps of the coating production line control method described above.
[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by the processor, the steps of the coating production line control method described above are implemented.
[0015] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the coating production line control method described above.
[0016] This application provides a coating production line control method. First, this application monitors production environment parameters and equipment operating parameters in real time to obtain the current status of the equipment in a timely manner. Then, using the monitored production environment parameters and equipment operating parameters, the deep reinforcement learning model is used to adjust the parameters to dynamically optimize the equipment operating parameters based on the real-time monitoring data. Finally, the equipment operating parameters are adjusted in real time according to the adjustment amount to ensure that the equipment operates in the optimal state. Combined with the deep reinforcement learning algorithm, real-time monitoring and dynamic adjustment of equipment parameters solves the problem of fixed equipment parameters in existing coating production lines affecting the quality of coating production, significantly improves the efficiency and quality of coating production, reduces energy consumption and defective product rates, and improves equipment utilization and production stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 A schematic flow chart of the first embodiment of the coating production line control method of the present application; Figure 2 A schematic flow chart of the second embodiment of the coating production line control method of the present application; Figure 3 This is a schematic diagram of the module structure of the coating production line control device according to an embodiment of the present application; Figure 4 Schematic diagram of the equipment structure of the hardware operating environment involved in the coating production line control method in the embodiment of the present application.
[0020] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0021] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0022] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0023] The main solution of the embodiment of the present application is: monitoring the production environment parameters and equipment operating parameters of the coating production line equipment; adjusting the parameters according to the production environment parameters and equipment operating parameters through a pre-trained deep reinforcement learning model to obtain the equipment parameter adjustment amount, and the pre-trained deep reinforcement learning model is obtained by training a preset initial deep reinforcement learning model based on the pre-acquired historical coating production data; adjusting the equipment operating parameters of the coating production line equipment according to the equipment parameter adjustment amount, and controlling the coating production line equipment after parameter adjustment to perform coating production.
[0024] Because existing glass coating production lines generally use fixed equipment parameters and process flows, parameters such as temperature, time, and chemical reagent concentration in the cleaning and pretreatment stages are usually pre-set before production and cannot be adjusted in real time according to the dynamic changes of contaminants on the glass surface; during the coating process, key parameters such as the air pressure of the vacuum chamber, target material power, and deposition rate remain unchanged, making it difficult to adapt to the requirements of different coating materials and film thicknesses; parameters such as annealing and tempering temperature and time in the heat treatment stage are also mostly fixed values and cannot be optimized according to the stress state of the film layer; parameters in the detection and post-processing stages (such as cutting speed and edge grinding accuracy) also lack adaptability, resulting in limited production efficiency and quality.
[0025] This fixed parameter production method results in unstable film quality and is unable to monitor and adjust parameters such as temperature, pressure, and flow during the coating process in real time, which can easily lead to problems such as uneven film thickness and insufficient adhesion.
[0026] In this application, first, the production environment parameters and equipment operating parameters are monitored in real time to obtain the current status of the equipment in a timely manner. Then, the monitored production environment parameters and equipment operating parameters are used to adjust the parameters through a deep reinforcement learning model to dynamically optimize the equipment operating parameters based on the real-time monitoring data. Finally, the equipment operating parameters are adjusted in real time according to the adjustment amount to ensure that the equipment operates in the optimal state. Combined with the deep reinforcement learning algorithm, real-time monitoring and dynamic adjustment of equipment parameters solves the problem of fixed equipment parameters in existing coating production lines affecting coating production quality, significantly improving the efficiency and quality of coating production, reducing energy consumption and defective product rates, and improving equipment utilization and production stability.
[0027] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of implementing the above functions, a coating production line control device, etc. The following uses the coating production line control device as an example to illustrate this embodiment and the following embodiments.
[0028] Based on this, the embodiment of the present application provides a coating production line control method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the coating production line control method of the present application.
[0029] In this embodiment, the coating production line control method includes steps S10 to S30: Step S10, monitoring the production environment parameters and equipment operating parameters of the coating production line equipment; It should be noted that coating production line equipment can be equipment involved in coating-related processes, including core coating equipment with coating functions and auxiliary equipment (such as cleaning machines, annealing furnaces, testing instruments, etc.). Coating equipment can be divided into physical vapor deposition equipment and chemical vapor deposition equipment according to the coating process. The coating equipment will provide a high vacuum or low-pressure environment for coating, including vacuum pumps, valves, and cavities. The glass substrate will be coated in the cavity of the coating equipment. Production environment parameters refer to the environmental conditions in which the coating production line equipment is located during operation, including equipment load, substrate specifications, temperature, and pressure. Equipment operating parameters refer to the parameters that need to be controlled during the operation of the coating production line equipment, such as magnetron sputtering power, reaction gas flow rate, and annealing temperature. The specific parameter categories are determined by the current coating equipment. Parameters that need to be controlled during the operation of physical vapor deposition equipment include magnetron sputtering power, and parameters that need to be controlled during the operation of chemical vapor deposition equipment include reaction gas flow rate. These are not limited here.
[0030] It is understood that by using various sensors installed on the coating equipment to monitor and collect production environment parameters and equipment operating parameters in real time, the current status of the equipment can be promptly obtained, providing data support for subsequent parameter adjustments. For example, a temperature sensor collects the temperature within the coating equipment cavity, a pressure sensor collects the pressure within the coating equipment cavity, and a flow sensor collects the flow rate of the reaction gas output from the coating equipment.
[0031] For example, in a glass coating production line, the system monitors the temperature in real time through a thermocouple sensor installed in the cavity, monitors the pressure in the coating equipment cavity through a pressure transmitter, and monitors the flow rate of the reaction gas through a mass flow controller.
[0032] In a feasible embodiment, the production environment parameters include one or more of equipment load, substrate specifications, temperature and pressure, and the equipment operation parameters include one or more of magnetron sputtering power, reaction gas flow and annealing temperature.
[0033] It should be noted that the production environment parameters mainly involve parameters related to the coating, annealing and other processes in the glass coating production process. The equipment load is calculated by monitoring the current changes of the equipment using the current sensor installed in the coating equipment, and represents the load of the coating equipment. The substrate specification refers to the size and initial thickness of the glass substrate to be coated in the coating equipment. The temperature refers to the temperature inside the coating equipment cavity during the coating process. The pressure refers to the pressure inside the coating equipment cavity during the coating process. The magnetron sputtering power refers to the output power of the magnetron sputtering power supply of the magnetron sputtering method in the coating process. The reactive gas flow rate refers to the flow rate of the reactive gas in the chemical vapor deposition process in the coating process. The annealing temperature refers to the temperature inside the annealing furnace during the annealing process of the coating production.
[0034] Step S20, adjusting parameters according to the production environment parameters and equipment operating parameters using a pre-trained deep reinforcement learning model to obtain equipment parameter adjustment amounts, wherein the pre-trained deep reinforcement learning model is obtained by training a preset initial deep reinforcement learning model based on pre-acquired historical coating production data; It should be noted that the deep reinforcement learning model is an algorithmic model based on deep learning and reinforcement learning, which learns optimal strategies through the interaction between the agent and the environment. The pre-trained deep reinforcement learning model is the result of training the initial model based on historical coating production data.
[0035] It's understandable that, in order to leverage models trained with historical data to adjust equipment parameters in real time and ensure production process stability and film quality uniformity, the collected production environment parameters and equipment operating parameters are fed into a pre-trained deep reinforcement learning model. The model then outputs equipment parameter adjustments based on the strategies learned from historical data. For example, if the glass requires a physical vapor deposition coating process and the coating equipment has magnetron sputtering capabilities, the model might output a magnetron sputtering power adjustment plan based on the deviation in the coating equipment's current temperature and pressure.
[0036] Step S30 , adjusting the equipment operating parameters of the coating production line equipment according to the equipment parameter adjustment amount, and controlling the coating production line equipment after parameter adjustment to perform coating production.
[0037] It should be noted that the equipment parameter adjustment refers to the extent to which the equipment's operating parameters need to be adjusted based on the output of the deep reinforcement learning model. Coating production refers to the manufacturing process of depositing one or more thin films on a material surface through a specific process, aiming to impart new functional properties (such as wear resistance, corrosion resistance, conductivity, and optical properties) to the substrate. This refers to the process of depositing a functional thin film layer that better meets quality requirements on the new glass surface after the coating production line equipment has readjusted its operating parameters.
[0038] It can be understood that the control system adjusts the equipment's operating parameters based on the adjustments output by the model. By adjusting these parameters in real time based on the adjustments, the equipment can be ensured to operate at its optimal state, improving production efficiency and film quality. Applying the model's optimization results to actual production ensures a highly efficient and stable coating process. For example, adjustments can be made to the magnetron sputtering power supply, the gas flow rate controller, or the annealing furnace temperature. After these adjustments, the equipment performs coating production according to the new parameters.
[0039] In a feasible embodiment, after the step of adjusting the equipment operating parameters of the coating production line equipment according to the equipment parameter adjustment amount and controlling the coating production line equipment after the parameter adjustment to perform coating production, the method further includes: Step S401, performing film layer detection on the glass produced by the coating production line equipment to obtain film layer detection parameters; It should be noted that film layer inspection refers to the quality inspection of coated glass, and film layer inspection parameters refer to the specific parameters in the inspection results, such as film layer thickness, adhesion, transmittance, reflectivity, etc.
[0040] It is understandable that after coating, the coating layer is tested using equipment such as spectrometers and ellipsometers to obtain the actual quality parameters of the coated glass film. This provides a basis for subsequent model updates, ensures that product quality during the production process meets standards, and promptly identifies and corrects production problems. For example, after the glass is coated, the film thickness is measured using a spectrometer, adhesion is tested using the cross-cut method, and transmittance is measured using a transmittance tester.
[0041] Step S402, calculating a deviation between the film layer detection parameter and a preset film layer standard parameter, and comparing the deviation value with a preset deviation threshold; It should be noted that the preset deviation threshold refers to the preset allowable deviation range of the film layer detection parameter, and the deviation value refers to the difference between the detection parameter and the standard parameter.
[0042] It is understood that in order to assess quality fluctuations during the coating production process and promptly identify any need for model adjustment, the detected film parameters are compared with the standard parameters, a deviation value is calculated, and then the deviation value is compared with a preset deviation threshold to determine whether it exceeds the preset deviation threshold. For example, if the standard film thickness is 300nm, the detected thickness is 305nm, the deviation value is 5nm, and the preset threshold is 10nm. If 5nm is less than the threshold, then the deviation is determined to be within the allowable range.
[0043] Step S403: If the deviation value exceeds a preset deviation threshold, the reward function of the pre-trained deep reinforcement learning model is updated based on the deviation value, and the pre-trained deep reinforcement learning model is retrained.
[0044] It should be noted that updating the reward function refers to adjusting the parameters of the reward function according to the new deviation value, and retraining refers to the process of retraining the pre-trained deep reinforcement learning model using historical coating production data and / or newly monitored production environment parameters and equipment operating parameters.
[0045] It is understandable that in order to dynamically adjust the optimization target of the model and adapt to changes in the production process, when the deviation value exceeds the preset deviation threshold, the weight parameters of the reward function are adjusted according to the deviation value between the film detection parameters and the preset film standard parameters, and the historical coating production data and / or newly monitored production environment parameters and equipment operation parameters and other coating production data are reused to train the model, ensuring that the model can dynamically control the coating production in real time while continuously optimizing the production process and improving product quality and production efficiency.
[0046] In this embodiment, by real-time detection of film layer detection parameters after glass coating, the reward function of the model is dynamically adjusted and the model is retrained when there is a quality deviation in the coating production, thereby improving the adaptability and accuracy of the model and ensuring the stability of the coating quality.
[0047] In a feasible embodiment, after the step of performing film layer detection on the glass produced by the coating production line equipment and obtaining film layer detection parameters, the step further includes: Step S4011, storing the production environment parameters, equipment operation parameters and film layer detection parameters in an online experience replay pool; It should be noted that the online experience replay pool is a data structure that stores historical data. It is used to save information such as production environment parameters, equipment operating parameters, and film layer detection parameters for subsequent model fine-tuning.
[0048] It is understood that the production environment parameters, equipment operating parameters, and film layer detection parameters collected during each coating production process are stored in real time in the online experience playback pool. For example, parameters such as temperature, pressure, magnetron sputtering power, reaction gas flow, film thickness, and adhesion are stored in the database to ensure that the model can be optimized using the latest production data, improving its adaptability and accuracy.
[0049] Step S4012: Read historical data from the online experience replay pool based on a preset model fine-tuning cycle; It should be noted that the model fine-tuning cycle refers to a pre-set time interval used to regularly read historical data and perform model fine-tuning.
[0050] It is understandable that in order to improve the real-time and adaptability of the model and ensure that the model can maintain optimal performance under different working conditions, historical data is read from the online experience replay pool according to the preset model fine-tuning cycle.
[0051] Step S4013: Perform online fine-tuning on the model parameters of the pre-trained deep reinforcement learning model based on the historical data.
[0052] It should be noted that online fine-tuning refers to the process of further optimizing model parameters using the latest historical data based on the pre-trained model.
[0053] It is understandable that historical data is fed into a pre-trained deep reinforcement learning model, which then undergoes online fine-tuning using the reinforcement learning algorithm. For example, the model parameters are adjusted using the environmental parameters in the historical data as states, the control parameters as actions, and the detection data as reward signals. Through fine-tuning, the model learns new strategies to optimize magnetron sputtering power and reaction gas flow to reduce film thickness deviations, ensuring that the model consistently maintains optimal performance in actual coating production.
[0054] In this embodiment, through online experience replay and periodic model fine-tuning, the model can be continuously optimized to adapt to changes in the production process, improve the adaptability and accuracy of the model, and ensure the stability of the coating quality.
[0055] This embodiment provides a coating production line control method. First, production environment parameters and equipment operating parameters are monitored in real time to promptly obtain the current status of the equipment. Then, using the monitored production environment parameters and equipment operating parameters, a deep reinforcement learning model is used to adjust the parameters, dynamically optimizing the equipment operating parameters based on the real-time monitoring data. Finally, the equipment operating parameters are adjusted in real time based on the adjustment amount to ensure that the equipment operates in the optimal state. Incorporating a deep reinforcement learning algorithm, real-time monitoring and dynamic adjustment of equipment parameters solves the problem of fixed equipment parameters in existing coating production lines affecting coating production quality. This significantly improves the efficiency and quality of coating production, reduces energy consumption and defective product rates, and enhances equipment utilization and production stability.
[0056] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above introduction and will not be described in detail later. Figure 2 , Figure 2 This is a flow chart of the second embodiment of the coating production line control method of the present application.
[0057] In this embodiment, before the step of adjusting parameters of the pre-trained deep reinforcement learning model according to the production environment parameters and equipment operating parameters, the following steps are further included: Step S201, acquiring historical coating production data, wherein the historical coating production data includes historical production environment data, historical coating control data, and historical coating detection data; It should be noted that historical coating production data refers to various data recorded in the past coating production process, including historical production environment data (such as equipment load, temperature, pressure, etc.), historical coating control data (such as power, flow, etc.) and historical coating detection data (such as film thickness, adhesion, transmittance, reflectivity, etc.).
[0058] It is understandable that relevant data on historical coating production can be obtained from the historical database. For example, environmental data such as equipment load, temperature, and pressure during the operation of the coating equipment can be extracted from the production log; control data such as power and flow rate during the operation of the coating equipment can be extracted from the control log; and film layer inspection data such as film thickness, adhesion, transmittance, and reflectivity after glass coating can be extracted from the film layer inspection records.
[0059] Step S202: Based on the historical production environment data, historical coating control data, and historical coating detection data, a preset initial deep reinforcement learning model is trained to obtain a pre-trained deep reinforcement learning model.
[0060] It should be noted that the initial deep reinforcement learning model refers to an untrained deep reinforcement learning model, and model training refers to optimizing the model through historical data so that it can learn the optimal parameter adjustment strategy.
[0061] It is understood that to ensure that the model can effectively optimize the production process in real-world applications, the acquired historical data is fed into the initial deep reinforcement learning model. This model is then trained using a reinforcement learning algorithm to learn the optimal parameter adjustment strategy, thereby improving the model's accuracy and reliability. For example, the model can be trained to learn the optimal strategy using environmental parameters from historical data as states, control parameters as actions, and detection data as reward signals.
[0062] For example, a deep Q-network algorithm was used to train the initial model, using historical data on temperature and pressure as states, magnetron sputtering power and gas flow as actions, and the deviation of film thickness from the standard value as a reward signal. After multiple iterations, the model learned how to adjust power and flow to optimize film thickness under different environmental conditions.
[0063] In a feasible embodiment, before the step of training a preset initial deep reinforcement learning model based on the historical production environment data, historical coating control data, and historical coating detection data to obtain a pre-trained deep reinforcement learning model, the step further includes: Step S2021, defining a state space according to the production environment parameter range of the coating production line equipment; It should be noted that the production environment parameter range refers to the value range of the environmental parameters when the equipment is operating normally, and the state space refers to the parameter range that represents the environmental state in the model.
[0064] It is understandable that in order to ensure that the model can cover all possible production environment conditions, the production environment parameter range is determined according to the specifications and historical data of the coating production line equipment, the state space is defined, and the environmental parameter range that the model needs to handle is clarified, providing clear input boundaries for subsequent model training and optimization.
[0065] Step S2022, defining an action space according to a coating control parameter range of the coating production line equipment; It can be understood that the coating control parameter range refers to the value range of the control parameters of the coating production line equipment during normal operation, and the action space refers to the parameter range in the model that represents the control actions that can be taken.
[0066] It is understandable that in order to ensure that the control parameters output by the model are within the actual capabilities of the equipment, the coating control parameter range is determined based on the control capabilities and historical data of the coating production line equipment, the action space is defined, and the range of control actions that the model can take is clarified, providing clear output boundaries for subsequent strategy learning.
[0067] Step S2023, defining a reward function based on preset film layer standard parameters and production target parameters; It should be noted that the standard parameters of the film layer refer to the standard values of the film layer quality, the production target parameters refer to the target values that need to be achieved during the production process, and the reward function refers to the function used to evaluate the quality of the model output action.
[0068] It's understandable that to ensure the model optimizes toward improving film quality and production efficiency during training, a reward function is defined based on the film standard parameters and production target parameters. This provides the model with an optimization objective and guides it in learning the optimal strategy. For example, the reward function could be the negative of the deviation of the film thickness from the standard value, with smaller deviations resulting in higher rewards.
[0069] For example, based on the standard value of the film thickness (such as 300nm) and the production goal (such as minimizing energy consumption), the reward function is defined as: reward = -|actual thickness - 300nm| + energy consumption weight term. The model will try to minimize thickness deviation and reduce energy consumption during the learning process.
[0070] Step S2024: construct a model based on the state space, action space, and reward function to obtain an initial deep reinforcement learning model.
[0071] It is understood that building a deep reinforcement learning model based on the state space, action space, and reward function provides the infrastructure for subsequent model training, ensuring that the model can effectively process data in the state space and action space and optimize according to the reward function. Specifically, a neural network model can be built using a deep learning framework, defining a state input layer, an action output layer, and a reward calculation module. For example, a deep Q-network model can be constructed, where the model input layer receives state space parameters, the output layer outputs action space parameters, and the intermediate layer uses a neural network to extract features and learn policies.
[0072] For example, a deep Q-network model is constructed. The input layer receives temperature, pressure, substrate specifications, and device load, and the output layer outputs adjustments for magnetron sputtering power, reaction gas flow, and annealing temperature. The intermediate layers, including multiple fully connected layers, are used to extract features and learn policies. A reward function is used to evaluate the quality of the output actions and guide model learning.
[0073] In this embodiment, based on the coating production conditions of the coating production line equipment and the film layer standards and production goals of the coating production, by defining the state space, action space and reward function, a deep reinforcement learning model adapted to the coating production line can be constructed, providing a basis for subsequent model training and real-time adjustment.
[0074] In a feasible implementation, the step of training a preset initial deep reinforcement learning model based on the historical production environment data, historical coating control data, and historical coating detection data to obtain a pre-trained deep reinforcement learning model includes: Step S2025, performing time series alignment processing on the historical production environment data, the historical coating control data, and the historical coating detection data to generate initial training samples; It should be noted that time series alignment refers to aligning the timestamps of data collected at different times. Initial training samples refer to organizing the aligned data into the sample format required for model training.
[0075] It can be understood that the timestamp alignment algorithm aligns historical production environment data, historical coating control data, and historical coating inspection data in time series to generate initial training samples. This ensures temporal consistency of data from different sources, improves the accuracy of training data, and provides a reliable data foundation for subsequent model training. For example, the timestamp of each data item is converted to a unified time base, and missing data is interpolated to generate training samples containing states, actions, and rewards.
[0076] Specifically, the system aligns the time series of production environment data (temperature, pressure), coating control (data power, flow) and coating detection data (film thickness, adhesion) in historical data. That is, for a certain time point t, the system extracts the temperature and pressure at time t, the power and flow at time t, and the film thickness and adhesion at time t+Δt to generate the initial data samples for model training.
[0077] Step S2026, performing data enhancement on the initial training samples using a sliding window method to obtain data-enhanced training samples; It should be noted that the sliding window method is a data augmentation technique that generates more training samples by sliding a window over time series data. Data augmentation refers to improving the generalization ability of a model by increasing the number or diversity of samples.
[0078] It can be understood that using the sliding window method to slide the window on the time series of the initial training samples to generate more training samples, increase the number and diversity of training samples, enable the model to better adapt to different production scenarios, and improve the robustness and adaptability of the model.
[0079] For example, for a sample sequence containing 100 time steps, assuming the window size is 10 time steps, the step size is 1, and a new sample is generated each time the slide is performed, 91 new training samples can be generated, each sample containing state, action, and reward data for 10 consecutive time steps.
[0080] Step S2027: Input the data-augmented training samples into the initial deep reinforcement learning model for offline iterative training to obtain a pre-trained deep reinforcement learning model.
[0081] It should be noted that the supervisory signal refers to the external signal used to guide model training, and offline iterative training means that the model is not updated in real time during the training process, but multiple iterative training is performed using historical data.
[0082] It is understandable that the data-augmented training samples are input into the initial deep reinforcement learning model, and offline iterative training is performed through the reinforcement learning algorithm until the model converges to obtain a pre-trained deep reinforcement learning model, ensuring that the model can learn the optimal parameter adjustment strategy and provide support for real-time adjustments in actual production.
[0083] For example, the augmented training samples were fed into an initial deep Q-network model, and the deviation of the film thickness from the standard value was used as the reward signal. The model was trained for 1,000 iterations. After each iteration, the model updated the parameters of the policy network, gradually learning the optimal action under different states.
[0084] In this embodiment, through time alignment and data enhancement, the generated training samples are used to perform offline iterative training on the initial deep reinforcement learning model, thereby improving the training effect and generalization ability of the model and making the model more stable in practical applications.
[0085] In this embodiment, historical coating production data is obtained and the initial deep reinforcement learning model is trained using the historical data, so that the model can learn the optimal parameter adjustment strategy, improve the accuracy and reliability of the model, and provide strong support for real-time adjustment.
[0086] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the coating production line control method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0087] This application also provides a coating production line control device, please refer to Figure 3 , the coating production line control device includes: Monitoring module 10, used to monitor the production environment parameters and equipment operating parameters of the coating production line equipment; An adjustment module 20 is configured to adjust parameters according to the production environment parameters and equipment operating parameters using a pre-trained deep reinforcement learning model to obtain an equipment parameter adjustment amount, wherein the pre-trained deep reinforcement learning model is obtained by training a preset initial deep reinforcement learning model based on pre-acquired historical coating production data; The control module 30 is used to adjust the equipment operating parameters of the coating production line equipment according to the equipment parameter adjustment amount, and control the coating production line equipment after the parameter adjustment to perform coating production.
[0088] Optionally, the adjustment module 20 is further configured to: Acquiring historical coating production data, wherein the historical coating production data includes historical production environment data, historical coating control data, and historical coating detection data; Based on the historical production environment data, historical coating control data and historical coating detection data, a preset initial deep reinforcement learning model is trained to obtain a pre-trained deep reinforcement learning model.
[0089] Optionally, the adjustment module 20 is further configured to: Defining a state space according to a production environment parameter range of the coating production line equipment; Defining an action space according to a coating control parameter range of the coating production line equipment; Define the reward function based on the preset film layer standard parameters and production target parameters; A model is constructed based on the state space, action space, and reward function to obtain an initial deep reinforcement learning model.
[0090] Optionally, the adjustment module 20 is further configured to: Performing time series alignment processing on the historical production environment data, the historical coating control data, and the historical coating detection data to generate initial training samples; Performing data enhancement on the initial training samples by a sliding window method to obtain data-enhanced training samples; The data-augmented training samples are input into the initial deep reinforcement learning model for offline iterative training to obtain a pre-trained deep reinforcement learning model.
[0091] Optionally, the production environment parameters include one or more of equipment load, substrate specifications, temperature and pressure, and the equipment operation parameters include one or more of magnetron sputtering power, reaction gas flow and annealing temperature.
[0092] Optionally, the control module 30 is further configured to: Performing film layer detection on the glass produced by the coating production line equipment to obtain film layer detection parameters; Calculating a deviation between the film layer detection parameter and a preset film layer standard parameter, and comparing the deviation value with a preset deviation threshold; If the deviation value exceeds a preset deviation threshold, the reward function of the pre-trained deep reinforcement learning model is updated based on the deviation value, and the pre-trained deep reinforcement learning model is retrained.
[0093] The coating production line control device provided in this application utilizes the coating production line control method described in the aforementioned embodiments, resolving the technical issue of fixed parameters in existing coating production lines affecting coating production quality. Compared to the prior art, the coating production line control device provided in this application offers the same beneficial effects as the coating production line control method described in the aforementioned embodiments. Other technical features of the coating production line control device are the same as those disclosed in the aforementioned embodiments and are not further elaborated upon here.
[0094] The present application provides a coating production line control device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the coating production line control method in the above-mentioned first embodiment.
[0095] Reference below Figure 4, which shows a schematic diagram of the structure of a coating production line control device suitable for implementing the embodiments of the present application. The coating production line control device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The coating production line control equipment shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.
[0096] like Figure 4 As shown, the coating line control device may include a processing device 1001 (e.g., a central processing unit, graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the coating line control device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, a magnetic tape or hard disk; and a communication device 1009. Communication device 1009 can allow the coating line control device to communicate with other devices wirelessly or wired to exchange data. Although the figure shows a coating line control device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented or provided instead.
[0097] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0098] The coating production line control device provided in this application, utilizing the coating production line control method described in the aforementioned embodiment, can resolve the technical issue of fixed parameters in existing coating production line equipment, which impacts coating production quality. Compared to the prior art, the coating production line control device provided in this application offers the same beneficial effects as the coating production line control method described in the aforementioned embodiment. Other technical features of the coating production line control device are the same as those disclosed in the aforementioned embodiment and are not further elaborated upon here.
[0099] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0100] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0101] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, and the computer-readable program instructions are used to execute the coating production line control method in the above-mentioned embodiment.
[0102] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0103] The computer-readable storage medium may be included in the coating production line control device; or it may exist independently without being assembled into the coating production line control device.
[0104] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the coating production line control equipment, the coating production line control equipment is enabled to: monitor the production environment parameters and equipment operation parameters of the coating production line equipment; adjust the parameters according to the production environment parameters and equipment operation parameters through a pre-trained deep reinforcement learning model to obtain the equipment parameter adjustment amount, and the pre-trained deep reinforcement learning model is obtained by training a preset initial deep reinforcement learning model based on pre-acquired historical coating production data; adjust the equipment operation parameters of the coating production line equipment according to the equipment parameter adjustment amount, and control the coating production line equipment after parameter adjustment to perform coating production.
[0105] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0107] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0108] The computer-readable storage medium provided herein stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned coating production line control method. This computer-readable storage medium can address the technical issue of fixed parameters in existing coating production lines, which impacts coating production quality. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided herein are similar to those of the coating production line control method provided in the aforementioned embodiments, and are not further elaborated here.
[0109] The present application also provides a computer program product, comprising a computer program, which implements the steps of the above-mentioned coating production line control method when executed by a processor.
[0110] The computer program product provided in this application can address the technical issue of fixed parameters in existing coating production lines, which impacts coating production quality. Compared to the prior art, the computer program product provided in this application offers the same beneficial effects as the coating production line control method provided in the aforementioned embodiments, and will not be further elaborated here.
[0111] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A coating production line control method, characterized in that: The coating production line control method includes: Monitor the production environment parameters and equipment operation parameters of coating production line equipment; By using a pre-trained deep reinforcement learning model, parameter adjustment is performed according to the production environment parameters and equipment operating parameters to obtain an equipment parameter adjustment amount, wherein the pre-trained deep reinforcement learning model is obtained by training a preset initial deep reinforcement learning model based on pre-acquired historical coating production data; The equipment operating parameters of the coating production line equipment are adjusted according to the equipment parameter adjustment amount, and the coating production line equipment after the parameter adjustment is controlled to perform coating production.
2. The coating production line control method according to claim 1, wherein: Before the step of adjusting parameters of the pre-trained deep reinforcement learning model according to the production environment parameters and equipment operating parameters, the method further includes: Acquiring historical coating production data, wherein the historical coating production data includes historical production environment data, historical coating control data, and historical coating detection data; Based on the historical production environment data, historical coating control data and historical coating detection data, a preset initial deep reinforcement learning model is trained to obtain a pre-trained deep reinforcement learning model.
3. The coating production line control method according to claim 2, wherein: Before the step of training a preset initial deep reinforcement learning model based on the historical production environment data, the historical coating control data, and the historical coating detection data to obtain a pre-trained deep reinforcement learning model, the method further includes: Defining a state space according to a production environment parameter range of the coating production line equipment; Defining an action space according to a coating control parameter range of the coating production line equipment; Define the reward function based on the preset film layer standard parameters and production target parameters; A model is constructed based on the state space, action space, and reward function to obtain an initial deep reinforcement learning model.
4. The coating production line control method according to claim 2, wherein: The step of training a preset initial deep reinforcement learning model based on the historical production environment data, the historical coating control data, and the historical coating detection data to obtain a pre-trained deep reinforcement learning model includes: Performing time series alignment processing on the historical production environment data, the historical coating control data, and the historical coating detection data to generate initial training samples; Performing data enhancement on the initial training samples by a sliding window method to obtain data-enhanced training samples; The data-augmented training samples are input into the initial deep reinforcement learning model for offline iterative training to obtain a pre-trained deep reinforcement learning model.
5. The coating production line control method according to claim 1, wherein: The production environment parameters include one or more of equipment load, substrate specifications, temperature and pressure, and the equipment operation parameters include one or more of magnetron sputtering power, reaction gas flow and annealing temperature.
6. The coating production line control method according to claim 1, wherein: After the step of adjusting the equipment operating parameters of the coating production line equipment according to the equipment parameter adjustment amount and controlling the coating production line equipment after the parameter adjustment to perform coating production, the method further includes: Performing film layer detection on the glass produced by the coating production line equipment to obtain film layer detection parameters; Calculating a deviation between the film layer detection parameter and a preset film layer standard parameter, and comparing the deviation value with a preset deviation threshold; If the deviation value exceeds a preset deviation threshold, the reward function of the pre-trained deep reinforcement learning model is updated based on the deviation value, and the pre-trained deep reinforcement learning model is retrained.
7. A coating production line control device, characterized in that: The coating production line control device includes: Monitoring module, used to monitor the production environment parameters and equipment operating parameters of the coating production line equipment; an adjustment module, configured to adjust parameters according to the production environment parameters and equipment operating parameters using a pre-trained deep reinforcement learning model to obtain an equipment parameter adjustment amount, wherein the pre-trained deep reinforcement learning model is obtained by training a preset initial deep reinforcement learning model based on pre-acquired historical coating production data; A control module is used to adjust the equipment operating parameters of the coating production line equipment according to the equipment parameter adjustment amount, and control the coating production line equipment after the parameter adjustment to perform coating production.
8. A coating production line control device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the coating production line control method according to any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the coating production line control method according to any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the coating production line control method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
New energy hybrid power supply method and system based on energy unloading and storage medium
CN116581799A
Intelligent electroplating control system and control method
CN117684243A
Method and system for optimizing parameters of vacuum coating equipment
CN118308702A
Fetal heart rate monitoring image classification method and system based on neural network
CN119559451A
Cited By
Production line management method and system for mixed production of konjak
CN121329039A
Particle swarm optimization design method for coating parameters of heat-insulating glass
CN121859754A
A Particle Swarm Optimization Design Method for Thermal Insulation Glass Coating Parameters
CN121859754B