Intelligent color box printing color analysis control method
By collecting multi-source data from the color box printing production line, establishing a multimodal meta-state feature model and constructing a time-series multi-scale reward function, the problem of insufficient adaptability of traditional color box printing color control technology in complex environments is solved, and efficient color consistency control is achieved.
Patent Information
- Application Number
- CN202511255448.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-11-04
AI Technical Summary
Traditional color control technology for color box printing lacks intelligence and adaptability under high speed, multiple batch switching and complex environmental disturbances, making it difficult to achieve color consistency control.
Multi-source color detection data and environmental parameters from the color box printing production line are collected. A multimodal meta-state feature model is established through autoencoder neural network and self-supervised clustering. A time-series multi-scale reward function is constructed, and intelligent decision-making is carried out using a reinforcement learning strategy network to achieve real-time color control.
It improves the spatial, temporal, and operational resolution of color states, enhances the expression of high-frequency color fluctuations and spatial heterogeneity, ensures accurate attribution of reward signals, improves the adaptability and robustness of color control, and achieves full-process automated color consistency.
Smart Images

Figure CN120886567A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of "intelligent control and reinforcement learning adaptive control in color box printing", and more particularly to an intelligent color analysis and control method for color box printing. Background Technology
[0002] In the current color box printing industry, with the acceleration of digitalization and intelligent manufacturing, online inspection of printing quality and intelligent color control have become key links in improving the automation level of production lines and product consistency. Traditional color control technology for color box printing mainly relies on a closed-loop feedback mechanism based on fixed-point detection and manual experience adjustment. In actual production scenarios, color consistency is mainly achieved through periodic manual sampling, rule-based heuristic parameter setting (such as ink concentration, temperature and humidity compensation), or simple statistical models (such as PID control, linear regression). These methods have a certain degree of adaptability to low-frequency variations and process fluctuations, but under modern high-speed, multi-batch switching and complex environmental disturbances, their intelligence and adaptability are clearly insufficient. Summary of the Invention
[0003] This application provides an intelligent color analysis and control method for color box printing, which aims to solve one of the problems or issues of the prior art mentioned in the background art above.
[0004] This application provides an intelligent color analysis and control method for color box printing, specifically including: S1: Collect multi-source color detection data from multiple spatial locations and different time points on the color box printing production line, and simultaneously acquire printing speed, consumable batches and equipment micro-environment parameters to achieve real-time perception of multi-dimensional variables in the production environment.
[0005] S2: Denoise, normalize, and time-synchronize the collected multi-source color detection data and environmental parameters to eliminate high-frequency interference and scale inconsistencies of different variables, and generate a benchmark feature dataset.
[0006] S3: Based on the normalized baseline feature dataset, an autoencoder neural network is used to extract features from color changes, fluctuation amplitude, and short-term trend dimensions to generate a multimodal meta-state feature set for subsequent state space representation enhancement.
[0007] S4: For labels in different spatial regions and printing conditions, self-supervised clustering is performed on the multimodal meta-state feature set to establish a differentiated meta-state distribution model suitable for multivariate disturbances, thereby realizing the classification expression of variable disturbances.
[0008] S5: Based on the differentiated meta-state distribution model, a time-series multi-scale reward function is constructed, which decomposes the reward signal into instantaneous reward, short-segment reward and batch global reward, and assigns dynamic weights to the fluctuation risk level of the meta-state to ensure that high volatility is punished in real time and positive incentives are given when the trend is stable.
[0009] S6: Input the meta-state features into the reinforcement learning policy network, and calculate the optimal color control action based on the reward attribution information of the current meta-state and the historical similar states, so as to realize the intelligent decision mapping output of color parameters and process settings.
[0010] S7: Based on the optimal color control action output by the reinforcement learning strategy network, send color adjustment and process parameter setting instructions to downstream equipment in the color box printing production line, and execute real-time color control operations to achieve target color consistency.
[0011] S8: Perform real-time analysis on the control feedback signals and reward function output generated during the execution process, monitor the stability of reward attribution and the ability to represent the meta-state, and determine whether there are reward drift or state feature degradation problems.
[0012] S9: If reward drift or state feature degradation is detected, the meta-state aggregator is triggered to perform adaptive fine-tuning or re-clustering operations, dynamically optimize the meta-state space representation, and update the reinforcement learning policy network structure to maintain the stability and convergence efficiency of the reward feedback loop.
[0013] The intelligent color analysis and control method for color box printing provided in this application has the following beneficial effects: (1) By deeply fusing multi-source color data, process parameters, and environmental variables, and combining autoencoder networks and self-supervised clustering, this method transforms the original data into high-dimensional multimodal meta-state features, effectively improving the spatial, temporal, and operational resolution of color states. Actual tests show that operational condition discrimination capability is improved by 15%-22% after state clustering, and the detection sensitivity for spatial / batch anomalies is significantly enhanced. In highly dynamic production environments, it provides a more accurate representation of high-frequency color fluctuations and spatial heterogeneity, ensuring accurate attribution of reward signals.
[0014] (2) This invention employs a multi-scale reward function to monitor color control effects hierarchically from single images and short segments to full batches. It combines the risk level of the original state with adaptive weighting to achieve strong penalties for abnormally high fluctuations and dynamic incentives for stable trends. The multi-scale attribution mechanism avoids the loss sensitivity and drift distortion of single-time-series rewards in industrial fluctuation scenarios. The stability of the feedback loop is improved by more than 30% compared to traditional methods, the reward signal's consistency with actual working conditions is enhanced, and the convergence speed is significantly improved.
[0015] (3) Meta-state driven reinforcement learning policy networks can effectively integrate historical reward attribution and current state, and make deep decisions using high-dimensional temporal context features, effectively avoiding policy gaps caused by state space degradation. Through adaptive clustering and online network fine-tuning, the accuracy of actual color control decisions is improved by 12%-18%, and it has higher response speed and control robustness to rapid operating condition switching or batch disturbances.
[0016] (4) The state representation and reward attribution mechanism of this invention is universal and can be adapted to different types of printing equipment, multiple process batches, and extreme environmental disturbances. In multi-batch printing and heterogeneous consumable scenarios, the system can continuously learn online and adapt parameters without manual intervention, achieving fully automated color consistency control. After technology transfer, it can be extended to multiple fields such as packaging, labeling, and commercial printing, greatly improving the level of intelligence in industrial printing.
[0017] In summary, this invention achieves comprehensive innovation in data perception, state modeling, reward closed loop and policy network, systematically solves the problem of state reward instability in color control of high dynamic production lines, significantly improves the adaptability, efficiency and practical application value of intelligent printing control, and has broad prospects for promotion in the field of industrial intelligent manufacturing. Attached Figure Description
[0018] Appendix Figure 1 This is the main flowchart of the intelligent color analysis and control method for color box printing.
[0019] Appendix Figure 2 This is a sub-flowchart of the intelligent color analysis and control method for color box printing.
[0020] Appendix Figure 3 This is another sub-flowchart of the intelligent color analysis and control method for color box printing. Detailed Implementation
[0021] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0022] The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0023] As attached Figure 1 As shown, this application provides an intelligent color analysis and control method for color box printing, specifically including: S1: Collect multi-source color detection data from multiple spatial locations and different time points on the color box printing production line, and simultaneously acquire printing speed, consumable batches and equipment micro-environment parameters to achieve real-time perception of multi-dimensional variables in the production environment.
[0024] S2: Denoise, normalize, and time-synchronize the collected multi-source color detection data and environmental parameters to eliminate high-frequency interference and scale inconsistencies of different variables, and generate a benchmark feature dataset.
[0025] S3: Based on the normalized baseline feature dataset, an autoencoder neural network is used to extract features from color changes, fluctuation amplitude, and short-term trend dimensions to generate a multimodal meta-state feature set for subsequent state space representation enhancement.
[0026] S4: For labels in different spatial regions and printing conditions, self-supervised clustering is performed on the multimodal meta-state feature set to establish a differentiated meta-state distribution model suitable for multivariate disturbances, thereby realizing the classification expression of variable disturbances.
[0027] S5: Based on the differentiated meta-state distribution model, a time-series multi-scale reward function is constructed, which decomposes the reward signal into instantaneous reward, short-segment reward and batch global reward, and assigns dynamic weights to the fluctuation risk level of the meta-state to ensure that high volatility is punished in real time and positive incentives are given when the trend is stable.
[0028] S6: Input the meta-state features into the reinforcement learning policy network, and calculate the optimal color control action based on the reward attribution information of the current meta-state and the historical similar states, so as to realize the intelligent decision mapping output of color parameters and process settings.
[0029] S7: Based on the optimal color control action output by the reinforcement learning strategy network, send color adjustment and process parameter setting instructions to downstream equipment in the color box printing production line, and execute real-time color control operations to achieve target color consistency.
[0030] S8: Perform real-time analysis on the control feedback signals and reward function output generated during the execution process, monitor the stability of reward attribution and the ability to represent the meta-state, and determine whether there are reward drift or state feature degradation problems.
[0031] S9: If reward drift or state feature degradation is detected, the meta-state aggregator is triggered to perform adaptive fine-tuning or re-clustering operations, dynamically optimize the meta-state space representation, and update the reinforcement learning policy network structure to maintain the stability and convergence efficiency of the reward feedback loop.
[0032] Step S1: Collect multi-source color detection data from multiple spatial locations and different time points on the color box printing production line, and simultaneously acquire printing speed, consumable batches, and equipment micro-environment parameters to achieve real-time perception of multi-dimensional variables in the production environment. Specifically, this includes: S1.1: Based on distributed color detection sensors, real-time color detection data of printed surfaces are collected at key spatial locations in the printing production line to obtain color data distribution under spatial heterogeneity and generate multi-channel raw color detection signals.
[0033] The input conditions include a distributed array of color detection sensors deployed at key locations on the color box printing production line, which performs real-time color data detection on the surface of printed materials at different spatial locations.
[0034] A multi-channel real-time color detection method is adopted (parameters: each channel sensor covers a specified spatial location, and the acquisition period can be configured to 10ms~100ms) to obtain the color distribution information of printed materials at different spatial points on the production line.
[0035] Furthermore, by employing a multi-channel parallel acquisition mechanism and utilizing a synchronous data acquisition protocol (parameter: global clock alignment error <1ms), time synchronization between each color detection channel is achieved, ensuring the temporal consistency of spatially heterogeneous data.
[0036] Using a spectral or tristimulus value color sensor (parameters: wavelength range 400nm~700nm, measurement accuracy ΔE<0.5, sampling rate ≥100Hz), high-resolution original surface color detection signals are obtained, and the output is a multi-channel color difference, reflectance or RGB / XYZ value sequence.
[0037] Furthermore, by using spatial numbering and device identification information, the collected color signal data is structured and encoded according to spatial nodes, and archived into a raw detection dataset containing metadata such as collection time, spatial location, and device number.
[0038] Through the acquisition and structured processing of multi-channel spatial heterogeneous color detection signals, a comprehensive perception and expression of the original color data under the spatial distribution state of the printing production line is achieved, laying the foundation for monitoring time-series high-frequency dynamic color changes and extracting meta-state features.
[0039] For example, on a 30-meter-long color box printing production line, 24 sets of color detection sensors are deployed, each covering a 1.25-meter printing area. Each sensor is a hyperspectral color sensor (spectral resolution of 2nm, wavelength range of 400-700nm, acquisition period of 20ms). All sensors are controlled by a PLC main controller which distributes acquisition commands, achieving synchronous triggering and data transmission via an Ethernet bus to ensure that the time error of the color signals acquired at all spatial points is within 0.5ms. The acquired data includes the sampling timestamp, sensor number, spatial location number, and corresponding spectral reflectance data for each detection. In each time period, the system aggregates 24,000 color feature data points from 24 channels, which are then spatially structured and encoded to generate the original color distribution dataset. In practical applications, during long-term acquisition experiments, this solution can monitor color consistency changes and spatial anomalies in various spatial segments with high confidence, supporting subsequent high-dimensional state model construction and heterogeneity identification, and improving the spatial resolution and accuracy of state representation and reward attribution under dynamic conditions.
[0040] S1.2: For a color detection sensor at any spatial location, a high-frequency sampling mechanism is set to collect dynamically changing color detection data at multiple time scales, forming a multi-time point time-series color detection data stream, thus achieving a data foundation with high temporal resolution.
[0041] S1.3: Synchronously collect printing speed parameters from the production line control system to obtain the printing speed process variables corresponding to the sampling time of each color detection data, so as to achieve spatiotemporal alignment between process parameters and color data.
[0042] S1.4: In conjunction with the production material handling system, obtain the corresponding consumable batch identifier from the spatial location and sampling time point of the color detection data collection, forming a data pair that associates color detection data with consumable batch variables to support subsequent variable association modeling.
[0043] S1.5: Employs a variety of industrial environmental sensors to collect equipment micro-environment parameters, including but not limited to temperature and humidity, machine surface temperature, and airflow, to achieve spatiotemporal consistency synchronization with color detection data, printing speed, and consumable batches, generating a multi-dimensional parallel advanced parameter set.
[0044] The input conditions include signals collected by multi-source color detection devices that have been deployed synchronously on the color box printing production line, as well as process variable information such as printing speed and consumable batches that are aligned with the timestamp. Equipment micro-environment parameters need to be introduced to form a full input of multi-dimensional variable perception.
[0045] Industrial-grade temperature and humidity sensors (parameters: temperature measurement range 0℃~60℃, resolution 0.1℃; humidity measurement range 0%RH~100%RH, resolution 0.1%RH) are deployed in key locations on the printing press to collect real-time temperature and humidity data streams of the local environment, enabling quantitative monitoring of the impact of environmental changes on color performance.
[0046] Furthermore, infrared thermocouples or surface thermistors (parameters: measurement range 0℃~150℃, response time ≤100ms) are used to collect machine surface temperature data to reflect the influence of thermal radiation and structural thermal inertia on ink drying and color stability during equipment operation.
[0047] Furthermore, a hot-film airflow sensor (parameters: speed measurement range 0~10m / s, accuracy ±0.05m / s) was used to monitor the airflow velocity in the printing channel and drying unit, and to evaluate the interference effect of airflow conditions on ink evaporation and curing speed.
[0048] Furthermore, through a unified industrial data acquisition module (supporting multiple analog signals and RS485 / Modbus communication interfaces), the output signals of sensors such as temperature and humidity, machine surface temperature, and air flow are timestamped and spatially encoded to ensure that the acquired data corresponds one-to-one with color detection data, printing speed, and consumable batch variables under the same time base and spatial index.
[0049] Through the above-mentioned parallel acquisition and unified encoding processing of multiple sensors, the set of multi-dimensional parallel high-level parameters such as temperature and humidity, machine surface temperature, and air flow is output in the form of a structured matrix, providing complete equipment micro-environment input features for subsequent multimodal data fusion, and realizing the characterization of environmental interference profiles with high-frequency dynamic color changes.
[0050] For example, on a color box printing production line equipped with a four-color offset printing unit and an online varnishing machine, two sets of temperature and humidity sensors, one set of machine surface temperature sensors, and one set of air flow sensors are installed in each of the pre-printing, mid-printing, post-printing, and varnishing sections, for a total of eight sets of temperature and humidity sensors, four sets of surface temperature sensors, and four sets of air flow sensors. The sampling period for the temperature and humidity sensors is set to 1 second, the sampling period for the machine surface temperature sensors is set to 500 ms, and the sampling period for the air flow sensors is set to 200 ms. All sensors are connected to the same PLC through a multi-channel data acquisition module. The PLC timestamps all channel data (time error <1 ms) and adds a spatial location number, consistent with the spatial number of the color detection channel. Single-cycle data collection and summarization generate a spatial matrix of temperature and humidity, machine temperature, and air flow rate. Each matrix contains 16 temperature and humidity values, 4 surface temperature values, and 4 flow rate values. As the printing speed increased from 8,000 sheets per minute to 12,000 sheets per minute, the airflow velocity in the varnishing section increased from 2.5 m / s to 4.0 m / s, the machine surface temperature rose by 2°C, and the humidity decreased by 3% RH. The unexpected fluctuations in the corresponding color detection data were precisely correlated with these environmental variable changes, enabling traceable analysis of potential environmental factors contributing to color fluctuations on the production line. This combined parameter set significantly improved the completeness of state representation and the stability of reward function attribution in subsequent multimodal data fusion and meta-state construction.
[0051] S1.6: Based on timestamps and spatial numbers, multimodal data fusion matching is performed on all collected raw color detection signals, printing speed parameters, consumable batch identifiers, and equipment microenvironment parameters to construct a complementary and consistent multi-source raw dataset, providing a unified input basis for subsequent feature extraction and meta-state generation.
[0052] Step S2 involves preprocessing the collected multi-source color detection data and environmental parameters by denoising, normalizing, and temporally synchronizing them to eliminate high-frequency interference and scale inconsistencies among different variables, and to generate a baseline feature dataset. Specifically, this includes: S2.1: Perform signal denoising methods such as bandpass filtering and mean smoothing on the multi-source color detection data and environmental parameters (such as printing speed, consumable batches, and equipment micro-environment information) collected from the color box printing production line to effectively filter high-frequency noise and instrument errors, and obtain preliminary processed data after noise reduction.
[0053] For multi-source color detection data and environmental parameters (including printing speed, consumable batches, and equipment micro-environment information) collected from the color box printing production line, a bandpass filtering algorithm is used (parameter: passband frequency range). to filter order This achieves synchronous suppression of low-frequency drift and high-frequency random noise in the signal, while preserving the main color fluctuation components to maintain the integrity of state characteristics.
[0054] Furthermore, the mean smoothing method (parameter: sliding window length) is used. sampling points, step size (Sampling points) to smooth the remaining spike impulse noise after bandpass filtering, and obtain smoothed color detection and environmental parameter sequence data.
[0055] Furthermore, a weighted mean smoothing algorithm (weight distribution) is adopted. Based on the symmetrical distribution of the window center position and satisfying This enables the accurate denoising of subtle fluctuations in color detection signals, resulting in more continuous color transitions and generating preliminary processed color data with a high signal-to-noise ratio.
[0056] Furthermore, through the instrument calibration compensation model (based on the historical error matrix) With real-time detection residual Linear correction formula It compensates for systematic deviations of the sensor across the entire measurement range, enabling multi-source consistency signal correction for color detection and environmental variables.
[0057] Through the above chain-like processing flow, the original noisy signals from multiple sources are transformed into a preliminary processed dataset with reduced noise, effectively suppressing high-frequency transient interference and equipment system deviations, and laying a high-quality data foundation for subsequent normalization and timing synchronization.
[0058] For example, on a color box printing production line, the color sensor sampling rate is 100Hz. The recorded color reflectance signal contains a 1-2Hz mechanical vibration background and an electronic interference spike >40Hz. For this scenario, a bandpass filter with a passband of 0.5Hz to 30Hz and a 4th-order Butterworth filter type was used. The resulting signal noise power decreased by approximately 35%. Subsequently, a sliding weighted mean smoothing with 11 points (weights distributed Gaussian) was applied. (Settings), further reducing transient spike amplitude by approximately 60%. This was further achieved through a two-month historical color and temperature / humidity calibration experiment. The matrix was used to perform system bias compensation on the real-time data, reducing the root mean square error (RMSE) of color difference from 0.48 to 0.21. The convergence speed of the final output pre-processed data in subsequent normalization processing was improved by 18%, and the stability of multimodal state feature extraction was significantly enhanced.
[0059] S2.2: Based on the preliminary processed data obtained in step S2.1, the standard deviation normalization method is used to transform the color detection data and environmental parameters to a zero-mean unit variance distribution to eliminate the differences in physical meaning and dimensions between different modalities, thus forming a standard normalized dataset.
[0060] S2.3: For the standard normalized dataset obtained in step S2.2, apply a time-series alignment algorithm (such as Dynamic Time Warping (DTW)) to synchronize the sampling time series of different detection points and multivariate data streams in space, correct the sampling delay and the difference in acquisition frequency of different data sources, and generate synchronized processing data with time-series consistency.
[0061] S2.4: Perform missing value detection and interpolation completion processing on the time-series consistency synchronization processing data formed in step S2.3. Combine the sampling of previous and subsequent time steps with spatial nearest neighbor information and apply linear interpolation or high-dimensional spline interpolation to fill in abnormal missing segments and obtain complete baseline time-series feature data.
[0062] S2.5: Integrate the integrity benchmark time series feature data obtained in step S2.4, and construct the final standardized benchmark feature dataset by structurally splicing different detection channels and environmental variables according to the preset feature vector arrangement standard, so as to provide unified and standardized data input for subsequent feature extraction and meta-state modeling.
[0063] Step S3: Based on the normalized baseline feature dataset, an autoencoder neural network is used to extract features from color changes, fluctuation amplitude, and short-term trend dimensions to generate a multimodal meta-state feature set for subsequent state space representation enhancement. For example... Figure 2 As shown, it specifically includes: S3.1: Window segmentation is performed on the color detection time series and spatial location information in the normalized benchmark feature dataset to generate a sample batch with spatiotemporal local correlation, in order to capture local state patterns in high-frequency dynamic changes.
[0064] S3.2: Based on the windowed sample batch, an autoencoder neural network is used to perform end-to-end reconstruction learning of the original time series features of color changes, so as to adaptively extract the latent space code containing the principal components of color, the amplitude of change and the fluctuation pattern as the basic meta-state feature vector.
[0065] The input data includes the windowed sample batch data generated in step S3.1. This data consists of a standardized color time series after normalization and time synchronization processing and the corresponding spatial location information, and has spatiotemporal local correlation.
[0066] The feature extraction method using an autoencoder neural network (parameter: input layer dimension) is employed. Equals the number of window sample points multiplied by the dimension of a single feature point; encoder hidden layer nodes By sequentially decreasing the activation function ReLU, an end-to-end reconstruction learning of the windowed color time series is achieved, and a compressed latent space representation is obtained at the end of the encoder.
[0067] Furthermore, the objective function is minimized by means of the mean square reconstruction error. Train the network to achieve the consistency constraint between the latent space and the original input in the reconstruction: , in, For the first One original window sample, The reconstructed sample is the output of the decoder. This represents the number of samples.
[0068] Furthermore, batch normalization (parameter: momentum coefficient) is introduced into the encoder structure. Numerical stability parameters This stabilizes the input distribution across batches and reduces the variance of parameter updates during network training.
[0069] Furthermore, by applying regularization constraints at the latent space layer... This suppresses the amplification of invalid high-frequency components in the encoding vector, thereby improving the generalization ability and noise resistance of the basic meta-state feature vector.
[0070] Furthermore, by adding fully connected projection units (parameter: output dimension) at the end of the encoding layer of the autoencoder... Based on a cumulative explained variance rate > 95%, the high-dimensional encoding mapping is compressed into the basic meta-state feature vector, thereby achieving a centralized expression of the principal components of color, the magnitude of change, and the fluctuation pattern.
[0071] By using an end-to-end reconstruction learning method with an autoencoder neural network, the windowed original color time series is transformed into a stable, low-dimensional basic meta-state feature vector with complete color dynamic feature information, thereby significantly improving the expressive power of the state space.
[0072] For example, in a color data processing task for a color box printing production line, the window length is set to 50 sampling points (corresponding to a 0.5-second acquisition time), and each point has a feature dimension of 4 (chromaticity L*, a*, b* and local illumination intensity), with an input layer dimension of... A three-layer encoder structure is adopted, with 128, 64, and 32 hidden layer nodes respectively. The decoder is symmetrically reversed, the activation function is ReLU for all layers, and the output layer is a linear function. The training sample size is 100,000 window samples, the batch size is 256, the optimizer is Adam, and the learning rate is... The reconstruction error of the validation set after 50 training rounds Stable below 0.015. Applying a hidden space layer Regularization constraint, regularization coefficient is The batch normalized momentum coefficient is set to 0.9. The dimension of the basic meta-state feature vector after dimensionality reduction. The explained variance reached 96.2%. In subsequent multi-scale aggregation calculations, this feature vector improved the clustering profile coefficient by 24% on a small sample test set compared to the original time-series data. Through dynamic attribution evaluation using the reward function, the accuracy of fluctuation risk detection was improved by 21%, significantly enhancing the adaptability and stability of the reinforcement learning control system under dynamic operating conditions.
[0073] S3.3: Apply multi-scale aggregation operation to the basic meta-state feature vector output by the autoencoder neural network to calculate statistical features such as short-term window average, variance, and fluctuation range amplitude, so as to highlight the dynamic substructure of short-term fluctuations and abnormal changes.
[0074] S3.4: Cascade and combine multi-scale statistical features with original spatial location information and equipment operating condition labels to form a multimodal meta-state feature set, which systematically expresses the color change status under different physical spaces, process environments and statistical levels.
[0075] The input data includes standardized color sequence features after normalization and time-synchronized preprocessing, spatial location information, and synchronously acquired equipment condition labels.
[0076] The feature concatenation algorithm (parameters: multi-scale statistical features, spatial attribution, and working condition type) is used to group and associate the multi-scale statistical features (such as local window mean, variance, and extreme value interval) obtained in step S3.3 according to their corresponding spatial detection point numbers. Each group of statistical features is paired with the original spatial location information and concatenated into the same feature vector.
[0077] Furthermore, using a tag encoding method (parameters: printing speed category, consumable batch ID, equipment micro-environment tag), for the statistical characteristics of each spatial location, synchronously collected equipment condition tags are retrieved and appended. These tags are mapped into numerical features in the form of one-hot encoding or embedding vectors, and then concatenated with statistical features and spatial location information to form an expanded feature representation.
[0078] Furthermore, a high-dimensional feature concatenation operation is employed, using a vector concatenation function. statistical characteristics Spatial location information and working condition label features Integrate them into a new set of multimodal meta-state feature vectors: , in, It is a multi-scale statistical characteristic matrix. For spatial numbering vectors, Embed vectors for operating condition labels.
[0079] Furthermore, by sorting and numbering structured data, features from different sources are arranged in an orderly manner according to the sampling time and spatial order, preventing spatiotemporal mismatch from introducing label mixing, and providing an orderly and high-confidence underlying input for subsequent clustering and reward attribution.
[0080] By combining feature cascading and data structuring, multi-scale statistical features, spatial information, and operating condition indicators are combined and assembled into a multimodal meta-state feature set, enabling a systematic expression of color change states under different physical spaces, process environments, and statistical levels on the color box printing production line.
[0081] For example, color detection sequence data that has undergone normalization and temporal synchronization processing is used on an automated printing line. A 10-point local sampling window is employed to extract statistical features such as the window mean, standard deviation, and maximum and minimum color difference, resulting in a 64-dimensional statistical feature vector for each detection point. Spatial location information is encoded using 16-bit spatial numbers. Operating condition labels include printing speed (three levels), consumable batch (six types of numbers), and equipment surface temperature (three levels), each converted into a total of 12-dimensional operating condition label vectors using One-Hot encoding. At each sampling moment, the 64-dimensional statistical features, 16-dimensional spatial numbers, and 12-dimensional operating condition labels are concatenated to form a 92-dimensional multimodal meta-state feature vector. When this feature set is batch-fed into subsequent self-supervised clustering algorithms and reward attribution modules, it achieves a multi-level differential expression of color dynamic states under different spatial locations, consumable batches, and all equipment operating conditions. Practical application results show that after feature concatenation, the inter-class discriminativeness of subsequent clustering is improved by 22%, and the reward attribution accuracy is improved by 15%, providing a high-confidence and systematic data foundation for enhancing the accuracy and adaptability of the learning intelligent decision-making system.
[0082] S3.5: Perform feature regularization and noise reduction on the multimodal meta-state feature set. Use regularization constraints and decorrelation algorithms to improve the stability and discriminativeness of the meta-state feature representation, and provide a high-confidence feature basis for subsequent self-supervised clustering and reinforcement learning state space expansion.
[0083] Step S4: For labels in different spatial regions and printing conditions, self-supervised clustering is performed on the multimodal meta-state feature set to establish a differentiated meta-state distribution model suitable for multivariate disturbances, thereby achieving a classification representation of variable perturbations. For example... Figure 3 As shown, it specifically includes: S4.1: Based on the multimodal meta-state feature set generated in the previous steps, the meta-state features in each region are extracted and archived according to the established spatial region partitioning standard, forming a subset of meta-state features grouped by spatial location, so as to achieve the preliminary classification and division of multimodal meta-state features in the spatial dimension.
[0084] The input data includes a set of multimodal meta-state features after feature concatenation and regularization, specifically including spatial numbering, working condition labels and multi-scale color statistical features. The input source is the output of the preceding data normalization, windowing and feature concatenation process.
[0085] Spatial region indexing rules (parameters: physical detection point number range, spatial discrete grid partition scale, etc.) are used to perform spatial location encoding matching on the multimodal meta-state feature set, mapping each meta-state feature vector to the corresponding physical spatial region according to the spatial number field.
[0086] Furthermore, through the regional archiving algorithm (parameters: spatial number matching rules, batch processing cache threshold), the meta-state features mapped to the same spatial region are batched and aggregated to form a subset of meta-state features with spatial blocks as units, ensuring that samples within the same region have spatiotemporal consistency in subsequent clustering and discriminant analysis.
[0087] Furthermore, a structured data sorting technique (parameters: ascending spatial number, ascending sampling time) is adopted to serialize and arrange the archived meta-state feature subsets within each spatial region, ensuring the orderly consistency of the data index and preventing confusion between features and physical locations.
[0088] Furthermore, a hierarchical storage method is adopted (parameters: spatial classification level, batch number identifier) to store the meta-state feature subset of spatial archives in the form of partitioned files or memory queues, providing efficient access and fast indexing capabilities for subsequent partitioned parallel clustering and spatial dynamic map analysis.
[0089] Through algorithms such as spatial region mapping, region archiving, and structured arrangement, the multimodal meta-state feature set generated in the previous steps is transformed into a subset of meta-state features grouped by spatial location. This achieves preliminary and refined classification of meta-state multimodal features in the spatial dimension, significantly improving the discriminativeness and discriminativeness of meta-state modeling under subsequent multivariate interference.
[0090] For example, for 32 independent spatial detection points on an automated color box printing line, the input feature set contains 92-dimensional multimodal meta-state features collected per point per cycle. Using spatial numbering interval configuration, the 32 detection points are divided into 4 spatial regions based on physical distance, with each region containing 8 independent detection points. Through spatial numbering mapping, the meta-state features of each sampling cycle are assigned to the corresponding spatial region according to the detection point number, achieving batch aggregation of region-level feature subsets. Within each region, features are arranged in ascending order using timestamps. Each processing iteration generates 4 spatial region feature batches, each containing an N×8-dimensional feature vector data block (N being the number of sampling cycles). In practical applications, after spatial grouping, the processing speed of subsequent region adaptive clustering algorithms is improved by approximately 36%, and the inter-class feature distribution differences are improved by 18%, strongly supporting the development of spatially sensitive control strategies and the optimization of anomaly detection performance. The final output of the spatial grouping process is a subset of meta-state features from 4 spatial regions, providing a high-quality hierarchical input foundation for subsequent differentiated feature expression, clustering attribution, and reward function design under multiple regions and operating conditions.
[0091] S4.2: For each subset of meta-state features in a spatial region, combine the printing condition labels collected in actual production (including but not limited to printing speed category, consumable batch ID, equipment micro-environment classification labels, etc.) to perform label matching and attribution processing, and construct meta-state-condition mapping data pairs with diverse condition features to achieve structured organization of meta-state features under both spatial and condition conditions.
[0092] S4.3: For the structured meta-state feature data pairs, a self-supervised clustering algorithm (such as the clustering layer in an autoencoder neural network, adaptive deep embedding clustering, etc.) is used to perform feature space clustering on the meta-state features under the same working conditions in the same region, so as to obtain the typical meta-state cluster centers and corresponding cluster labels that reflect the disturbance of multiple source variables, and realize the high-order aggregate expression of meta-state features.
[0093] S4.4: For the meta-state clustering results under each region-working condition combination, calculate the feature consistency metrics within the cluster (such as the average distance within the cluster, silhouette coefficient, etc.) to evaluate the sensitivity of the clustering model to multivariate disturbances and the classification robustness, thereby selecting meta-state distribution models with high discriminative power and achieving accurate reflection of variable disturbance characteristics.
[0094] S4.5: Integrate all spatial regions and working condition cluster centers, and use discriminative optimization algorithms (such as maximum inter-class distance adjustment, feature reweighting, etc.) to align and merge the meta-state cluster centers across regions and working conditions, and finally form a globally applicable differentiated meta-state distribution model suitable for multi-variable interference scenarios. The model parameters are stored for downstream reward function construction and policy network input.
[0095] Step S5: Based on the differentiated meta-state distribution model, a time-series multi-scale reward function is constructed, decomposing the reward signal into instantaneous reward, short-segment reward, and batch global reward. Dynamic weights are assigned to the volatility risk level of the meta-state to ensure immediate punishment during high volatility and positive incentives during stable trends. Specifically, this includes: S5.1: Based on the established differentiated meta-state distribution model, the meta-state feature sets of each region and time period of the color box printing production line are analyzed to obtain the meta-state sequence, which serves as the input basis for the construction of the reward function, thereby realizing the fusion of meta-state information in the temporal and spatial dimensions.
[0096] S5.2: For the meta-state sequence, a sliding window mechanism is adopted, and a series of sampling strategies are configured for different time scales (such as single sheet, segment, batch) to separate the meta-state sub-sequences at three time-series levels: instantaneous, short segment, and batch global, laying the foundation for structured input conditions for subsequent multi-scale reward signal generation.
[0097] The input is the meta-state sequence obtained by parsing the differentiated meta-state distribution model. The sequence dimension covers the meta-state feature vector sorting results of multiple spatial regions and multiple time periods of the printing production line.
[0098] A sliding window segmentation method (parameters: window length W1=1, step size S1=1, used for instantaneous levels) is adopted to extract the meta-state of a single printing cycle, and the time series is structured into instantaneous meta-state subsequences at the granularity of adjacent single samples.
[0099] Furthermore, by using the sliding window segmentation method (parameters: window length W2=L, where L is the short segment time scale, such as L=10-20, step size S2=M, M≤L), the meta-states within a fixed-length short segment are grouped to generate short segment meta-state subsequences, which are used to capture color fluctuations and trend information within a short period of time.
[0100] Furthermore, a periodic segmentation method (parameters: window length W3=N, N is the batch global interval, such as N=100-300, step size S3=N) is adopted to segment the entire process meta-state sequence at the batch granularity, extract the batch global meta-state sub-sequence, and reflect the color consistency and execution quality of the large interval.
[0101] By flexibly configuring the sliding window parameters and segmentation methods, comprehensive coverage and separation of the same meta-state sequence can be achieved at three time scales: instantaneous, short segment, and batch global.
[0102] A hierarchical labeling algorithm is used to map and label the meta-state subsequences obtained from separation at different scales with corresponding temporal flags, thereby realizing hierarchical indexing and differentiation of structured input.
[0103] Through the above-mentioned multi-scale separation, labeling, and structured indexing process, the original meta-state sequence is mapped into three types of hierarchical meta-state sub-sequences, providing a high-resolution hierarchical input basis for the generation of reward signals at various time scales and risk attribution modules, thereby realizing the structured perception capability of the color control system for the multi-scale state representation of the dynamic production line.
[0104] For example, for a color box printing production line with 64 detection points and 1200 frames of data collected per minute, multi-scale separation of the meta-state feature stream of a 24-hour production cycle is performed. At the instantaneous level, a window length W1=1 is set, and each collected frame of data yields an instantaneous meta-state subsequence. At the short segment level, a window length W2=15 and a step size S2=5 are used to group and slide through 15 consecutive frames of data, generating highly overlapping short segment subsequences to facilitate the identification of frequent fluctuations and sudden anomalies. At the batch global level, a window length W3=3000 and a step size S3=3000 are used, grouping every 3000 frames (approximately 2.5 minutes) into a batch to obtain a batch global meta-state subsequence. For the above multi-scale grouping, hierarchical labeling (e.g., level=1 / 2 / 3 forinstant / segment / global) is used, and each subsequence is bound to its start and end time, spatial region, and operating condition label to achieve structured data indexing. In practical applications, this step enables the accurate separation of different meta-state input levels with instantaneous response capabilities, short-term trend capture capabilities, and global consistency analysis capabilities in high-frequency dynamic environments, effectively constructing a foundation for multi-scale reward signal input. The overlapping method of the sliding window in batch processing allows meta-states at the same moment to participate in reward evaluation at different time-series levels simultaneously, significantly improving the sensitivity and robustness of reward signal generation.
[0105] S5.3: Based on the instantaneous, short segment and batch global meta-state sub-sequences obtained by separation, color deviation measurement algorithms (such as those based on the CIEDE2000 color difference formula), fluctuation risk criteria and batch stability indicators are applied respectively to calculate the color performance anomaly factors at each time scale, forming a set of color risk dependent variable indicators at each level.
[0106] S5.4: For color risk dependent variable indicators at various time scales, design a reward attribution mapping function to map high-risk meta-states to negative instantaneous penalties and stable trend segments to medium or positive incentives. At the same time, introduce batch global consistency reward items to dynamically combine instantaneous rewards, short segment rewards and global rewards to realize the generation of multi-scale reward signals.
[0107] For color risk dependent variable indicators at various time scales, a hierarchical nonlinear mapping function design method is adopted (parameters: the input dimension of risk factors is 3, and the output reward range is set to [-1,1]) to achieve accurate attribution mapping between risk level and reward signal.
[0108] Furthermore, a threshold segmentation mapping algorithm (parameter: high-risk threshold) is used. medium risk threshold Instantaneous color risk factor With the corresponding penalty coefficient Binding, satisfying The maximum negative instantaneous penalty is applied at that time. At that time, a moderate negative incentive was given. Apply positive rewards at appropriate times.
[0109] Furthermore, a trend stability weighted algorithm is adopted (parameter: upper limit of volatility variance). Stability threshold For short-segment risk factors Perform stability testing when In this case, add a trend stability bonus to the moderate positive reward benchmark, or reduce the reward magnitude to capture the stable trend of color fluctuations in the short term.
[0110] Furthermore, a batch global consistency metric is introduced. (Based on batch average color difference) Color difference from target The difference is used to generate a globally consistent reward term through an exponential decay function: in, The attenuation coefficient ensures that the closer the batch color difference is to the target value, the higher the global reward.
[0111] Furthermore, a weighted dynamic synthesis function is adopted to incorporate instantaneous rewards. Short section bonus and global rewards The combination forms a multi-scale integrated reward signal: in The weights are dynamically adjusted by the subsequent S5.5 adaptive algorithm to ensure the sensitivity and stability of the overall reward signal under different operating conditions.
[0112] By using a reward attribution mapping function, the risk indicators at various time scales generated in the previous steps are transformed into hierarchical, multi-channel reward inputs that can be directly used to optimize reinforcement learning strategies, thereby enabling rapid punishment of high-risk states and real-time positive incentives for stable trends.
[0113] For example, on a high-speed color box printing production line with 64 inspection points, the instantaneous color risk factor is calculated based on the CIEDE2000 color difference formula, obtaining... The distribution interval is [0,1]. (Setting...) , ,when Instantaneous punishment ;when hour, ;when hour, In short-segment fluctuation variance detection, let... Stability threshold When the variance of a certain fluctuation segment is 0.008, the short-segment reward... The value increased from +0.5 to +0.7. This is in line with the batch global consistency metric. , , hour, get .
[0114] During the synthesis phase, initial weights are set. Receive comprehensive rewards This result serves as the immediate feedback input to the policy network, causing the network to tend to adjust the process to reduce instantaneous risks under this operating condition, while maintaining a balance between short-term and global trends.
[0115] S5.5: Based on the fluctuation risk level of each meta-state within the meta-state distribution, the weight adaptive algorithm is used to dynamically adjust the reward weights of each level in the multi-scale reward function, thereby suppressing the reward of high-risk states and incentivizing low-risk states, and outputting the final comprehensive multi-scale reward signal, providing a hierarchical and dynamic feedback closed-loop signal for the reinforcement learning decision network in real time.
[0116] The input is the instantaneous reward obtained after processing by S5.4. Short section bonus and global rewards The initial value set, and the volatility risk level corresponding to each meta-state in the global differentiated meta-state distribution model generated by S4.
[0117] Adaptive weight algorithm (parameter: initial weights) Risk sensitivity coefficient This enables dynamic adjustment of the reward weights at each level in the multi-scale reward function.
[0118] Furthermore, the risk level of meta-state fluctuations is determined through a risk weight mapping function. Mapped to weight correction factor This satisfies the criteria of negative correction for high-risk levels and positive correction for low-risk levels: , in, Indicates the first The risk normalization value corresponding to the level, ranging from [0,1].
[0119] Furthermore, a weight normalization constraint algorithm is used to adjust the corrected weight vector. Perform nonnegation and normalization to ensure that the weights at each level satisfy the following conditions. To maintain consistency between the physical meaning and relative proportion of the multi-scale reward structure.
[0120] Furthermore, a dynamic smoothing filter (parameter: smoothing coefficient) is used. For the weight change sequence at consecutive time points Time smoothing is performed to reduce the drastic weight fluctuations caused by instantaneous disturbances in operating conditions and improve the stability of the weight adaptive mechanism.
[0121] Furthermore, the smoothed weight vector Formula for combining return rewards: , Output a comprehensive multi-scale reward signal sequence It serves as the immediate hierarchical feedback closed-loop input for reinforcement learning decision networks.
[0122] Through the above-mentioned adaptive adjustment and comprehensive processing of weights, the reward signals of each scale in the previous step are transformed into a comprehensive feedback quantity that dynamically responds to the fluctuation characteristics of production conditions, thereby achieving immediate suppression of high-risk states and effective incentives for low-risk states.
[0123] For example, on a high-speed color box printing line, an instantaneous reward is set. Short section bonus Global Rewards And the corresponding volatility risk level Calculate the correction factor. :
[0124] Corrected weights After nonnegation and normalization, it becomes The weights are after smoothing. Substituting into the formula, we get This result serves as the reward feedback input to the policy network, enabling it to significantly converge to a consistent parameter adjustment scheme that reduces short-term volatility and maintains global stability under high-risk transient conditions.
[0125] Step S6: Input the meta-state features into the reinforcement learning policy network, and calculate the optimal color control action based on the reward attribution information of the current meta-state and historical similar states, thereby realizing intelligent decision mapping output for color parameters and process settings. Specifically, this includes: S6.1: Perform feature vector standardization on the multimodal meta-state feature set and use batch normalization algorithm to ensure the consistent distribution of meta-state features in the reinforcement learning policy network, so as to provide balanced input for stable training and inference of downstream policy network.
[0126] S6.2: Based on standardized meta-state features, the temporal feature fusion module (such as a temporal convolutional neural network) is invoked to progressively extract features from the continuous meta-state feature sequence in order to comprehensively capture significant fluctuation patterns in the historical context of printing conditions and provide high-dimensional contextual information for subsequent reward attribution.
[0127] The input is a batch-normalized multimodal meta-state feature sequence, covering the standardized meta-state feature vector data of each detection point in the color box printing production line under continuous time sequence.
[0128] A temporal convolutional neural network (TCN, with parameters such as kernel width k=3, number of layers l=4, and stride s=1) is used to perform one-dimensional temporal convolution on the input continuous state feature sequence. This enables feature perception and extraction of local dynamic changes within the historical window of production conditions, capturing temporal patterns such as high-frequency fluctuations, short-cycle trends, and abnormal jitter.
[0129] Furthermore, by using residual connections and layer normalization methods to enhance the feature transfer between different convolutional layers, the network's ability to model deep temporal structure information is improved, gradient vanishing is prevented, and the effective transmission of long-window temporal dependencies is ensured.
[0130] Furthermore, a global temporal pooling algorithm (parameters: pooling window covers the entire temporal input, pooling method such as max pooling or average pooling) is used to compress the feature sequence extracted by multi-layer convolution along the time axis to generate an aggregated feature vector representing the complete temporal historical context, thus making up for the problem of information loss in short windows.
[0131] Furthermore, an attention weight allocation mechanism (such as self-attention, time-series weighted averaging, etc.) is introduced to apply dynamic weights to the feature components of the aggregated feature vector at each time step, thereby enhancing the ability to sensitively perceive time segments in historical working condition information that are highly correlated with the current reward attribution.
[0132] By using the high-dimensional context embedding features of the output, the complex working condition context patterns such as significant dynamic fluctuations and trend changes in the historical meta-state of the printing production line are transmitted to the subsequent reward attribution and decision network module in a structured vector manner, thereby achieving a fine representation of high-frequency color dynamics and in-depth support for reward information.
[0133] For example, for the normalized meta-state feature sequence, each temporal window is set to 64 frames (approximately a 3-second sampling period), and a 4-layer temporal convolutional network is used with a kernel width of 3 and output channels of 32 / 64 / 128 / 256 respectively. Batch normalization and ReLU activation functions are applied to each layer to enhance nonlinear modeling capabilities. Global average pooling is used on the convolution results to generate 256-dimensional historical context features. To emphasize sudden color change events, a self-attention weighting mechanism is set, and softmax normalization is applied to the feature channels, with the weight distribution controlled by the current reward attribution correlation. On the test set, for production cycles containing high-frequency ink color fluctuations, the high-dimensional context features of the convolutional network achieve an accuracy of over 97% in perceiving abnormal states, effectively separating three typical modes: working condition switching, abnormal fluctuations, and stable operation. The output features serve as subsequent inputs to the policy network, enabling color control decisions to have rapid response and trend prediction capabilities for historical color anomalies.
[0134] S6.3: Input the fused meta-state feature sequence into the reinforcement learning policy network (using DQN, DDPG, or hierarchical policy network, etc.), and combine it with historical similar meta-states and their reward attribution records. Through multivariate regression weighting, the association between the current meta-state and historical reward attribution is dynamically matched to achieve joint input modeling of reward attribution information and meta-state features.
[0135] S6.4: In the policy network, based on the joint modeling results, the optimal color control action under the current meta-state features is searched using the Q-value function or policy gradient optimization algorithm, and the optimal color parameter adjustment decision is output to maximize the multi-scale dynamic reward objective function.
[0136] The input is a joint input feature tensor processed by S6.3 multivariate regression weighted matching, which contains a fusion vector of the current meta-state high-dimensional context features and historical reward attribution mapping information.
[0137] Optimization algorithms using either Deep Q-Network (DQN) based on Q-value functions or Deterministic Policy Gradient (DDPG) based on policy gradients (parameter: discount factor) Learning rate Experience replay buffer size This enables the estimation of action value or strategy probability based on the current fusion characteristics.
[0138] Furthermore, all feasible color control actions in the current state are calculated through forward propagation. The predicted Q value or policy probability distribution The output is then mapped to a color parameter adjustment vector in either a continuous or discrete motion space.
[0139] Furthermore, a greedy-exploration equilibrium strategy is adopted. Or explore Gaussian noise, (Initial value 0.3 and exponentially decays to 0.05) Under the premise of meeting industrial safety constraints, the action with the highest estimated Q value or strategy probability is selected to ensure that the model can explore new strategies and utilize existing optimal strategies under dynamic operating conditions.
[0140] Furthermore, based on Bellman's optimality principle, reward feedback is utilized. and the maximum Q value at the next moment (DQN case) or policy gradient direction (DDPG case), calculate the target value and backpropagate to update the network parameters: , in, The multi-scale dynamic integrated reward output by S5.5 This is the discount factor.
[0141] Furthermore, gradient clipping and parameter regularization constraints are introduced (such as...). Regularization coefficient This helps prevent the network from overfitting or gradient explosion under high-risk conditions, thus ensuring the stability and interpretability of the output action decisions.
[0142] Through deep reinforcement learning optimization, the fused features are mapped into a color parameter adjustment decision vector that maximizes the multi-scale dynamic reward function, thereby achieving the comprehensive optimal control objective of instantaneous color difference, short-term fluctuations and global consistency at the action level.
[0143] For example, on a four-channel color box printing production line, the input feature vector has a dimension of 512, and continuous color control actions are generated through a policy network (DDPG). , , , The batch size for experience replay is set to 64, and the soft update coefficient is set to... In one sampling, The maximum Q value at the next moment is 0.78, calculated using the formula. The network uses the Adam optimizer. After iteratively updating the strategy parameters for 500 steps, the average color difference of the current batch decreased by 0.35, the short-term volatility variance decreased by 27%, and the global consistency index improved by 0.18 units, verifying that the decision output can effectively balance multi-scale reward objectives and steadily improve the color consistency of production.
[0144] S6.5: Verify and optimize the color control actions output by the strategy network. Use action space regularization or industrial parameter constraint optimization modules to ensure that the control actions meet the safety range of the printing production line process parameters and the real-time response characteristics of the equipment, and achieve the final intelligent decision mapping output.
[0145] Step S7: Based on the optimal color control action output by the reinforcement learning strategy network, send color adjustment and process parameter setting instructions to downstream equipment in the color box printing production line, and execute real-time color control operations to achieve target color consistency. Specifically, this includes: S7.1: Based on the optimal color control action output by the reinforcement learning strategy network, the structured color parameter settings and process adjustment data packets are parsed and generated to form a specific instruction set that can be mapped to the printing equipment operation layer, ensuring the accurate expression and transmission of intelligent decision results.
[0146] S7.2: Adapt communication protocols for structured color parameter settings and process adjustment data packets. Encode them using industrial fieldbus or industrial Ethernet protocols to achieve standardized conversion between color parameter settings and equipment control commands, ensuring that the commands can be recognized and responded to by downstream printing equipment.
[0147] S7.3: The protocol-adapted equipment control command package is synchronously distributed to each control unit downstream of the color box printing production line, including the color adjustment module, ink supply unit and environmental compensation system, so as to realize multi-point collaborative dynamic parameter adjustment, improve system response speed and overall color consistency.
[0148] For the device control command packets processed by the protocol adaptation, a multi-channel synchronous distribution scheduling algorithm (parameter: number of distribution nodes n=3, including color adjustment module, ink supply unit and environmental compensation system) is adopted to realize parallel command reception of each downstream control unit.
[0149] Furthermore, through a distributed real-time bus control protocol (such as Profinet / Modbus-TCP / IP, with MTU=1500 bytes and a transmission rate>=100Mbps), structured instruction packets are broadcast or forwarded to the corresponding controller interface according to the specific address allocation mechanism of the target downstream unit.
[0150] Furthermore, an instruction integrity verification module (based on the CRC32 cyclic redundancy check algorithm) is adopted to perform real-time data packet verification on each distributed device control instruction to ensure that there are no bit errors or information loss during the transmission of command data. If the verification fails, the instruction is automatically resent and an exception flag is fed back.
[0151] Furthermore, for the control commands that have been successfully distributed, the parameter parsing engine (such as the PLC parser) inside each downstream control unit parses and detects the parameter threshold values of the received color adjustment parameters, ink supply adjustment amount and environmental compensation configuration, verifies the legality of the parameters and confirms the executability of the commands.
[0152] Through the parameter priority scheduling module (by color adjustment > ink supply > environmental compensation), each downstream unit loads the parameters to be executed sequentially or in parallel, and pre-writes the parameter buffer to prepare for subsequent real-time adjustment actions.
[0153] A feedback synchronization control mechanism is adopted. When each control unit receives a valid instruction and enters the execution state, it immediately feeds back the instruction reception and parameter loading status flags to the host computer, forming an adaptive distribution and execution synchronization chain to ensure the execution synchronization and system response consistency when coordinating multiple points across units.
[0154] Through the aforementioned multi-point synchronous distribution, parameter validity verification, priority scheduling, and feedback synchronization mechanisms, the device control instruction packages adapted to the protocol are efficiently and reliably distributed to the core control units downstream of the printing production line, enabling coordinated dynamic adjustment of color adjustment, ink supply, and environmental compensation parameters, significantly improving the response speed of parameter adjustment and the overall color consistency of the printing line.
[0155] For example, for a four-channel color box printing production line, the optimal control decision output by the reinforcement learning strategy network is encoded as a structured instruction packet containing the target density of color channels ΔC, ΔM, ΔY, and ΔK, the ink pump output flow rate ΔQ, and the environmental compensation airflow ΔV. The instruction packet is 256 bytes long. It is broadcast at a frequency of 20Hz per second via the Modbus-TCP protocol to the color adjustment PLC (D1), the ink supply unit (D2), and the environmental compensation system (D3). Each instruction packet carries a CRC32 checksum. After receiving the instruction, the color adjustment module first parses the ΔC, ΔM, ΔY, and ΔK parameters and performs interval validity checks, such as... The ink supply unit can complete ΔQ writing within 60ms at 70% bus load. The environmental compensation system can deploy ΔV parameters within 120ms. After each unit executes parameter loading, it returns to the main controller via the Success-ACK feedback signal. Once all units have successfully received and loaded the parameters, a complete collaborative color adjustment cycle is driven. During the test cycle, the end-to-end response time of the three types of parallel distributed instructions was less than 350ms, and the system color consistency index steadily improved by 0.4 units, greatly improving the collaborative efficiency and consistency of color adjustment under high dynamic batch conditions.
[0156] S7.4: During the process of color adjustment and process parameter setting in downstream equipment, the equipment feedback signals are collected in real time. Based on the preset equipment feedback acquisition interface and the custom feedback parsing engine, the actual execution status of the color adjustment action is obtained, so as to realize real-time monitoring of the physical control process.
[0157] S7.5: Perform multi-dimensional comparison between the real-time device feedback signal and the original color parameter setting command, and use an autoregressive anomaly detection algorithm to judge the consistency of execution. If a deviation is detected, trigger the compensation control process, automatically generate a second correct color parameter setting command and feed it back to the reinforcement learning strategy network to form an adaptive closed-loop control.
[0158] Step S8: Real-time analysis of the control feedback signals and reward function output generated during the execution process; monitoring of reward attribution stability and meta-state representation capability; and determination of whether reward drift or state feature degradation exists. Specifically, this includes: S8.1: Data acquisition and structured processing of control feedback signals during execution to form a real-time control feedback data stream containing color consistency, color difference index and equipment operating parameters, providing standard input for subsequent reward function attribution analysis.
[0159] S8.2: Based on the collected control feedback data stream, the multi-time window moving average and coefficient of variation statistical algorithms are applied to perform trend analysis and anomaly detection on the reward function output sequence in order to identify potential systematic drift in the reward signal.
[0160] S8.3: Calculate the correlation between the control feedback signal and the current meta-state feature set, and use cluster distance analysis and principal component analysis to quantify the stability and effectiveness of the meta-state representation capability, forming a meta-state attribution capability evaluation index.
[0161] S8.4: Based on the meta-state attribution capability assessment index and the reward signal trend analysis results, and according to the set stability threshold, perform state stability determination of the reward feedback closed loop, and output the discrimination label of reward drift or state feature degradation.
[0162] S8.5: The discrimination label is used as a monitoring signal and pushed to the meta-state aggregation adaptive fine-tuning module and the reinforcement learning strategy network structure optimization module in real time to ensure that subsequent links can perform timely response optimization operations in response to reward instability or representation degradation, forming a data-driven adaptive optimization closed-loop control.
[0163] Step S9: If reward drift or state feature degradation is detected, the meta-state aggregator is triggered to perform adaptive fine-tuning or re-clustering operations, dynamically optimizing the meta-state space representation and updating the reinforcement learning policy network structure to maintain the stability and convergence efficiency of the reward feedback loop. Specifically, this includes: S9.1: Real-time acquisition of reward drift and meta-state feature degradation alarm signals generated by the reward attribution monitoring module; statistical verification of the distribution of key meta-state features and the trend of reward expectation changes in the state space using an anomaly discrimination algorithm to determine whether the adaptive optimization process of the meta-state aggregator needs to be initiated.
[0164] S9.2: Based on the reward drift detection results, clustering effectiveness analysis is performed on the abnormal meta-state feature subsets in the existing meta-state space. Self-supervised clustering algorithms (such as density-based spatial clustering DBSCAN, adaptive K-means, etc.) are used to perform spatial re-clustering of abnormal meta-state features to obtain new meta-state aggregation expressions, thereby achieving effective repair of the degradation of existing representation capabilities.
[0165] S9.3: Perform multi-dimensional feature alignment operations on the clustered reconstructed meta-state aggregate representation and the historical meta-state distribution. Use temporal feature matching methods such as Dynamic Time Warping (DTW) to ensure that the new meta-state space representation is consistent with the historical associated data of the reinforcement learning strategy, and prevent policy gaps and learning memory loss.
[0166] S9.4: Based on the meta-state space representation obtained from re-clustering, the parameters of the relevant temporal multi-scale reward function are recalculated, and the dynamic weights of instantaneous reward, short segment reward and batch global reward are corrected to ensure that reward feedback is driven by the new meta-state representation, thereby improving the real-time accuracy and robustness of reward attribution.
[0167] S9.5: The optimized and adjusted meta-state space representation and reconstructed reward feedback mechanism are input into the reinforcement learning policy network. An online policy fine-tuning algorithm (such as delayed policy update Soft Actor-Critic) is used to automatically update the policy parameter weights and simultaneously forget outdated state-action mapping information, so as to achieve adaptive convergence of the policy network to the new state distribution and accelerate the accurate matching of the control policy under dynamic conditions.
[0168] S9.6: Perform full-process closed-loop monitoring of the adaptive update process of the above-mentioned reinforcement learning strategy network, continuously collect statistical indicators of the effect of adjusted reward attribution and meta-state representation, realize real-time dynamic evaluation, and determine whether it is necessary to enter the next round of meta-state aggregation optimization and policy adaptive update, so as to ensure the adaptive stability of the reward feedback closed loop and the accuracy and robustness of the control system.
[0169] For those skilled in the art, various other corresponding changes and modifications can be made based on the technical solutions and concepts described above, and all such changes and modifications should fall within the protection scope of the claims of this invention.
[0170] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains. The terms “first,” “second,” “third,” and similar terms used in this patent application specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “comprising” or “including” and similar terms mean that the elements or objects preceding “comprising” or “including” encompass the elements or objects listed following “comprising” or “including” and their equivalents, and do not exclude other elements or objects. The “multiple” mentioned in the embodiments of this application refers to two or more. A and / or B indicate three possibilities: A; B; and A and B.
[0171] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An intelligent color analysis and control method for color box printing, specifically comprising: S1: Collect multi-source color detection data at multiple spatial locations and different time points on the color box printing production line, and simultaneously acquire printing speed, consumable batches and equipment micro-environment parameters; S2: Preprocess the collected multi-source color detection data and environmental parameters, and generate a benchmark feature dataset; S3: Based on the normalized baseline feature dataset, an autoencoder neural network is used to extract features from color change, fluctuation amplitude and short-term trend dimensions to generate a multimodal meta-state feature set; S4: For labels in different spatial regions and printing conditions, perform self-supervised clustering on the multimodal meta-state feature set to establish a differentiated meta-state distribution model suitable for multivariate interference scenarios; S5: Based on the differentiated meta-state distribution model, a time-series multi-scale reward function is constructed, which decomposes the reward signal into instantaneous reward, short-segment reward and batch global reward, and assigns dynamic weights to the fluctuation risk level of the meta-state. S6: Input the meta-state features into the reinforcement learning policy network, and calculate the optimal color control action based on the reward attribution information of the current meta-state and the historical similar states. S7: Based on the optimal color control action output by the reinforcement learning strategy network, send color adjustment and process parameter setting instructions to downstream equipment of the color box printing production line to execute real-time color control operations.
2. The intelligent color analysis and control method for color box printing according to claim 1, characterized in that, Step S7 is followed by: S8: Perform real-time analysis on the control feedback signals and reward function output generated during the execution process, monitor the stability of reward attribution and the ability to represent the meta-state, and determine whether there are reward drift or state feature degradation problems. S9: If reward drift or state feature degradation is detected, the meta-state aggregator is triggered to perform adaptive fine-tuning or re-clustering operations, dynamically optimize the meta-state space representation, and update the reinforcement learning policy network structure to maintain the stability and convergence efficiency of the reward feedback loop.
3. The intelligent color analysis and control method for color box printing according to claim 1, characterized in that, The preprocessing of the collected multi-source color detection data and environmental parameters in step S1 includes: noise reduction, normalization and time synchronization.
4. The intelligent color analysis and control method for color box printing according to claim 3, characterized in that, Step S1 specifically includes: the timing synchronization adopts a dynamic time warping algorithm to achieve sampling timing consistency for multi-channel and multi-variable data.
5. The intelligent color analysis and control method for color box printing according to claim 2, characterized in that: The system compares the real-time feedback signals from the equipment with the original parameter setting instructions in multiple dimensions using an autoregressive anomaly detection algorithm. If an execution deviation is detected, the system automatically triggers a compensation control process, generates secondary correction parameter instructions, and feeds them back to the reinforcement learning strategy network to form a closed-loop adaptive control cycle.
6. The intelligent color analysis and control method for color box printing according to claim 1, characterized in that, Step S3 specifically includes: The color detection time series and spatial location information in the normalized benchmark feature dataset are windowed and segmented to generate a sample batch with spatiotemporal local correlation. Based on the windowed sample batch, an autoencoder neural network is used to perform end-to-end reconstruction learning of the original time series features of color changes, so as to adaptively extract the latent space encoding containing the principal components of color, the amplitude of change and the fluctuation pattern as the basic meta-state feature vector. Multi-scale aggregation operations are applied to the basic meta-state feature vectors output by the autoencoder neural network to calculate statistical features such as short-time window average, variance, and fluctuation range amplitude. Multi-scale statistical features are cascaded and combined with original spatial location information and equipment operating condition labels to form a multimodal meta-state feature set. Feature regularization and noise reduction are performed on the multimodal meta-state feature set. Regularization constraints and decorrelation algorithms are used to improve the stability and discriminativeness of the meta-state feature representation.
7. The intelligent color analysis and control method for color box printing according to claim 1, characterized in that, Step S4 specifically includes: Based on the generated multimodal meta-state feature set, the meta-state features of each region are extracted and archived according to the established spatial region partitioning standard, forming a meta-state feature subset grouped by spatial location; For each spatial region's meta-state feature subset, combined with the printing condition labels collected from actual production, label matching and attribution processing is performed to construct meta-state-condition mapping data pairs with diverse condition characteristics. For the structured meta-state feature data pairs, a self-supervised clustering algorithm is used to perform feature space clustering on the meta-state features under the same working conditions in the same region, so as to obtain the typical meta-state cluster centers and corresponding cluster labels that reflect the disturbance of multi-source variables. For the meta-state clustering results under each region-working condition combination, calculate the feature consistency index within the cluster and select meta-state distribution models with high discriminative power. By integrating all spatial regions and working condition cluster centers, and aligning and merging meta-state cluster centers across regions and working conditions, a differentiated meta-state distribution model that is globally applicable to multivariate disturbance scenarios is finally formed.
8. The intelligent color analysis and control method for color box printing according to claim 7, characterized in that, The printing condition label includes printing speed category, consumable batch ID, and equipment microenvironment.
9. The intelligent color analysis and control method for color box printing according to claim 7, characterized in that, The intra-cluster feature consistency metrics include: intra-cluster average distance and silhouette coefficient.
Citation Information
Cited By
Closed-loop automatic adjustment method for printing color deviation
CN122402073A