Multi-model fusion winding coating defect prediction method and system and medium
By employing multi-model fusion technology and physical consistency constraints, the real-time performance and robustness of defect prediction during roll-to-roll coating were addressed, enabling high-precision online defect prediction and closed-loop control, thereby improving the yield and uniformity of roll-to-roll coating.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI INST OF CERAMIC CHEM & TECH CHINESE ACAD OF SCI
- Filing Date
- 2025-12-09
- Publication Date
- 2026-05-01
AI Technical Summary
Existing roll-to-roll coating defect prediction technologies are difficult to implement in real time, traditional single-variable control methods are difficult to establish accurate prediction models, image processing algorithms lose key geometric information in images with extreme aspect ratios, deep learning models have high computational complexity and are difficult to implement in real-time inference on resource-constrained edge devices, and lack robustness guarantee mechanisms.
Employing multi-model fusion technology, combined with ontology semantic driving mechanism and physical consistency constraints, a wrinkle trend prediction model is constructed through predictive gated recurrent units. This model is then integrated with an anomaly detection model to perform multimodal defect detection. Furthermore, N:M structured sparse pruning technology enables lightweight model deployment and supports real-time inference on edge computing devices.
It achieves high-precision online prediction of defects in the roll-to-roll coating process, significantly improving product quality, reducing defective products and raw material waste, increasing the yield to over 82%, and stabilizing coating uniformity within ±3%, realizing the transformation from passive detection to proactive prediction and intervention.
Smart Images

Figure CN121962001A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of roll-to-roll coating equipment, and in particular to a multi-model fusion method, system, and medium for predicting defects in roll-to-roll coating. Background Technology
[0002] Roll-to-roll coating, a core technology in modern thin-film manufacturing, is widely used in high-end manufacturing fields such as flexible thin films for spacecraft, solar panels, and flexible displays. This process deposits functional thin films on a continuously moving substrate using magnetron sputtering technology, offering advantages such as high production efficiency and stable film quality. Traditional quality control in roll-to-roll coating relies primarily on offline inspection and manual judgment, identifying defects through sampling inspections after production. While this post-production inspection method can identify product quality issues, it cannot intervene in real-time during defect formation, leading to a large number of defective products and wasted raw materials.
[0003] Existing defect prediction technologies face multiple challenges. First, the roll-to-roll coating process involves complex nonlinear coupling of multi-dimensional process parameters such as sputtering power, gas flow rate, substrate temperature, and winding speed, making it difficult for traditional single-variable control methods to establish accurate prediction models. Second, film surface images often exhibit extreme aspect ratios, and existing image processing algorithms are prone to losing crucial geometric information when processing such data, affecting prediction accuracy. Furthermore, interference factors in industrial environments, such as varying lighting and background noise, can lead to misjudgments in prediction models, and existing technologies lack effective robustness mechanisms. Simultaneously, traditional deep learning models have a large number of parameters and high computational complexity, making real-time inference difficult to achieve on resource-constrained edge devices, thus limiting their application in actual production lines. Summary of the Invention
[0004] In view of the shortcomings of the prior art described above, the purpose of this invention is to provide a multi-model fusion method, system and medium for predicting defects in roll coating. By combining multi-model fusion technology with ontology semantic driving mechanism and physical consistency constraints, high-precision online prediction of defects in the roll coating process is achieved, which can significantly improve product quality and reduce raw material waste.
[0005] To achieve the above objectives, the present invention adopts the following technical solution.
[0006] In a first aspect, the present invention provides a multi-model fusion method for predicting defects in roll-to-roll coatings, employing the following technical solution: Acquire film surface image sequences and process parameter data during the roll-to-roll coating process; A wrinkle trend prediction model is constructed based on a predictive gated recurrent unit. The prediction model integrates an ontology semantic driving mechanism and a physical consistency constraint mechanism. The ontology semantic driving mechanism guides the model to focus on the deformation characteristics of wrinkle defects through a spatial attention map generated by an external proxy model. The physical consistency constraint mechanism establishes a physical constraint loss function by associating the image gradient field with the material stress field. An anomaly detection model is constructed, generating negative anomaly prompts through a single-class prompting learning mechanism, and fusing 2D images and 3D point cloud data for multimodal defect detection; and By integrating the outputs of the wrinkle trend prediction model and the anomaly detection model, defect prediction information is generated and control commands are output to adjust the coating process parameters.
[0007] Furthermore, in the above-mentioned method for predicting defects in roll-to-roll coating, the acquisition of film surface image sequences and process parameter data during the roll-to-roll coating process includes: A sequence of membrane surface images with an 8:1 aspect ratio was acquired using an industrial line scan camera; Data on process parameters such as sputtering power, gas flow rate, substrate temperature, and winding speed were collected; and The membrane image sequence and process parameter data are aggregated into the SCADA platform for data integration.
[0008] Furthermore, in the above-mentioned roll-to-roll coating defect prediction method, the ontology semantic driving mechanism includes: Contrast-limited adaptive histogram equalization is used to preprocess the membrane image to overcome uneven illumination. Pseudo-labels for wrinkled defect regions are generated using Canny edge detection and morphological dilation operations; and The spatial attention map K is obtained by gradient-weighted class activation mapping, and the spatial attention map K is weighted as a continuous intermediate supervision signal and added to the lateral information flow of the predictive gated recurrent unit.
[0009] Furthermore, in the above-mentioned roll-to-roll coating defect prediction method, the physical consistency constraint mechanism includes: The gradients of the predicted and ground images are calculated using the Sobel operator, implicitly modeling the stress field on the membrane surface; and Construct a fusion loss function L=WMSE+λ(1-IoUsobel), where WMSE is the mean square error weighted by the spatial attention map K, and λ(1-IoUsobel) is the physical consistency loss term established by the cross-union ratio of the predicted stress field and the actual stress field.
[0010] Furthermore, in the above-mentioned roll-to-roll coating defect prediction method, the single-class cue learning mechanism includes: By semantically concatenating learnable normal cues with anomalous suffixes, a set of negative cues with anomalous semantics is constructed; and An explicit anomaly interval loss is introduced, which forces the distance from normal sample features to normal prototypes to be less than the distance from them to anomalous prototypes, thus establishing a discrimination boundary in the absence of real anomalous samples.
[0011] Furthermore, in the above-mentioned method for predicting defects in roll-to-roll coatings, the multimodal defect detection by fusing two-dimensional images and three-dimensional point cloud data includes: The 3D point cloud features are registered and transformed to a 2D image plane using inverse distance weighted interpolation and camera projection model; Feature fusion using unsupervised contrastive learning; and Independent memory banks were constructed for RGB images, point clouds, and fused features, and a single-class support vector machine was used for decision layer fusion.
[0012] Furthermore, the above-mentioned method for predicting defects in roll-to-roll coatings also includes: The wrinkle trend prediction model and anomaly detection model are lightweighted using N:M structured sparse pruning technology, where N:M structured sparse pruning adopts a 2:4 sparse mode, retaining 2 non-zero weights in 4 consecutive weights.
[0013] Furthermore, in the above-mentioned method for predicting defects in roll-to-roll coatings, the lightweighting process includes: Use a sparse refined pass-through estimator to maintain the N:M constraint during training; Deploying the lightweight model on edge computing devices enables millisecond-level real-time inference; and The prediction results are converted into control commands through a PyQt-based visualization platform, and then linked in real time with the tension control system and sputtering power control system to achieve closed-loop control.
[0014] Secondly, the present invention provides a multi-model fusion system for predicting defects in roll-to-roll coatings, which employs the following technical solution: The data acquisition module is used to acquire the film surface image sequence and process parameter data during the roll-to-roll coating process; The wrinkle trend prediction module constructs a wrinkle trend prediction model based on a predictive gated recurrent unit. The prediction model integrates an ontology semantic driving mechanism and a physical consistency constraint mechanism. The ontology semantic driving mechanism guides the model to focus on the deformation characteristics of wrinkle defects through a spatial attention map generated by an external proxy model. The physical consistency constraint mechanism establishes a physical constraint loss function by associating the image gradient field with the material stress field. The anomaly detection module generates negative anomaly prompts through a single-class prompting learning mechanism and integrates 2D images and 3D point cloud data for multimodal defect detection; and The model fusion module is used to fuse the output results of the wrinkle trend prediction model and the anomaly detection model to generate defect prediction information and output control commands to adjust the coating process parameters.
[0015] Thirdly, the present invention provides a readable storage medium, which adopts the following technical solution: A readable storage medium storing computer instructions that, when executed by a processor, implement the roll-to-roll coating defect prediction method as described in any one of the first aspects above.
[0016] In summary, compared with the prior art, the present invention has at least one of the following beneficial technical effects: This invention combines a wrinkle trend prediction model with an anomaly detection model to enable early prediction and real-time intervention of defects during the coating process. Compared with traditional post-delivery inspection methods, this significantly reduces the rate of defective products and raw material waste. Specifically, the ontology semantic-driven mechanism guides the model to accurately identify the deformation characteristics of wrinkle defects through spatial attention maps; the physical consistency constraint mechanism improves the physical rationality of the prediction by associating the image gradient field with the material stress field; the single-class cue learning mechanism can effectively establish discrimination boundaries even in the absence of anomaly samples; and multimodal data fusion further enhances the robustness of detection, thereby achieving high-precision online defect prediction and supporting closed-loop control optimization. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart of an embodiment of a multi-model fusion roll-to-roll coating defect prediction method of the present invention is shown.
[0019] Figure 2 A flowchart of another embodiment of the multi-model fusion roll coating defect prediction method of the present invention is shown.
[0020] Figure 3 A flowchart of a multimodal defect detection method according to the present invention is shown.
[0021] Figure 4 A flowchart of a model lightweighting and deployment process according to the present invention is shown.
[0022] Figure 5 The diagram shows a structural block diagram of a multi-model fusion roll-to-roll coating defect prediction system according to the present invention.
[0023] Figure 6 The diagram shows a hierarchical architecture of a multi-model fusion roll-to-roll coating defect prediction system according to the present invention. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Furthermore, it should be understood that the specific embodiments described herein are only for illustration and explanation of this application and are not intended to limit this application.
[0025] It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments of this application. Furthermore, the descriptions of each embodiment in the following embodiments have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0026] The method steps described in this embodiment of the invention can be executed in the order described in the specific implementation, or the execution order of each step can be adjusted according to actual needs, provided that the technical problem can be solved. These are not listed one by one here.
[0027] The present invention will be further described in detail below with reference to the accompanying drawings.
[0028] Reference Figure 1 A multi-model fusion method for predicting defects in roll-to-roll coating includes several sequentially executed steps for predicting and controlling defects during the roll-to-roll coating process. This method 100 provides a comprehensive defect prediction solution by integrating various advanced machine learning techniques and physical constraint mechanisms.
[0029] In some implementations, method 100 begins at step 102, acquiring a sequence of film surface images and process parameter data during the roll-to-roll coating process. Step 102 involves collecting various types of data from the actual roll-to-roll coating production line, including visual information of the film surface captured by industrial imaging equipment and various process parameters reflecting the state of the production process. This data provides the basic input for subsequent prediction and detection algorithms.
[0030] The process then proceeds to step 104, where a wrinkle trend prediction model is constructed based on predictive gated recurrent units. The prediction model built in step 104 employs an improved recurrent neural network architecture, which exhibits lower computational complexity and faster inference speed compared to traditional long short-term memory networks. The predictive gated recurrent units process time-series data through a simplified gating mechanism, enabling them to capture the temporal evolution of membrane wrinkle formation.
[0031] In step 106, method 100 integrates the ontology semantic driving mechanism and the physical consistency constraint mechanism into the wrinkle trend prediction model. The ontology semantic driving mechanism guides the model to focus on the deformation characteristics of wrinkle defects through a spatial attention map generated by an external proxy model, rather than being disturbed by non-critical factors such as illumination changes. The physical consistency constraint mechanism establishes a physical constraint loss function by associating the image gradient field with the material stress field, ensuring that the prediction results conform to the basic laws of materials mechanics.
[0032] Method 100 then proceeds to step 108, constructing an anomaly detection model. The anomaly detection model in step 108 is specifically designed to identify various defect patterns on the film surface, including wrinkles, uneven deposition, and other problems. This model employs a single-class classification method, enabling effective defect detection even in the absence of a large number of anomaly samples.
[0033] In step 110, method 100 generates anomalous negative cues through a single-class cue learning mechanism. Step 110 addresses the problem of scarce negative samples in traditional anomaly detection by combining normal cues with anomalous suffixes using semantic concatenation technology to automatically construct a set of negative cues with anomalous semantics, thereby establishing clear discrimination boundaries during training.
[0034] Method 100 then proceeds to step 112, which fuses two-dimensional images and three-dimensional point cloud data for multimodal defect detection. Step 112 improves the accuracy and robustness of defect detection by integrating information from different sensor modalities. Two-dimensional images provide surface texture and color information, while three-dimensional point cloud data supplements geometric shape and depth information; the combination of the two enables the detection of complex defects that are difficult to identify with a single modality.
[0035] In step 114, method 100 fuses the outputs of the wrinkle trend prediction model and the anomaly detection model. Step 114 intelligently fuses the forward-looking information from the prediction module with the real-time perception results from the detection module to generate more comprehensive and accurate defect analysis results. This fusion strategy combines the advantages of predictive analysis and real-time detection.
[0036] Finally, in step 116, defect prediction information is generated and control commands are output to adjust the coating process parameters. Step 116 converts the fused analysis results into operable control commands, which can adjust process parameters such as winding speed, tension control, and sputtering power in real time, realizing a shift from a passive detection to an active prevention quality control mode.
[0037] Reference Figure 2 The method 200 details the specific implementation process of data acquisition and wrinkle trend prediction. It achieves precise acquisition and processing of membrane surface images and process parameters through a series of sequentially executed steps.
[0038] In some implementations, method 200 begins at step 202, acquiring a sequence of membrane surface images with an aspect ratio of 8:1 using an industrial line scan camera. The 8:1 extreme aspect ratio image format used in step 202 preserves key geometric information such as the overall distribution pattern and spatial correlation of wrinkles compared to the standard 1:1 square image format. This aspect ratio design avoids the severe loss of geometric information during traditional image compression, enabling the prediction model to more accurately capture the global structural features of wrinkle defects.
[0039] Method 200 then proceeds to step 204, acquiring process parameter data for sputtering power, gas flow rate, substrate temperature, and winding speed. The process parameter sensors in step 204 also monitor vacuum and tension parameters; these high-dimensional process parameters exhibit complex nonlinear coupling relationships. Sputtering power controls the thin film deposition rate, gas flow rate affects the deposition environment, substrate temperature determines the thin film crystallization quality, and winding speed is directly related to the mechanical stress distribution on the thin film surface.
[0040] In step 206, method 200 aggregates the membrane surface image sequence and process parameter data to the SCADA platform for data integration. Step 206 achieves unified management and synchronous processing of heterogeneous data sources. The SCADA platform, as the data integration center, coordinates the multi-source information flow from vision sensors and process parameter sensors, providing a standardized data interface for subsequent intelligent analysis.
[0041] The ontology semantic-driven mechanism of Method 200 is implemented starting from step 208. It employs contrast-limited adaptive histogram equalization to preprocess the membrane image to overcome uneven illumination. The CLAHE algorithm in step 208 serves as an unsupervised preprocessing link, eliminating brightness unevenness caused by changes in illumination conditions through local contrast enhancement techniques. This prevents the model from misinterpreting brightness changes as deformation semantics and thus generating spurious regression phenomena.
[0042] In step 210, method 200 uses Canny edge detection and morphological dilation to generate pseudo-labels for the wrinkled defect region. Step 210 identifies the deformation boundary on the membrane surface using an edge detection algorithm, and then applies morphological dilation transformation to expand the detection area, adaptively generating a rough bounding box for the wrinkled defect region as a pseudo-label, providing spatial localization information for subsequent supervised learning.
[0043] In step 212 of method 200, a spatial attention map K is obtained through gradient-weighted class activation mapping, and the spatial attention map K is used as a continuous intermediate supervision signal and weighted into the lateral information flow of the predictive gated recurrent unit. Step 212 uses an external defect detection proxy model to obtain a feature map K of the defect region of interest through Grad-CAM technology. This feature map K is continuously incorporated into the lateral information flow of the PredGRU unit during training. This guides the model to shift its learning focus from unstable background brightness to physically meaningful wrinkle deformation semantics.
[0044] The physical consistency constraint mechanism of Method 200 uses the Sobel operator in step 214 to calculate the gradients of the predicted and ground images, implicitly modeling the stress field of the membrane surface. Step 214, based on the assumption that the image gradient field is closely related to the material stress field, uses the discrete differential operator Sobel to implicitly model the image gradient, through calculation... An approximate stress field distribution is obtained. The predictive gated recurrent unit uses a large 9×9 convolutional kernel to expand the receptive field, ensuring that the network can capture the global structural information of the folds.
[0045] In step 216, method 200 constructs a fusion loss function L = WMSE + λ(1 - IoUsobel). The fusion loss function in step 216 comprises two components: WMSE is the mean squared error weighted using a spatial attention map K, emphasizing the prediction accuracy of the defect region; λ(1 - IoUsobel) is a physical consistency loss term established by the cross-union ratio (CUNR) of the predicted stress field and the actual stress field, where λ is a weight hyperparameter used to balance the two losses. In a specific embodiment, this weight coefficient λ can be set to 0.001. By minimizing the difference in the CUNR between the predicted stress field and the approximate actual stress field, the model is forced to learn deformation patterns that conform to physical laws during backpropagation, achieving a deep integration of data-driven approaches and physical laws.
[0046] Reference Figure 3 This paper demonstrates the specific implementation process of a single-class prompting learning mechanism and multimodal data fusion. The method 300 achieves robust defect detection even in the absence of anomalous samples by integrating semantic splicing technology, explicit anomaly interval loss, and multimodal feature fusion.
[0047] In some implementations, method 300 begins at step 302, where learnable normal cues are concatenated with anomalous suffixes via semantic concatenation. The semantic concatenation mechanism in step 302 concatenates learnable normal cues {P1...PEN~obj.} with manually designed anomalous suffixes (MAPs) or learnable anomalous suffixes (LAPs), automatically constructing a set of negative cues with anomalous semantics. This step employs a VV-CLIP backbone network, enhancing the extraction capability of local features through the VV attention mechanism to meet the requirements of the localization task.
[0048] Method 300 then proceeds to step 304, which constructs a set of negative cue messages with anomalous semantics. The negative cue message set constructed in step 304 solves the problem of missing negative samples in industrial single-class anomaly detection. By combining normal semantic cue messages with anomalous semantic suffixes, it generates diverse anomaly descriptions, providing a semantic basis for the subsequent establishment of discrimination boundaries.
[0049] In step 306, method 300 introduces an explicit anomaly margin loss, which forces the distance from a normal sample feature to a normal prototype to be less than its distance to an anomalous prototype, thus establishing a discrimination boundary in the absence of real anomalous samples. The explicit anomaly margin (EAM) loss L_ema in step 306, through a regularization term, explicitly forces the distance from a normal sample feature z to a normal prototype w_n to be less than its distance to an anomalous prototype w_a during training, effectively defining the discrimination boundary between the normal and anomalous feature spaces in the absence of real anomalous samples.
[0050] Method 300 then reaches decision point 308, determining whether 3D point cloud data processing is required. Step 308, based on the specific defect detection task requirements and available sensor configuration, decides whether to enable the multimodal fusion processing path of the 3D point cloud data to detect defects related to the 3D geometry.
[0051] When decision point 308 determines that 3D point cloud data processing is required, method 300 proceeds to step 310, where the 3D point cloud features are registered and transformed to a 2D image plane using inverse distance weighted interpolation and a camera projection model. In step 310, PointTransformer and Vision Transformer are used to extract deep features from the point cloud and the image, respectively. Considering the sparsity of point cloud features, this step obtains the group centers of the point cloud through farthest point sampling (FPS). Then, inverse distance weighted interpolation (IDW) is used to interpolate the sparse features back to each point in the original point cloud. Finally, the camera's intrinsic and extrinsic parameters are used to accurately project these 3D feature points onto the 2D image plane, resulting in a point cloud feature map F_pt that is strictly aligned with the image feature space.
[0052] Method 300 then proceeds to step 312, which utilizes unsupervised contrastive learning for feature fusion. Step 312 employs the InfoNCE loss function to achieve unsupervised fusion of RGB image features and 3D point cloud features. Through a contrastive learning mechanism, the semantic correspondence between the two modalities is learned, avoiding information loss that may occur during early fusion.
[0053] In step 314, method 300 constructs independent memory banks for the RGB image, point cloud, and fused features, and uses a single-class support vector machine (OCSVM) for decision-level fusion. Step 314 adopts a hierarchical late-stage fusion strategy, constructing independent memory banks for the RGB image, point cloud, and fused features, and then using an OCSVM for decision-level fusion. This method preserves the discriminative power of a single modality to the greatest extent.
[0054] When decision point 308 determines that 3D point cloud data processing is not required, method 300 directly proceeds to step 316 to perform 2D image anomaly detection. Step 316 specifically handles the defect detection task based solely on 2D RGB images, identifying surface defects on the film surface through a single-modal anomaly detection algorithm.
[0055] Method 300 finally outputs the multimodal defect detection results in step 318. Step 318 integrates the detection results from different processing paths to generate a comprehensive defect detection report, including information such as the location, type, and severity of the defects, providing a basis for decision-making in subsequent process parameter adjustments.
[0056] Reference Figure 4 This paper demonstrates the specific implementation process of N:M structured sparse pruning technology and the deployment process of the closed-loop control system. This method 400 achieves real-time deployment of high-precision models on edge devices through a series of optimization steps and constructs a complete perception-prediction-decision-execution closed loop.
[0057] In some implementations, method 400 begins at step 402, applying N:M structured sparse pruning techniques to lightweight the wrinkle trend prediction model and the anomaly detection model. The N:M structured sparse pruning technique in step 402 innovatively combines fine-grained sparsity with structured constraints. Compared to traditional unstructured pruning methods, this technique achieves hardware acceleration while maintaining model performance.
[0058] Method 400 then executes step 404, employing a 2:4 sparse pattern to retain two non-zero weights in every four consecutive weights (using a 2:4 sparse pattern to retain two non-zero weights). The 2:4 sparse pattern in step 404 achieves 50% structured sparsity. This pattern, which retains two non-zero weights in every four consecutive weights, has hardware-friendly advantages, gaining native acceleration support from the NVIDIA Ampere architecture GPU sparse tensor cores, greatly improving actual inference speed.
[0059] In step 406, method 400 uses a sparse refined pass-through estimator (SR-STE) to maintain the N:M constraint during training. The SR-STE mechanism in step 406 ensures that the network always satisfies the N:M constraint during backpropagation. By applying a sparse mask in forward propagation and maintaining gradient flow in backpropagation, efficient training under sparse constraints is achieved. This mechanism allows the performance degradation to be successfully controlled to below 5% even when the number of parameters pruned exceeds 50%.
[0060] Method 400 then proceeds to step 408, which deploys the lightweight model on an edge computing device to achieve millisecond-level real-time inference. Step 408's edge computing deployment places the optimized algorithm model on an edge computing platform in the industrial field, serving as a bridge between the intelligent analysis layer and the digital twin layer, enabling the high-precision model to operate efficiently in resource-constrained environments.
[0061] Method 400 then proceeds to decision point 410 to determine whether the current system meets the millisecond-level real-time inference requirements. Step 410 evaluates the model's inference performance to confirm whether it meets the 50-millisecond inference threshold set by the production line. If the inference latency exceeds the threshold due to insufficient computing power or system lag, the system will automatically activate a downsampling strategy to maintain real-time response capability while ensuring the adjustment of key process parameters and defect prevention capabilities.
[0062] When decision point 410 determines that the real-time inference requirements are met, method 400 proceeds to step 412, where the prediction results are converted into control commands through a PyQt-based visualization platform. The PyQt visualization platform in step 412 serves as a human-computer interaction interface, fusing forward-looking information from the wrinkle trend prediction module with real-time perception results from the anomaly detection module. The intelligent analysis layer then generates real-time control commands accordingly.
[0063] When decision point 410 determines that the real-time inference requirements are not met, method 400 proceeds to step 414 to optimize model performance and then returns to step 408 for redeployment. Step 414 implements an iterative optimization mechanism, improving the model's inference efficiency by adjusting the sparsity ratio, optimizing the network structure, or improving the quantization strategy, ensuring that the finally deployed model meets the real-time requirements.
[0064] Method 400 then executes step 416 to achieve real-time linkage with the tension control system and the sputtering power control system. The control commands in step 416 are linked in real time with the servo motor and tension control system through an online visualization platform to achieve adaptive adjustment and active intervention of coating process parameters, including precise adjustment of key parameters such as winding speed, tension control, and sputtering power.
[0065] Method 400 ultimately achieves closed-loop control to optimize coating process parameters in step 418. Step 418 constructs a millisecond-level closed loop from sensing and prediction to decision execution. Through this closed-loop control, the yield of flexible films is significantly improved from 35% to over 82%, while the uniformity of the coating is stabilized within ±3%, realizing the transformation of quality control from passive "post-event detection" to proactive "pre-event prediction and intervention".
[0066] In some implementations, the continuous learning framework employs an instance-aware cue tuning (IPT) mechanism, which leverages a task-shared cue pool and instance-level query mechanism to achieve task-specific adaptation. IPT matches the most relevant cue to the current image x using the task-shared cue pool P and instance query mechanism, concatenating it with the input. This allows the network to learn task-specific adaptations without updating the core feature extractor, and the encoder (ViT) parameters are frozen to prevent catastrophic forgetting.
[0067] The continuous learning framework employs a gradient-aware parameter decoupling (GPD) mechanism, which projects the gradients of the new task onto a null space that minimizes the impact on the old task through singular value decomposition (SVD). GPD calculates the decentralized covariance of the feature space of the old task and then performs SVD decomposition. When learning a new task, the gradient of the new task will be... Effectively project onto the null space that has the least impact on prior knowledge. This ensures that parameter updates only occur in a subspace that does not interfere with the knowledge of the old task.
[0068] In some implementations, the model fusion module includes a trainable transition matrix M that corrects for strong biased predictions in few-shot incremental learning by applying new knowledge preference constraints. The transition matrix M addresses the "strong bias" problem caused by the tendency for new classes to be overwritten by memories of older classes in few-shot incremental learning by imposing new knowledge preference constraints (ensuring that the diagonal element values corresponding to the new class are larger), thus shifting the biased prediction distribution. Corrected to unbiased distribution This improves the efficiency and accuracy of learning new categories.
[0069] In summary, the multi-model fusion roll-to-roll coating defect prediction method in the above embodiments achieves high-precision online prediction of defects during the roll-to-roll coating process by integrating predictive gated recurrent units and anomaly detection models, combined with ontology semantic driving mechanisms and physical consistency constraint mechanisms. This method uses 8:1 aspect ratio images to retain key geometric information, addresses the problem of scarce negative samples through a single-class cue learning mechanism, fuses 2D images and 3D point cloud data for multimodal defect detection, and employs N:M structured sparse pruning technology to achieve lightweight model deployment. This method can significantly improve the yield of flexible films from 35% to over 82%, and stabilize the coating uniformity within ±3%, realizing a shift in quality control from passive "post-event detection" to proactive "pre-event prediction and intervention." It effectively reduces the generation of defective products and raw material waste, providing a reliable intelligent quality control solution for high-end manufacturing fields such as flexible films for spacecraft.
[0070] This invention also discloses a multi-model fusion system for predicting defects in roll-to-roll coatings. Reference Figure 5 A multi-model fusion roll-to-roll coating defect prediction system 500 adopts a hierarchical architecture design, achieving intelligent monitoring and defect prediction of the roll-to-roll coating process through the collaborative work of multiple functional modules. This system 500 integrates functions such as data acquisition, predictive analysis, anomaly detection, and control execution, providing comprehensive quality assurance for the roll-to-roll coating process.
[0071] Specifically, the multi-model fusion roll-to-roll coating defect prediction system 500 includes a data acquisition module 502, used to acquire film surface image sequences and process parameter data during the roll-to-roll coating process. As the system's data input terminal, the data acquisition module 502 is responsible for collecting multi-source heterogeneous data from the production site, providing basic data support for subsequent intelligent analysis. This data acquisition module 502 ensures unified management and synchronous processing of data from different types of sensors through a standardized data interface.
[0072] The data acquisition module 502 includes an industrial line scan camera 504 for acquiring visual information about the membrane surface. The industrial line scan camera 504 is specifically designed to capture image sequences of the membrane surface with an extreme 8:1 aspect ratio. Compared to traditional area scan cameras, this line scan camera 504 can acquire continuous membrane surface images during high-speed winding, preserving the complete spatial distribution information of wrinkles and defects. The industrial line scan camera 504 uses high-resolution imaging technology to ensure that the image quality meets the accuracy requirements of subsequent defect detection algorithms.
[0073] The data acquisition module 502 also includes process parameter sensors 506, which are used to monitor key process parameters during the roll-to-roll coating process. The process parameter sensors 506 include various types of sensors such as sputtering power sensors, gas flow sensors, substrate temperature sensors, winding speed sensors, vacuum sensors, and tension sensors. These sensors monitor various process parameters affecting film quality in real time, providing a data foundation for correlation analysis between process parameters and defect formation.
[0074] The data acquisition module 502 further includes a SCADA platform 508, which acts as a data integration center to coordinate multi-source data streams. The SCADA platform 508 receives data from the industrial line scan camera 504 and the process parameter sensor 506, enabling unified management, storage, and preprocessing of heterogeneous data sources. This SCADA platform 508 connects to various sensor devices through standardized communication protocols, ensuring the real-time performance and reliability of data acquisition, while also providing a standardized data interface for subsequent intelligent analysis modules.
[0075] The multi-model fusion roll-to-roll coating defect prediction system 500 includes a wrinkle trend prediction module 510, which constructs a wrinkle trend prediction model based on a predictive gated loop unit. The wrinkle trend prediction module 510 is specifically used to analyze the temporal evolution of wrinkle defects on the film surface, and provides a forward-looking prediction of the future wrinkle state through a time-series prediction algorithm, realizing the transformation from passive detection to proactive prevention.
[0076] The wrinkle trend prediction module 510 includes a predictive gated recurrent unit 512, which serves as the core computational engine for time series prediction. The predictive gated recurrent unit 512 replaces the traditional long short-term memory network with a simplified gating mechanism, processing time series data through the synergistic effect of update and reset gates. Compared to the ST-LSTM architecture, it reduces the number of parameters by approximately 30% and improves inference speed by 35%. This predictive gated recurrent unit 512 receives time series data from the SCADA platform 508 and captures the temporal dependencies of wrinkle formation through the memory mechanism of the recurrent neural network.
[0077] The wrinkle trend prediction module 510 integrates the ontology semantic driving mechanism 514, which guides the model to focus on the deformation features of wrinkle defects through a spatial attention map generated by an external proxy model. The ontology semantic driving mechanism 514 obtains a spatial attention map K through gradient-weighted class activation mapping technology. This attention map K is used as a continuous intermediate supervision signal and weighted into the lateral information flow of the predictive gated recurrent unit 512, guiding the model to shift its learning focus from background brightness changes to the physically meaningful wrinkle deformation semantics, effectively avoiding the generation of spurious regression phenomena.
[0078] The wrinkle trend prediction module 510 also integrates a physical consistency constraint mechanism 516, which establishes a physical constraint loss function by associating the image gradient field with the material stress field. Based on the physical assumption that the image gradient field and the material stress field are closely related, the physical consistency constraint mechanism 516 uses the Sobel operator to calculate the gradients of the predicted image and the real image, implicitly modeling the stress field distribution on the membrane surface. This physical consistency constraint mechanism 516 works in conjunction with the ontology semantic driving mechanism 514 to ensure that the prediction results conform to both semantic logic and physical laws.
[0079] The multi-model fusion roll-to-roll coating defect prediction system 500 includes an anomaly detection module 518, which generates negative anomaly prompts through a single-class prompt learning mechanism and fuses two-dimensional images and three-dimensional point cloud data for multi-modal defect detection. The anomaly detection module 518 is specifically designed to identify various defect patterns on the coating surface, including geometric defects such as wrinkles, uneven deposition, and depressions. It improves the accuracy and robustness of defect detection through multi-modal data fusion.
[0080] The anomaly detection module 518 includes a single-class cue learning mechanism 520 to address the scarcity of negative samples in industrial anomaly detection. The single-class cue learning mechanism 520 automatically constructs a set of negative cue messages with anomalous semantics by concatenating learnable normal cue messages with anomalous suffixes using semantic concatenation technology. It also introduces an explicit anomaly margin loss to force the distance from normal sample features to normal prototypes to be less than their distance to anomalous prototypes. This single-class cue learning mechanism 520 is based on the VV-CLIP backbone network and enhances local feature extraction capabilities through the VV attention mechanism.
[0081] The anomaly detection module 518 includes a multimodal data fusion module 522, which integrates two-dimensional image and three-dimensional point cloud data for comprehensive defect detection. The multimodal data fusion module 522 registers and transforms the three-dimensional point cloud features to the two-dimensional image plane using inverse distance weighted interpolation and a camera projection model, and then performs feature fusion using unsupervised contrastive learning. This multimodal data fusion module 522 constructs independent memory banks for the RGB image, point cloud, and fused features, and employs a hierarchical late-stage fusion strategy to maximize the preservation of the discriminative power of each individual modality.
[0082] The anomaly detection module 518 also includes a continuous learning framework 524 for online adaptation and learning of new defect patterns. The continuous learning framework 524 avoids forgetting old knowledge when learning new tasks through instance-aware cue tuning (IPT) and gradient-aware parameter decoupling (GPD) mechanisms. This continuous learning framework 524 freezes encoder parameters, achieves task-specific adaptation through a task-shared cue pool and instance-level query mechanism, and projects the gradients of new tasks onto the null space that has the least impact on old tasks.
[0083] The multi-model fusion roll-to-roll coating defect prediction system 500 includes a model fusion module 526, which fuses the outputs of a wrinkle trend prediction model and anomaly detection model to generate defect prediction information and output control commands to adjust coating process parameters. The model fusion module 526 acts as the system's decision center, intelligently fusing forward-looking information from the prediction module with real-time sensing results from the detection module to generate comprehensive defect analysis results and control strategies. Specifically, the anomaly detection model generates a comprehensive anomaly score. An optimal anomaly detection threshold (e.g., 0.21) is set; when the comprehensive anomaly score exceeds this threshold, an anomaly is identified. To achieve closed-loop control, the system further converts this anomaly score into an analog control signal, such as a 0-5V voltage. This conversion can be achieved by normalizing the anomaly score from the range [0.21, 1.0] to the range [0, 1.0]. For example, the calculation formula is: voltage = ((overall_score - 0.21) / (1.0 - 0.21)) × 5.0. On the other hand, the wrinkle trend prediction module uses the Sobel operator (e.g., the edge threshold can be set to 140 or 145) to extract bounding boxes from the predicted image and the ground truth image, and calculates the intersection-over-union (IoU) between them. This IoU value can be used to assess the confidence of the prediction or to determine whether the prediction is anomalous. For example, an IoU greater than 0.9 can be marked as "very likely to be correct," an IoU greater than 0.8 and less than or equal to 0.9 can be marked as "moderate probability of prediction being anomalous," and an IoU less than 0.8 can be marked as "very likely to be anomalous," thus providing additional decision support for system self-checking or operator intervention. Furthermore, the wrinkle trend prediction module performs a Hough transform on the predicted image to detect wrinkle lines, with a Hough transform threshold set to 50, a minimum line length set to 50, and a maximum line gap set to 20. By calculating the average angle of all detected lines, the module converts it into a second control voltage representing the wrinkle trend, for example, within a range of -5V to +5V. The voltage can be linearly mapped using the tangent of the angle, such as voltage = 0.05 × tan(radians(avg_deg)), and clamped within a range of ±5.0V. These two voltage signals, along with their prediction information and anomaly detection results, are merged and sent to the decision layer fusion 528.
[0084] The model fusion module 526 includes a decision-level fusion module 528, which integrates the analysis results from the wrinkle trend prediction module 510 and the anomaly detection module 518. The decision-level fusion module 528 receives wrinkle trend prediction information from the predictive gated recurrent unit 512 and anomaly detection results from the continuous learning framework 524, and generates a comprehensive defect assessment report through a weighted fusion algorithm, including information such as the location, type, severity, and development trend of the defects.
[0085] The model fusion module 526 also includes a control command generation 530, which converts the fused analysis results into operable process parameter adjustment commands. Based on the defect analysis results provided by the decision-level fusion 528, and combined with the correlation model between process parameters and defect formation, the control command generation 530 generates real-time adjustment commands for key process parameters such as tension control, sputtering power, and winding speed, thereby achieving proactive defect prevention and process optimization.
[0086] Reference Figure 6 The multi-model fusion roll-to-roll coating defect prediction system 500 implements a layered architecture of digital twins, comprising five layers: a physical perception layer, a data integration layer, a digital twin layer, an intelligent analysis layer, and an application interaction layer. This layered architecture ensures seamless integration from bottom-level data acquisition to top-level intelligent decision-making, achieving deep integration between the physical production process and the digital analysis system.
[0087] Specifically, the physical sensing layer, located at the bottom of the architecture, contains various sensing devices for data collection. This layer includes a 3D point cloud module (3D) for acquiring the three-dimensional geometric information of the membrane surface. The 3D point cloud module obtains the three-dimensional point cloud data of the membrane surface through laser scanning or structured light technology, providing spatial information for detecting geometry-related defects. The physical sensing layer also includes other types of sensors, such as temperature sensors and pressure sensors, which together form a complete physical information acquisition network.
[0088] The data integration layer, situated above the physical sensing layer, contains the data collected by the SCADA platform. This layer is responsible for standardizing and uniformly managing heterogeneous data from different sensors, providing a high-quality data foundation for the upper-layer digital twin modeling. The data integration layer ensures data real-time performance and consistency through standardized data interfaces and communication protocols.
[0089] The digital twin layer comprises three types of digital twins: vacuum digital twin, coating digital twin, and radiation digital twin. This layer provides a high-fidelity virtual mapping of the physical production line. The vacuum digital twin simulates the dynamic changes in the vacuum environment, the coating digital twin models the thin film deposition process, and the radiation digital twin analyzes the impact of the radiation environment on thin film performance. These digital twins achieve accurate simulation of the actual process through a combination of physical modeling and data-driven approaches.
[0090] The intelligent analysis layer, situated above the digital twin layer, comprises two main analysis modules. The left module includes OntoNet and SPINet for predictive modeling, corresponding to the functionality of the wrinkle trend prediction module 510. The right module includes RGB+3D point clouds and GPD / IPT components for anomaly detection, corresponding to the multimodal fusion and continuous learning capabilities of the anomaly detection module 518. The intelligent analysis layer achieves intelligent prediction and detection of defects through deep learning algorithms and physical constraint mechanisms.
[0091] The application interaction layer, located at the top of the architecture, contains the PyQt visualization platform for user interface and control. This layer provides operators with an intuitive human-machine interface, displaying real-time system operating status, defect detection results, and process parameter adjustment suggestions. Through visualization technology, the application interaction layer presents complex analysis results in the form of charts, curves, and alarm information, supporting operator decision-making and intervention.
[0092] This layered architecture uses directional arrows to indicate the flow of information upwards from the bottom physical sensing layer, through various processing and analysis stages, and finally providing visualization and control functions at the application layer. This architectural design enables a complete data flow from physical sensors through various processing and analysis stages to ultimately provide visualization and control capabilities, ensuring the system's real-time performance and reliability.
[0093] In summary, the multi-model fusion roll-to-roll coating defect prediction system 500 in the above embodiments, through a hierarchical architecture design, integrates core components such as a data acquisition module 502, a wrinkle trend prediction module 510, an anomaly detection module 518, and a model fusion module 526, achieving intelligent monitoring and defect prediction of the roll-to-roll coating process. The system 500 employs a predictive gated loop unit 512 combined with an ontology semantic driving mechanism 514 and a physical consistency constraint mechanism 516 to construct a wrinkle trend prediction model. Anomaly detection is achieved through a single-class prompting learning mechanism 520 and multi-modal data fusion 522, and lightweight model deployment is achieved using N:M structured sparse pruning technology. The system 500 converts the prediction results into real-time control commands through decision-level fusion 528 and control command generation 530, linking with the tension control system and sputtering power control system to construct a millisecond-level closed-loop control from perception and prediction to decision execution. This system can significantly improve the yield of flexible films from 35% to over 82%, and stabilize the coating uniformity within ±3%. It realizes the transformation of quality control from passive "post-event inspection" to proactive "pre-event prediction and intervention", providing a reliable intelligent quality control solution for high-end manufacturing fields such as flexible films for spacecraft, effectively reducing the generation of defective products and waste of raw materials.
[0094] This invention also discloses a readable storage medium.
[0095] A computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the roll-to-roll coating defect prediction method described in any of the above embodiments. The computer-readable storage medium may include any entity or device capable of carrying a computer program, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), and a software distribution medium, etc. The computer program includes computer program code. The computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable storage medium may include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), and a software distribution medium, etc.
[0096] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
[0097] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a system including a processing module or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).
[0098] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-model fusion method for predicting defects in roll-to-roll coatings, characterized in that, include: Acquire film surface image sequences and process parameter data during the roll-to-roll coating process; A wrinkle trend prediction model is constructed based on a predictive gated recurrent unit. The prediction model integrates an ontology semantic driving mechanism and a physical consistency constraint mechanism. The ontology semantic driving mechanism guides the model to focus on the deformation characteristics of wrinkle defects through a spatial attention map generated by an external proxy model. The physical consistency constraint mechanism establishes a physical constraint loss function by associating the image gradient field with the material stress field. An anomaly detection model is constructed, anomaly negative prompts are generated through a single-class prompt learning mechanism, and multimodal defect detection is performed by fusing two-dimensional images and three-dimensional point cloud data; as well as By integrating the outputs of the wrinkle trend prediction model and the anomaly detection model, defect prediction information is generated and control commands are output to adjust the coating process parameters.
2. The method for predicting defects in roll-to-roll coating according to claim 1, characterized in that, The acquisition of film surface image sequences and process parameter data during the roll-to-roll coating process includes: A sequence of membrane surface images with an 8:1 aspect ratio was acquired using an industrial line scan camera; Data on process parameters such as sputtering power, gas flow rate, substrate temperature, and winding speed were collected; and The membrane image sequence and process parameter data are aggregated into the SCADA platform for data integration.
3. The method for predicting defects in roll-to-roll coating according to claim 2, characterized in that, The ontology semantic driving mechanism includes: Contrast-limited adaptive histogram equalization is used to preprocess the membrane image to overcome uneven illumination. Pseudo-labels for wrinkled defect regions are generated using Canny edge detection and morphological dilation operations; and The spatial attention map K is obtained by gradient-weighted class activation mapping, and the spatial attention map K is weighted as a continuous intermediate supervision signal and added to the lateral information flow of the predictive gated recurrent unit.
4. The method for predicting defects in roll-to-roll coating according to claim 3, characterized in that, The physical consistency constraint mechanism includes: The gradients of the predicted and ground images are calculated using the Sobel operator, implicitly modeling the stress field on the membrane surface; and Construct a fusion loss function L=WMSE+λ(1-IoUsobel), where WMSE is the mean square error weighted by the spatial attention map K, and λ(1-IoUsobel) is the physical consistency loss term established by the cross-union ratio of the predicted stress field and the actual stress field.
5. The method for predicting defects in roll-to-roll coatings according to claim 1, characterized in that, The single-class prompting learning mechanism includes: By semantically concatenating learnable normal cues with anomalous suffixes, a set of negative cues with anomalous semantics is constructed; and An explicit anomaly interval loss is introduced, which forces the distance from normal sample features to normal prototypes to be less than the distance from them to anomalous prototypes, thus establishing a discrimination boundary in the absence of real anomalous samples.
6. The method for predicting defects in roll-to-roll coatings according to claim 5, characterized in that, The multimodal defect detection by fusing two-dimensional images and three-dimensional point cloud data includes: The 3D point cloud features are registered and transformed to a 2D image plane using inverse distance weighted interpolation and camera projection model; Feature fusion using unsupervised contrastive learning; and Independent memory banks were constructed for RGB images, point clouds, and fused features, and a single-class support vector machine was used for decision layer fusion.
7. The method for predicting defects in roll-to-roll coating according to claim 1, characterized in that, Also includes: The wrinkle trend prediction model and anomaly detection model are lightweighted using N:M structured sparse pruning technology, where N:M structured sparse pruning adopts a 2:4 sparse mode, retaining 2 non-zero weights in 4 consecutive weights.
8. The method for predicting defects in roll-to-roll coatings according to claim 7, characterized in that, The lightweighting process includes: Use a sparse refined pass-through estimator to maintain the N:M constraint during training; Deploying the lightweight model on edge computing devices enables millisecond-level real-time inference; and The prediction results are converted into control commands through a PyQt-based visualization platform, and then linked in real time with the tension control system and sputtering power control system to achieve closed-loop control.
9. A multi-model fusion system for predicting defects in roll-to-roll coatings, characterized in that, include: The data acquisition module is used to acquire the film surface image sequence and process parameter data during the roll-to-roll coating process; The wrinkle trend prediction module constructs a wrinkle trend prediction model based on a predictive gated recurrent unit. The prediction model integrates an ontology semantic driving mechanism and a physical consistency constraint mechanism. The ontology semantic driving mechanism guides the model to focus on the deformation characteristics of wrinkle defects through a spatial attention map generated by an external proxy model. The physical consistency constraint mechanism establishes a physical constraint loss function by associating the image gradient field with the material stress field. The anomaly detection module generates negative anomaly prompts through a single-class prompt learning mechanism and integrates two-dimensional images and three-dimensional point cloud data for multimodal defect detection; as well as The model fusion module is used to fuse the output results of the wrinkle trend prediction model and the anomaly detection model to generate defect prediction information and output control commands to adjust the coating process parameters.
10. A readable storage medium, characterized in that, The readable storage medium stores computer instructions that, when executed by a processor, implement the roll-to-roll coating defect prediction method as described in any one of claims 1-8.
Citation Information
Patent Citations
Geomembrane defect detection system under soil and stone medium coverage condition and method thereof
CN120721154A
Film lamination defect detection method and system based on image processing
CN120876486A
A prediction-tuning capsule network for thermal anomaly identification on building envelopes as well as image classification and object detection
US20250148611A1