Industrial visual detection system and method with generation and discrimination collaborative optimization
By optimizing the generation and discrimination of an industrial vision inspection system, the problems of insufficient detection versatility and high cost in existing technologies are solved through the collaborative optimization of a unified encoder, reconstruction decoder and discriminator, achieving efficient and robust industrial vision inspection.
Patent Information
- Application Number
- CN202510985969.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-31
AI Technical Summary
Existing industrial vision inspection technologies lack versatility in terms of high-quality and high-efficiency inspection, and the models cannot be effectively reused. They also have high R&D costs, rely on large-scale labeled data which is difficult to obtain, and lack deep collaborative optimization between generation and discrimination modules, making it difficult to achieve breakthroughs in complex industrial scenarios.
An industrial vision inspection system employing co-optimization of generation and discrimination utilizes a unified encoder to extract common visual features, a reconstructed decoder to perform zero-sample anomaly detection, a conditional generation decoder to generate specified defect images, a discrimination head for recognition and training, and a collaborative loss function and a total loss function to optimize the model through a collaborative optimization module, thereby achieving adversarial and collaborative generation and discrimination.
It achieves efficient and robust industrial vision inspection with low data dependence, can detect difficult defects in complex scenes, improves inspection performance, reduces R&D costs, and has high versatility and self-optimization capabilities.
Smart Images

Figure CN120876407A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and artificial intelligence, specifically to an industrial vision inspection system and method that optimizes the generation and discrimination processes. Background Technology
[0002] The current field of industrial visual inspection faces a core contradiction: on the one hand, production lines urgently require high-quality, high-efficiency inspection; on the other hand, traditional visual inspection technologies suffer from a severe lack of versatility. Mainstream supervised learning methods heavily rely on large-scale, finely labeled defect samples, which are extremely difficult to obtain in industrial environments that strive for "zero defects." Simultaneously, the significant visual differences between different products and production lines result in models being "developed once, used once," unable to be effectively reused, and leading to persistently high R&D costs.
[0003] To address this issue, academia and industry have explored various approaches, including unsupervised anomaly detection (UAD) and data augmentation. However, UAD methods often suffer from high false positive rates and cannot classify defects; while simple data augmentation techniques struggle to achieve the required controllability and realism in generating defects. More importantly, existing technologies tend to treat pre-training, anomaly detection, data generation, and model fine-tuning as independent modules linked together. This simplistic, procedural approach, lacking deep information interaction and collaborative optimization, limits the overall system performance to the simple sum of the shortcomings of each module, making it difficult to achieve systemic breakthroughs in complex and ever-changing real-world industrial scenarios.
[0004] The above-mentioned problems urgently need to be solved. Summary of the Invention
[0005] The purpose of this invention is to overcome at least one technical problem existing in the prior art and to provide an industrial vision inspection system and method that optimizes the generation and discrimination processes.
[0006] On one hand, embodiments of the present invention provide an industrial vision inspection system for collaborative optimization of generation and discrimination. The system includes: a data input module, a model module, a collaborative optimization module, and an output module. The data input module receives an input image to be inspected. The model module includes a unified encoder, a reconstruction decoder, a conditional generation decoder, and a discriminator head. The unified encoder extracts general visual features based on the image to be inspected. The reconstruction decoder performs zero-sample anomaly detection and adaptive threshold segmentation based on the general visual features to obtain a reconstructed image without defects. The conditional generation decoder, based on the general visual features and multimodal instructions, diffuses defect patterns. The system generates a defect image containing specified defects; the discriminator head is used to identify the reconstructed image and output a confidence score based on the identification result; and the system trains on the defect image generated by the conditional generator decoder to generate defect categories and category probabilities; the collaborative optimization module is used to construct a collaborative loss function based on the confidence score of the reconstructed image output by the discriminator head, and obtain a task loss function based on the category probabilities generated by the discriminator head; and to construct a total loss function through the task loss function and the collaborative loss function, and back-update the discriminator head parameters, generator decoder parameters, and unified encoder parameters based on the gradient generated by the total loss function; the output module is used to output the optimized and updated model.
[0007] Furthermore, the unified encoder includes a pre-trained ViT backbone network and a trainable adapter module, which is inserted between the key layers of the unified encoder.
[0008] Furthermore, the reconstruction decoder is also used to calculate the difference between the reconstructed image and the image to be detected based on a preset difference calculation formula to obtain a difference map; and to output abnormal regions based on the difference map using a Gaussian adaptive thresholding algorithm.
[0009] Furthermore, the preset formula for calculating the degree of difference is as follows:
[0010]
[0011] In the formula, I is the image to be detected, I ' To reconstruct the image, |II ' | Calculates the absolute difference at the pixel level. The gradient of the image to be detected. To reconstruct the gradient of the image, The term calculates the difference in image gradients, where α is a weighting coefficient used to balance the importance of pixel differences and gradient differences.
[0012] Furthermore, the conditional generation decoder integrates a conditional encoding unit, a feature fusion unit, and a conditional generation unit; the conditional encoding unit is used to convert the multimodal defect generation instructions provided by the user into a unified mathematical representation vector that the model can understand and recognize; the feature fusion unit is used to fuse the unified mathematical representation vector with the features extracted from the input good product image by the unified encoder using a cross-attention mechanism to generate fused features; the conditional generation unit is used to generate the final defect image by using a diffusion model as the generator and the fused features as the core guide.
[0013] Furthermore, the collaborative optimization module for constructing a collaborative loss function based on the confidence level of the reconstructed image output by the discriminator includes: the discriminator scanning the reconstructed image and outputting a pixel-by-pixel confidence map P. norm Based on the pixel-by-pixel confidence map P norm The low-confidence region is located by comparing it with a preset confidence threshold; a collaborative loss function is constructed for the low-confidence region, the collaborative loss function including:
[0014] L syn =-E[log(P) norm )];
[0015] In the formula, L syn Here, E is the expectation operator, and the negative sign indicates that the direction of loss minimization is consistent with the direction of confidence maximization.
[0016] Furthermore, the total loss function is:
[0017] L fine-tune =L task +β·L syn ;
[0018] In the formula, β is the balance coefficient, and L task Let L be the task loss function. syn For the collaborative loss function, L fine-tune The total loss function is defined as follows: the collaborative optimization module is used to: perform targeted gradient updates on the total loss function to obtain the gradient of the task loss and the gradient of the collaborative loss; update the parameters of the Adapter module and the discriminator based on the gradient of the task loss; and update the parameters of the reconstruction decoder and the unified encoder based on the gradient of the collaborative loss.
[0019] Furthermore, the collaborative optimization module is also used to determine whether a preset convergence condition has been reached after the total loss function has completed one backpropagation and parameter update; in response to reaching the preset convergence condition, the collaborative optimization process ends and the model optimization update is completed.
[0020] Furthermore, the targeted gradient update of the total loss function to obtain the gradient of the task loss and the gradient of the collaborative loss includes: calculating the difference between the discriminant head prediction result and the true label using the cross-entropy loss function; calculating the gradient of the task loss function with respect to the discriminant parameters and the adapter module parameters through automatic differentiation; and calculating the gradient of the collaborative loss function with respect to the parameters of the reconstruction decoder and the unified encoder through automatic differentiation.
[0021] Secondly, embodiments of the present invention provide an industrial vision inspection method for collaborative optimization of generation and discrimination. This method is applied to the aforementioned industrial vision inspection system for collaborative optimization of generation and discrimination. The method includes: Step S1, receiving an input image to be inspected; Step S2, extracting general visual features based on the image to be inspected; Step S3, performing zero-sample anomaly detection and adaptive threshold segmentation based on the general visual features to obtain a reconstructed image without defects; Step S4, generating a defect image containing specified defects based on a diffusion defect model using the general visual features and multimodal instructions; Step S5, recognizing the reconstructed image and outputting a confidence score based on the recognition result; and training the defect image generated by the conditional generation decoder to generate defect categories and category probabilities; Step S6, constructing a collaborative loss function based on the confidence score of the reconstructed image output by the discriminator, and obtaining a task loss function based on the category probabilities generated by the discriminator; Step S7, constructing a total loss function using the task loss function and the collaborative loss function, and updating the discriminator parameters, generation decoder parameters, and unified encoder parameters in reverse based on the gradient generated by the total loss function; Step S8, outputting the optimized and updated model.
[0022] In another aspect, the present invention also provides a computer-readable storage medium storing one or more instructions for causing the computer to execute the above-described industrial vision inspection method for co-optimization of generation and discrimination.
[0023] In another aspect, the present invention provides an electronic device, comprising: a memory and a processor; the memory storing at least one program instruction; the processor loading and executing the at least one program instruction to implement the above-mentioned industrial vision inspection method for coordinated optimization of generation and discrimination.
[0024] The beneficial effects of this invention are:
[0025] (1) Directed evolution of generative models: The reconstruction decoder no longer blindly learns reconstruction, but is guided by the discriminator to overcome the most difficult negative samples, making its understanding of "normal" patterns far exceed that of conventional models.
[0026] (2) Adversarial evolution of the discriminant model: The discriminant head faces a continuously evolving generator (reconstruction decoder) that specifically targets its weaknesses, which forces it to learn decision boundaries that are far more robust and refined than those learned through conventional training.
[0027] (3) The generation and discrimination capabilities have achieved a spiral-like co-evolution in the process of adversarial and collaborative processes. It is this non-obvious synergy driven by a specific mechanism that enables the present invention to outperform existing technologies (which treat pre-training, anomaly detection, data generation, model fine-tuning, etc. as independent modules and connect them in series) when dealing with complex industrial scenarios (especially in terms of low false alarm rate and ability to detect difficult defects). Attached Figure Description
[0028] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0029] Figure 1 This is a structural diagram of an industrial vision inspection system for collaborative optimization of generation and discrimination provided in Embodiment 1 of the present invention.
[0030] Figure 2 This is a schematic diagram of the workflow of a reconstruction decoder provided in Embodiment 1 of the present invention.
[0031] Figure 3 This is a schematic diagram of the workflow of a conditional generation decoder provided in Embodiment 1 of the present invention.
[0032] Figure 4 This is a schematic diagram of the workflow of a collaborative optimization module provided in Embodiment 1 of the present invention.
[0033] Figure 5 This is a flowchart of an industrial visual inspection method for collaborative optimization of generation and discrimination provided in Embodiment 2 of the present invention.
[0034] Figure 6 This is a partial block diagram of the electronic device provided in Embodiment 4 of the present invention. Detailed Implementation
[0035] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subroutine, etc.
[0036] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0037] The present invention will now be described in detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0038] Example 1
[0039] To facilitate understanding, the working principle of this system will be explained in general before describing the embodiments of the invention in detail: This embodiment aims to solve the problems of unidirectional and fragmented information flow, blind optimization of generation models, and passive acceptance of input by discrimination models in the prior art. It provides an industrial vision inspection system and method that coordinates the optimization of generation and discrimination, so as to achieve a highly versatile, low data dependency, and self-optimization effect for large industrial vision models. This invention adopts a "unified encoder + multi-functional task head" architecture. All tasks share a ViT encoder pre-trained on a large industrial dataset to extract general, high-quality visual features. The specific architecture is described as follows: The unified encoder is the core, providing high-quality features for all downstream tasks. The reconstruction decoder is used for zero-shot anomaly detection. It reconstructs a "perfect" image based on general features and completes anomaly detection by comparing the differences between the reconstructed image and the original image, without directly relying on the discrimination head. The conditional generation decoder generates images with specific defects on demand, solving the problem of insufficient small-sample or long-tail defect data. It is implemented by optimizing the diffusion model to ensure the authenticity and controllability of the generated defects. The discriminator is a lightweight network that activates during fine-tuning with small samples. It classifies and identifies initially detected anomalous regions, outputting a category and a confidence score. The output confidence score serves as a key feedback signal for collaborative optimization. A dynamic, bidirectional feedback loop is constructed, crucially quantifying the "uncertainty" of the discrimination task and transforming it into a direct, real-time optimization signal for the generation task. An adversarial collaborative loss (L...) is designed in this way. syn During training, if the discriminator gives a low "normal" confidence level to the "perfect" image region reconstructed by the reconstruction decoder, this invention captures this signal and allows the generative model to reinforce learning in that region, generating a higher fidelity reconstruction.
[0040] The specific implementation method is as follows:
[0041] like Figure 1The diagram shown is a structural diagram of an industrial vision inspection system for collaborative optimization of generation and discrimination provided by the present invention.
[0042] As an example, the system includes: a data input module 1, a model module 2, a collaborative optimization module 3, and an output module 4; the data input module 1 is used to receive an input image to be detected; the model module 2 includes a unified encoder 20, a reconstruction decoder 21, a conditional generation decoder 22, and a discriminator 23; the unified encoder 20 is used to extract general visual features based on the image to be detected; the reconstruction decoder 21 is used to perform zero-sample anomaly detection and adaptive threshold segmentation based on the general visual features to obtain a reconstructed image without defects; the conditional generation decoder 22 is used to generate a model containing specified defects based on the general visual features and multimodal instructions using a diffusion defect model. The defect image is used for identification; the discriminator 23 is used to identify the reconstructed image and output a confidence score based on the identification result; and the defect image generated by the conditional generator decoder 22 is used for training to generate defect categories and category probabilities; the collaborative optimization module 3 is used to construct a collaborative loss function based on the confidence score of the reconstructed image output by the discriminator 23, and obtain a task loss function based on the category probabilities generated by the discriminator 23; and a total loss function is constructed through the task loss function and the collaborative loss function, and the discriminator parameters, generator decoder parameters and unified encoder parameters are updated in reverse based on the gradient generated by the total loss function; the output module 4 is used to output the optimized and updated model.
[0043] In some feasible implementations, the unified encoder 20 also includes a general pre-training process. Specifically, it is self-supervised pre-trained on approximately 200,000 unlabeled industrial images of various types through a masked image modeling (MIM) task. The model reconstructs occluded regions based on partially visible regions of the image. This task forces the model to learn the intrinsic patterns of the image from low-level texture to high-level structure. The composite loss function is as follows:
[0044] L pretrain =L reconstruction +λL contrastive ;
[0045] L reconstruction (Reconstruction Loss): This is the main loss in the MIM task, typically using L1 or L2 loss to calculate the pixel differences between the reconstructed image and the original image in the occluded regions. It ensures the model has fine-grained image reconstruction capabilities. contrastive(Contrastive Learning Loss): This loss function enhances the model's ability to distinguish subtle differences between different samples. By treating images of the same object under different viewpoints or lighting conditions as positive sample pairs and images of different objects as negative sample pairs, it narrows the distance between positive sample pairs in the feature space and widens the distance between negative sample pairs, thereby improving the robustness and discriminative power of the features. λ is a hyperparameter used to balance the weights of the two losses.
[0046] In some feasible implementations, the unified encoder 20 includes a pre-trained ViT backbone network 201 and a trainable adapter module 202, which is inserted between the key layers of the unified encoder 20. Preferably, the unified encoder 20 employs a pre-trained ViT-L / 16 model (input image size 224×224).
[0047] In some feasible implementations, the reconstruction decoder 21 is also used to calculate the difference between the reconstructed image and the image to be detected based on a preset difference calculation formula to obtain a difference map; and to output abnormal regions based on the difference map using a Gaussian adaptive thresholding algorithm.
[0048] Preferably, the preset formula for calculating the degree of difference is:
[0049]
[0050] In the formula, I is the image to be detected, I ' To reconstruct the image, |II ' | Calculates the absolute difference at the pixel level. The gradient of the image to be detected. To reconstruct the gradient of the image, The term calculates the difference in image gradients, where α is a weighting coefficient used to balance the importance of pixel differences and gradient differences.
[0051] Specifically, in combination Figure 2 As shown, the workflow of the reconstruction decoder 21 includes:
[0052] Step 1: Input the image to be inspected (I), where the image to be inspected I includes: the image to be inspected collected in the industrial field (which may contain defects, such as scratches, dirt, etc.). This image does not need to be pre-labeled and does not need defect samples.
[0053] Step 2: Feature extraction using a unified encoder. This includes inputting the image into a pre-trained VIT encoder and extracting high-dimensional visual features.
[0054] Step 3: Generate a defect-free reconstructed image using the reconstruction decoder (I 'This includes inputting the high-dimensional visual features obtained in step 2 into the reconstruction decoder to generate a theoretically defect-free image. Specifically, the reconstruction decoder learns only the distribution of normal samples during pre-training. When a defective image is input, the reconstruction decoder will spontaneously repair the abnormal areas (e.g., fill in scratches into a smooth surface) and output an image of "ideal normal state".
[0055] Step 4: Based on the image to be detected I and the reconstructed image I ' Calculate the difference plot. This includes:
[0056] In the formula, |II ' This function calculates absolute differences at the pixel level, which can effectively detect areas with abnormal color or brightness. This term calculates the difference in image gradients (edges). It is introduced because many industrial defects (such as minor scratches and cracks) show slight variations at the pixel level but significant variations at the local gradient level. This enhances the model's sensitivity to structural defects. α is a weighting coefficient used to balance the importance of pixel differences and gradient differences.
[0057] Step 5: Adaptive Threshold Segmentation. This includes: applying a Gaussian adaptive thresholding algorithm to the difference map, dynamically calculating local thresholds to overcome uneven illumination (a common industrial pain point), outputting a binary mask, and identifying suspected abnormal areas.
[0058] Step 6: Output the abnormal areas. This includes the abnormal areas that are highlighted in the final output.
[0059] Compared to traditional supervised learning models that require a large number of defect sample annotations, the above-mentioned reconstruction decoder 21 uses zero samples and requires no defect data. Compared to traditional supervised learning models that are only applicable to a single product / scenario, this solution is highly versatile and can be directly applied across product lines. Compared to traditional supervised learning models that have difficulty detecting unknown defect types, this solution can detect any abnormal situation that deviates from the normal pattern. Compared to traditional supervised learning models that are susceptible to lighting interference, this solution uses gradient difference + adaptive threshold to have strong anti-interference capabilities.
[0060] In some feasible implementations, the conditional generation decoder 22 integrates a conditional encoding unit 220, a feature fusion unit 221, and a conditional generation unit 222. The conditional encoding unit 220 is used to convert the multimodal defect generation instructions provided by the user into a unified mathematical representation vector that the model can understand and recognize. The feature fusion unit 221 is used to fuse the unified mathematical representation vector with the features extracted from the input good product image by the unified encoder 20 using a cross-attention mechanism to generate fused features. The conditional generation unit 222 is used to generate the final defect image by using a diffusion model as the generator and the fused features as the core guide.
[0061] Preferred, combined Figure 3 As shown, the subsequent collaborative optimization module requires a small number of labeled defect samples to start. When real defect samples are scarce or insufficient, this embodiment can utilize controllable defect generation technology to synthesize high-fidelity defect data on demand, providing crucial data support for fine-tuning. Users can issue commands through various methods such as text, parameters, or masks. Specifically, its core technical process can be decomposed into three main steps: conditional encoding, feature fusion, and conditional generation.
[0062] The conditional encoding step involves converting the diverse and multimodal defect generation instructions provided by the user into a unified mathematical representation (conditional embedding vector C) that the model can understand. The diverse and multimodal defect generation instructions input by the user include: text control: for natural language instructions (e.g., generating a circular stain with a diameter of 3mm at the center of the headlight lens), this invention uses a pre-trained text encoder CLIP's Text Encoder to convert it into a semantic embedding vector C. text Parameter control: For structured parameter instructions (e.g., {type: "scratch", position: [x, y], length: 5}), the system converts them into a numerical parameter embedding vector C using a mapping function. param Mask control: For a binary mask M of the defect area directly drawn by the user, its spatial features are extracted through a small convolutional network (CNN) to generate a spatial embedding vector C. mask Depending on the type of input instruction, select or combine the corresponding embedding vectors as the final condition C.
[0063] Feature fusion includes the following step: The goal of this step is to combine the defect conditions C, which describe "where and what is generated," with the image content F, which describes "on which image it is generated." img To achieve effective integration. imgFeatures are extracted from the input good-quality image by a unified encoder. This embodiment preferably employs a cross-attention mechanism to achieve feature fusion. Its mathematical principle can be simplified as follows:
[0064] F fused =Attention(Q=F) img (K = C, V = C);
[0065] In this mechanism, each feature location (Query) of the original image focuses on different parts of the conditional vector (Key / Value), thereby adaptively injecting the attribute information of the defect into the most relevant image feature regions to generate the fused feature F. fused .
[0066] The condition generation includes: this step uses the fused features F fused Guided by the core, the final defect image is generated. This embodiment uses a diffusion model as the generator. The inverse denoising process of the diffusion model receives F at each step. fused As a conditional input, its single-step denoising process can be expressed as:
[0067] I t-1 =DenoiseNet(I t ,t,C guidance =F fused );
[0068] Among them, I t Let I be the noisy image at step t, where t is the time step. DenoiseNet is a noise prediction network with a U-Net structure. The model starts from a pure Gaussian noise image I. T Beginning, in F fused Under continuous guidance, noise is gradually removed, and finally an image I0 with accurate, high-fidelity defects is recovered.
[0069] In some feasible implementations, after obtaining a small number of labeled samples (whether real or generated by a conditional generation decoder), this embodiment does not simply fine-tune the model, but initiates a sophisticated dual-path mechanism that combines "fine-tuning" and "optimization" in parallel, fundamentally improving the model's generation and discrimination capabilities while adapting to new scenarios.
[0070] Preferably, the collaborative optimization module 3 constructs a collaborative loss function based on the confidence level of the reconstructed image output by the discriminator, comprising: the discriminator scanning the reconstructed image and outputting a pixel-by-pixel confidence map P. norm Based on the pixel-by-pixel confidence map P normThe low-confidence region is located by comparing it with a preset confidence threshold; a collaborative loss function is constructed for the low-confidence region, the collaborative loss function including: L syn =-E[log(P) norm In the formula, L syn Here, E is the expectation operator, and the negative sign indicates that the direction of loss minimization is consistent with the direction of confidence maximization.
[0071] Preferably, the total loss function is: L fine-tune =L task +β·L syn ;
[0072] In the formula, β is the balance coefficient, and L task Let L be the task loss function. syn For the collaborative loss function, L fine-tune The total loss function is defined as follows: the collaborative optimization module 3 is used to: perform targeted gradient updates on the total loss function to obtain the gradient of the task loss and the gradient of the collaborative loss; update the parameters of the Adapter module and the discriminator based on the gradient of the task loss; and update the parameters of the reconstruction decoder and the unified encoder based on the gradient of the collaborative loss.
[0073] Specifically, in combination Figure 4 As shown, the overall workflow of collaborative optimization module 3 is as follows:
[0074] Core components and settings: Before fine-tuning begins, the system is configured as follows:
[0075] Unified Encoder (Parameter Freeze): To preserve the general feature extraction capabilities gained from pre-training on massive datasets, the encoder's backbone parameters are frozen during fine-tuning. Adapter Module (Trainable): Lightweight adapter modules are inserted between key layers of the encoder. During fine-tuning, the parameters of these modules are primarily updated, allowing the model to efficiently adapt to new tasks with extremely low computational cost. Discriminator Head (Trainable): Responsible for classifying or locating anomalies; its parameters are updated along with the adapter modules during fine-tuning. Reconstruction Decoder (Trainable): Responsible for image reconstruction; its parameters are specifically updated by the co-optimization path.
[0076] Detailed Explanation of Collaborative Optimization with Two Paths:
[0077] The fine-tuning process involves two parallel execution paths:
[0078] Path 1: Supervised Fine-tuning Path Based on Labeled Samples Figure 4 The process on the right (including the following):
[0079] Data input: Input a small number of pre-processed (e.g., image enhancement) labeled samples into the system.
[0080] Forward propagation: The sample is processed by a "frozen encoder + trainable adapter module" to extract features suitable for the current task, and then fed into the discriminator for defect classification or localization.
[0081] Calculate the task loss: Compare the prediction results of the discriminant head with the true labels of the samples to calculate the task loss (L). task The loss is typically cross-entropy loss. This loss directly measures the model's performance on the target task.
[0082] Path 2: Generation-Discrimination Co-optimization Path Based on Normal Samples ( Figure 4 The process on the left includes:
[0083] This path is an unsupervised adversarial loop whose core objective is not to directly detect industrial defects, but to use the discriminator as a high-quality evaluator to inversely improve the performance of the reconstruction decoder. By forcing the reconstruction decoder to generate a perfect image sufficient to "fool" the discriminator, it indirectly enhances the encoder's deep understanding of normal sample patterns.
[0084] Generating reconstructed samples and using a discriminant head to locate "generated defects": A batch of images is selected from unlabeled normal samples (or normal regions within labeled samples). norm The encoder and the reconstruction decoder generate the corresponding reconstructed image I'. norm =G recon (Encoder(I norm These reconstructed images I' norm Theoretically, these are perfect "negative samples" (defect-free samples), but they often contain flaws in the early stages of training. These reconstructed images are then fed into a discriminator. Here, the discriminator's role is completely different from the anomaly detection process: it doesn't look for industrial defects, but rather for reconstruction flaws within the generative model itself. Because the discriminator has learned the features of real defects in path one, it is extremely sensitive to any region deviating from the "perfectly normal" pattern (i.e., reconstruction flaws), outputting a low "normal" confidence score. These low-confidence regions precisely expose the weaknesses of the generative model.
[0085] Constructing the collaborative loss: Based on the "reconstruction flaws" (i.e., low-confidence regions) identified by the discriminant head, construct the collaborative loss L. syn The loss is designed to maximize the confidence level of the discriminant head, i.e., L. syn =-E[log(P) norm The goal of this loss function is to force the reconstruction decoder to repair the flaws identified by the discriminator in order to generate a more realistic normal image.
[0086] The dual-path gradient update and bidirectional feedback closed-loop process includes: the loss functions of the two paths are eventually integrated, but their gradients "intelligently" flow to different parts of the network during backpropagation, which is the key to achieving collaborative optimization.
[0087] Total Loss Calculation: The total loss function integrates task loss and collaboration loss: L fine-tune =L task +β·L syn , where β is the balance coefficient.
[0088] Targeted gradient update:
[0089] Gradient of task loss Its gradients only update the parameters of the Adapter module and the discriminator head. The goal of this update path is to "teach" the discriminator how to perform the labeling task more accurately.
[0090] gradient of collaborative loss Its gradients only update the parameters of the reconstruction decoder and the shared encoder (including its Adapter module). The goal of this update path is to "force" the generator (reconstruction decoder and encoder) to fix its reconstruction flaws in order to "fool" the discriminator.
[0091] Preferably, the collaborative optimization module 3 is further configured to determine whether a preset convergence condition has been met after the total loss function completes one backpropagation and parameter update; in response to meeting the preset convergence condition, the collaborative optimization process ends, and the model optimization update is completed. Specifically, the convergence condition determination includes: in each iteration (i.e., the total loss L... fine-tune After each backpropagation and parameter update, the system performs a judgment. Convergence conditions are pre-defined and typically include one or more of the following: reaching the maximum number of iterations: to prevent infinite training, an upper limit is set for the number of training cycles; loss function convergence: within multiple consecutive iteration cycles, the total loss L... fine-tune Or its loss on the validation set no longer decreases significantly; key performance indicators saturate: the model's accuracy, recall, or F1 score on the validation set does not improve for several consecutive periods.
[0092] Completion and Deployment: Once any convergence condition is met ("Yes"), the entire collaborative optimization and fine-tuning process is complete. The resulting model possesses both adaptability to new tasks and robustness derived from adversarial optimization. It can be saved and deployed to the target industrial scenario for rapid and efficient adaptation. If the condition is not met ("No"), the process returns to execute the next iteration.
[0093] Preferably, the targeted gradient update of the total loss function to obtain the gradient of the task loss and the gradient of the collaborative loss includes: calculating the difference between the discriminant head prediction result and the true label using the cross-entropy loss function; calculating the gradient of the task loss function with respect to the discriminant parameters and the adapter module parameters using automatic differentiation; and calculating the gradient of the collaborative loss function with respect to the parameters of the reconstruction decoder and the unified encoder using automatic differentiation.
[0094] In this embodiment, a cooperative loss L is introduced. syn Compared to traditional generator loss (such as MSE), which optimizes all regions on average and is insensitive to critical defects, the collaborative loss L in this embodiment... syn By penalizing only low-confidence regions, the system effectively addresses specific weaknesses, achieving precise and targeted optimization. Furthermore, the technical solution described in this embodiment continuously attacks the weaknesses of the discrimination head, forcing the generator to learn textures that the discrimination head easily confuses (such as brushed metal vs. scratches) and structures that the discrimination head easily misses (such as cracks within transparent materials).
[0095] To facilitate understanding, a specific example is provided here. Taking "defect detection on the surface of mobile phone screen glass" as an example, and combining the implementation steps of the above invention, the specific application process is demonstrated:
[0096] Scenario: A mobile phone factory needs to inspect three types of defects on the surface of its screen glass: scratches, bubbles, and dirt, and determine the "normal" state. The production line currently has 50 labeled samples (10 for each type of defect + 20 for each type of normal sample), and a real-time inspection system (requiring at least 30fps) needs to be deployed quickly.
[0097] Example of implementation steps:
[0098] 1. System Environment and Parameter Configuration: Hardware: NVIDIA Jetson AGX edge device (compatible with model deployment), training phase uses a server equipped with 8×NVIDIA A100 processors. Core Parameters: Unified Encoder: ViT-L / 16 (input 224×224, frozen after fine-tuning on the mobile phone glass dataset). Discriminator Head: Output channels = 3 (defects) + 1 (normal) = 4. Hyperparameters: α = 0.3, λ = 0.1, β = 0.5 (using general configuration).
[0099] 2. Phase One: General Pre-training: Dataset: Collect 100,000 unlabeled industrial images (including mobile phone glass, computer screens, watch faces, etc.), uniformly adjust to 224×224, and randomly crop to 192×192 to enhance data diversity. Training: Through masked image modeling (MIM), 50% of the area of each image is occluded, and the reconstruction decoder predicts the occluded part. Train for 100 epochs using L1 loss (reconstruction) + InfoNCE loss (contrast) to enable the model to master the general visual features of glass surfaces (such as smoothness and reflectivity).
[0100] 3. Phase Two: Zero-Sample Anomaly Detection (Rapid Trial): Detection Target: New batch of unlabeled mobile phone glass images (resolution 1280×720). Process: Preprocessing: After denoising, the image is cropped into a 224×224 sub-image. Reconstruction: The encoder extracts features, and the decoder generates a "perfect glass" reconstructed image (automatically repairing scratches, bubbles, and other anomalies). Anomaly Localization: The difference between the original image and the reconstructed image (including pixel differences and gradient differences, weight α = 0.3) is calculated, and a heatmap is generated. For example: A 1mm scratch appears in the upper right corner of an image. In the difference map, the value of this area is significantly higher than the threshold (μ+2σ), and it is marked as an anomaly area "(180,50)-(220,60), confidence level 0.91".
[0101] 4. Phase Three: Small Sample Fine-tuning (Precise Classification): Objective: To distinguish between "scratches," "bubbles," and "dirt," using only 50 labeled samples. Data Augmentation (Controllable Generation): Input instructions: "Generate a 0.5mm diameter bubble in the center of the glass," "Generate a 2mm long gray scratch at the edge," generating 500 synthetic samples (all PSNR > 30), which are then mixed with real samples. Collaborative Fine-tuning: Supervised Path: Train the discriminator using labeled samples, and apply cross-entropy loss (L... task Learn defect category features and update the Adapter module (without freezing the encoder backbone). Generate-discriminate collaboration: Input 1000 normal glass images, let the reconstruction decoder generate a "normal" reconstruction map, and the discriminator head needs to output a high "normal" confidence level, through collaborative loss (L... syn Optimize the generator to reduce false positives. Training stops: After 35 epochs, training stops when the validation set F1 score stabilizes at 0.96.
[0102] 5. Actual deployment and testing results:
[0103] Deployment: The model is exported in ONNX format and run on Jetson AGX, achieving an inference speed of 35fps, meeting real-time requirements. Detection Process: Production line cameras acquire images in real time, which are then cropped and input into the model. Output Results: For example, "Region (120,80)-(150,90): Scratch, confidence level 0.96", "Region (30,30)-(45,45): Bubble, confidence level 0.93". Anomaly information is uploaded to the production line system in real time, triggering a sorting robot arm to remove defective products. Results: After fine-tuning with a small sample size, the defect classification accuracy is 97.2%, the detection rate of 0.1mm fine scratches is 93.8%, and the false alarm rate is only 2.5%. Through the above process, this system achieves efficient detection of defects on the surface of mobile phone glass in low-label data scenarios, balancing speed and accuracy.
[0104] In summary, the above-described embodiments have the following technical effects:
[0105] (1) No need for large-scale labeled data: Phase 1 is pre-trained with unlabeled images, Phase 2 can directly detect unseen defects in zero-sample mode, and Phase 3 only requires 50 labeled samples to achieve accurate classification, which greatly reduces the cost of manual labeling in industrial scenarios (especially suitable for scenarios where defect samples are scarce).
[0106] (2) Controllable generation of supplementary data: A large number of realistic defect samples (such as scratches at specified locations and sizes) are synthesized through a conditional generation decoder to solve the problem of insufficient real defect samples.
[0107] (3) Excellent performance in zero-sample mode: For new products that have not been fine-tuned (such as different models of glass), the average detection rate is over 92%, and the false alarm rate is low (only 2.5% in the example), which is better than traditional methods.
[0108] (4) Strong ability to identify minute defects: The detection rate of minute scratches <0.1mm exceeds 93%, which can meet the needs of high-precision industrial inspection (such as minute defects on the surface of electronic components).
[0109] (5) High versatility: The unified encoder is pre-trained on multiple categories of industrial images and can be transferred to different scenarios such as electronic components and metal parts, without the need to train a basic model separately for each category.
[0110] (6) Targeted optimization: Through small sample fine-tuning, it can quickly adapt to the defect types (scratches, bubbles, etc.) of specific products (such as mobile phone glass) to achieve accurate classification.
[0111] (7) Fast inference speed: After the model is deployed to an edge device (such as NVIDIA Jetson AGX), the inference speed reaches more than 30fps, which meets the real-time detection requirements of the production line (without affecting the production cycle).
[0112] (8) Flexible deployment: Supports export in ONNX format, can be adapted to a variety of industrial edge devices, and is easy to integrate into existing production line systems.
[0113] (9) Generation and discrimination collaboration: Through the collaborative loss of “reconstruction decoder to generate normal image + discrimination head verification”, the model is forced to improve its ability to distinguish between “normal” and “abnormal”, and reduce misjudgment caused by changes in lighting and texture.
[0114] (10) Dynamic adaptability: During the detection process, the model can automatically repair abnormal areas to generate a "perfect reconstruction map" and accurately locate defects through difference comparison, without being disturbed by the complex texture of the product surface.
[0115] In summary, this implementation method balances the core requirements of "low cost, high precision, high speed, and wide adaptability" in industrial inspection, and is especially suitable for industrial scenarios with limited labeled data and diverse defect types.
[0116] It is worth mentioning that all modules involved in this embodiment are logical units. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this invention, this embodiment does not introduce units that are not closely related to solving the technical problem proposed by this invention; however, this does not mean that other units are absent from this embodiment.
[0117] Example 2
[0118] Please see Figure 5 The above is a flowchart of an industrial visual inspection method for collaborative optimization of generation and discrimination, provided by an embodiment of the present invention.
[0119] As an example, the method is applied to the industrial vision inspection system with coordinated optimization of generation and discrimination described in Example 1, and the method includes:
[0120] Step S1: Receive the input image to be detected.
[0121] Step S2: Extract general visual features based on the image to be detected.
[0122] Step S3: Perform zero-sample anomaly detection and adaptive threshold segmentation based on the general visual features to obtain a reconstructed image that does not contain defects.
[0123] Step S4: Based on the general visual features and multimodal instructions, generate a defect image containing the specified defect using a diffusion defect model.
[0124] Step S5: Recognize the reconstructed image and output the confidence level based on the recognition result; and train the defect image generated by the conditional generation decoder to generate defect categories and category probabilities.
[0125] Step S6: Construct a collaborative loss function based on the confidence of the reconstructed image output by the discriminator, and obtain a task loss function based on the class probabilities generated by the discriminator.
[0126] Step S7: Construct the total loss function using the task loss function and the collaborative loss function, and update the discriminant head parameters, generator decoder parameters, and unified encoder parameters in reverse based on the gradient generated by the total loss function.
[0127] Step S8: Output the optimized and updated model.
[0128] It is not difficult to see that this embodiment is a method embodiment corresponding to the first embodiment, and this embodiment can be implemented in conjunction with the first embodiment. The relevant technical details mentioned in the first embodiment are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the first embodiment.
[0129] Example 3
[0130] This invention also proposes a storage medium storing a collaboratively optimized industrial vision inspection method for generation and discrimination. When the program for generating and discriminating collaboratively optimized industrial vision inspection is executed by a processor, it implements the steps of the collaboratively optimized industrial vision inspection method as described above. Since this storage medium employs all the technical solutions of the above embodiments, it possesses at least all the beneficial effects brought about by the technical solutions of the above embodiments, which will not be elaborated upon further here.
[0131] Example 4
[0132] Please see Figure 6 The present invention also provides an electronic device, including: a memory and a processor; the memory stores at least one program instruction; the processor loads and executes the at least one program instruction to implement the industrial vision inspection method for generation and discrimination co-optimization provided in Embodiment 2.
[0133] The memory 702 and processor 701 are connected via a bus, which may include any number of interconnecting buses and bridges, connecting various circuits of one or more processors 701 and memory 702 together. The bus may also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. A bus interface provides an interface between the bus and the transceiver. The transceiver may be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 701 is transmitted over a wireless medium via an antenna, which further receives data and transmits it to processor 701.
[0134] Processor 701 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory 702 can be used to store data used by processor 701 during operation.
[0135] The above descriptions are merely embodiments of the present invention. Commonly known structures and characteristics are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, based on the guidance provided in this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.
Claims
1. An industrial vision inspection system with coordinated optimization of generation and discrimination, characterized in that, The system includes: a data input module, a model module, a collaborative optimization module, and an output module; The data input module is used to receive the input image to be detected; The model module includes a unified encoder, a reconstruction decoder, a conditional generation decoder, and a discriminant head; The unified encoder is used to extract common visual features based on the image to be detected; The reconstruction decoder is used to perform zero-sample anomaly detection and adaptive threshold segmentation based on the general visual features to obtain a reconstructed image without defects. The conditional generation decoder is used to generate a defect image containing a specified defect based on the general visual features and a diffusion defect model on the basis of multimodal instructions. The discriminator is used to identify the reconstructed image and output a confidence score based on the identification result; and to train the defect image generated by the conditional generation decoder to generate defect categories and category probabilities. The collaborative optimization module is used to construct a collaborative loss function based on the confidence of the reconstructed image output by the discriminator, and to obtain a task loss function based on the class probabilities generated by the discriminator; and The total loss function is constructed by the task loss function and the collaborative loss function, and the discriminant head parameters, generator decoder parameters and unified encoder parameters are updated in reverse based on the gradient generated by the total loss function. The output module is used to output the optimized and updated model.
2. The industrial vision inspection system for collaborative optimization of generation and discrimination according to claim 1, characterized in that, The unified encoder includes a pre-trained ViT backbone network and a trainable adapter module, which is inserted between the key layers of the unified encoder.
3. The industrial vision inspection system with coordinated optimization of generation and discrimination according to claim 1, characterized in that, The reconstruction decoder is also used to calculate the difference between the reconstructed image and the image to be detected based on a preset difference calculation formula to obtain a difference map. Based on the difference map, an abnormal region is output using a Gaussian adaptive threshold algorithm.
4. The industrial vision inspection system with coordinated optimization of generation and discrimination according to claim 3, characterized in that, The preset formula for calculating the degree of difference is: In the formula, I is the image to be detected, I ' To reconstruct the image, |II ' | Calculates the absolute difference at the pixel level. The gradient of the image to be detected. To reconstruct the gradient of the image, The term calculates the difference in image gradients, where α is a weighting coefficient used to balance the importance of pixel differences and gradient differences.
5. The industrial vision inspection system with coordinated optimization of generation and discrimination according to claim 1, characterized in that, The conditional generation decoder integrates a conditional coding unit, a feature fusion unit, and a conditional generation unit. The conditional coding unit is used to convert the multimodal defect generation instructions provided by the user into a unified mathematical representation vector that the model can understand and recognize. The feature fusion unit is used to fuse the unified mathematical representation vector with the features extracted from the input good image by the unified encoder using a cross-attention mechanism to generate fused features; The conditional generation unit is used to generate the final defect image by using a diffusion model as a generator and the fused features as the core guide.
6. The industrial vision inspection system with coordinated optimization of generation and discrimination according to claim 1, characterized in that, The collaborative optimization module is used to construct a collaborative loss function based on the confidence of the reconstructed image output by the discriminator, including: The discriminant head scans the reconstructed image and outputs a pixel-by-pixel confidence map P. norm ; Based on the pixel-by-pixel confidence map P norm The low-confidence region is located by comparing it with a preset confidence threshold; A collaborative loss function is constructed for the low-confidence region, and the collaborative loss function includes: L syn = -E[log(P norm )]; In the formula, L syn Here, E is the expectation operator, and the negative sign indicates that the direction of loss minimization is consistent with the direction of confidence maximization.
7. The industrial vision inspection system for collaborative optimization of generation and discrimination according to claim 1, characterized in that, The total loss function is: L fine-tune L task +β·L syn ; In the formula, β is the balance coefficient, and L task Let L be the task loss function. syn For the collaborative loss function, L fine-tune This is the total loss function; The collaborative optimization module is used for: The gradient of the task loss and the gradient of the collaborative loss are obtained by performing targeted gradient update on the total loss function; The parameters of the Adapter module and the discriminant head are updated based on the gradient of the task loss. The parameters of the reconstruction decoder and the unified encoder are updated based on the gradient of the cooperative loss.
8. The industrial vision inspection system with coordinated optimization of generation and discrimination according to claim 7, characterized in that, The collaborative optimization module is also used to determine whether the preset convergence condition has been met after the total loss function has completed one backpropagation and parameter update. Upon reaching the preset convergence condition, the collaborative optimization process ends, and the model is optimized and updated.
9. The industrial vision inspection system with coordinated optimization of generation and discrimination according to claim 7, characterized in that, The targeted gradient update of the total loss function to obtain the gradient of the task loss and the gradient of the collaborative loss includes: The difference between the discriminant head prediction and the true label is calculated using the cross-entropy loss function. The gradient of the task loss function with respect to the discriminant parameters and the Adapter module parameters is calculated using automatic differentiation; The gradient of the collaborative loss function with respect to the parameters of the reconstruction decoder and the unified encoder is calculated using automatic differentiation.
10. An industrial vision inspection method for co-optimization of generation and discrimination, the method being applied to the industrial vision inspection system for co-optimization of generation and discrimination as described in any one of claims 1-9, characterized in that, The method includes: Step S1: Receive the input image to be detected; Step S2: Extract general visual features based on the image to be detected; Step S3: Perform zero-sample anomaly detection and adaptive threshold segmentation based on the general visual features to obtain a reconstructed image that does not contain defects; Step S4: Based on the general visual features and multimodal instructions, generate a defect image containing the specified defect using a diffusion defect model; Step S5: Recognize the reconstructed image and output the confidence score based on the recognition result; and train the defect image generated by the conditional generation decoder to generate defect categories and category probabilities. Step S6: Construct a collaborative loss function based on the confidence of the reconstructed image output by the discriminator, and obtain a task loss function based on the class probabilities generated by the discriminator; Step S7: Construct the total loss function through the task loss function and the collaborative loss function, and update the discriminant head parameters, generator decoder parameters and unified encoder parameters in reverse based on the gradient generated by the total loss function; Step S8: Output the optimized and updated model.
Citation Information
Cited By
Urban safety abnormal event image generation method, system, equipment and medium
CN121392471A
Intelligent perception and collaborative optimization platform of industrial internet of things
CN121809985A