Zero-sample-driven dialogue type industrial defect detection system and method

Through a zero-sample-driven dialogue-based industrial defect detection system, the pre-trained model and object-independent prompt design is used to solve the problem of scarcity of samples and weak generalization capabilities in industrial defect detection, and achieve high-precision and strong generalization capabilities of industrial defect detection.

CN120107190APending Publication Date: 2025-06-06NEW TECH APPL INST BEIJING CITY

Patent Information

Application Number
CN202510167547.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-16
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing technology faces the problems of scarcity, high diversity and unpredictability in industrial defect detection, which leads to limited model generalization capabilities and is difficult to meet the needs of complex industrial environments.

Method used

A zero-sample-driven dialogue-based industrial defect detection system is adopted, which includes a query image encoding module, an object-independent text prompt module, a prompt learner module and a defect recognition dialogue module. Industrial defect detection is realized through knowledge transfer of pre-trained models, object-independent prompt design and dynamic feature fusion mechanism.

Benefits of technology

The system can complete industrial defect detection without the need for target scenario training data, solving the problem of traditional methods relying on labeled data and weak generalization capabilities. It has high precision and strong generalization capabilities, and can support multi-type defect detection in complex industrial environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107190A_ABST
    Figure CN120107190A_ABST
Patent Text Reader

Abstract

The invention discloses a zero-sample-driven dialogue type industrial defect detection system and method, and aims to solve the problems of sample scarcity, weak model generalization ability and the like in industrial defect detection. The system comprises a query image coding module, an object-independent text prompt module, a prompt learner module and a defect identification dialogue module. The industrial defect detection without target scene training data is realized by utilizing the prior knowledge of the pre-training model and combining the prompt design irrelevant to the object and the dynamic feature fusion mechanism. By utilizing the method, the industrial defect detection efficiency and accuracy are improved, natural language interaction is supported, and the method is particularly suitable for multi-type defect detection in a complex industrial environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a zero-sample driven conversational industrial defect detection system and also to a corresponding conversational industrial defect detection method, belonging to the technical field of industrial defect detection. Background Art

[0002] Industrial defect detection technology based on deep learning can significantly reduce the cost of manual quality inspection by learning defect features from labeled samples. However, such methods face fundamental limitations in actual industrial scenarios: first, industrial defect samples are scarce, highly diverse, and unpredictable, resulting in insufficient training data; second, existing models need to make a difficult trade-off between high precision (reducing false positives) and high recall (reducing missed detections). The lack of samples limits the generalization ability of the model, making it difficult to meet the needs of complex industrial environments.

[0003] To address the problem of sample scarcity, an anomaly detection method based on an unsupervised generative adversarial network (GAN) was proposed. It increases the sample size by synthesizing defect data, thereby reducing the reliance on labeled data. However, this method has obvious defects: the quality of generated samples directly affects the performance of the model. If the synthetic data deviates greatly from the actual defect distribution, the model will learn non-critical features, which will reduce the detection accuracy.

[0004] In recent years, multimodal methods based on large language models have been tried to be applied to industrial inspection, but the actual effect is still not ideal. The fundamental reason is that large language models tend to understand the overall object of the image through text descriptions, while ignoring local subtle anomalies (such as micron-level cracks and hidden scratches); in addition, industrial defect samples are difficult to obtain and some involve confidentiality agreements, resulting in a lack of sufficient defect prior knowledge in the pre-training stage of large language models, making it difficult to establish accurate abnormal attribute associations.

[0005] In the prior art, a Chinese invention patent with the authorization announcement number CN116777906B proposed an anomaly detection scheme based on good product image generation. This scheme constructs a multi-scale data set by randomly cropping good product images, combines the diffusion model (DDPM) with the super-resolution model to repair the image, and finally realizes defect recognition through similarity comparison. However, the core limitation of this scheme is that it only optimizes image clarity, and does not solve key problems such as insufficient diversity of defect samples and weak ability to extract complex defect features. It cannot meet the needs of accurate positioning and classification of multiple types of defects in industrial scenarios. Summary of the invention

[0006] The primary technical problem to be solved by the present invention is to provide a zero-sample driven conversational industrial defect detection system.

[0007] Another technical problem to be solved by the present invention is to provide a zero-sample driven conversational industrial defect detection method.

[0008] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:

[0009] According to a first aspect of an embodiment of the present invention, there is provided a zero-sample driven conversational industrial defect detection system, comprising a query image encoding module, an object-independent text prompt module, a prompt learner module, and a defect recognition dialogue module;

[0010] The query image encoding module encodes and embeds features of the query image to generate two branches: the first branch inputs visual features into the multimodal large language model in the defect recognition dialogue module; the second branch integrates the output of the object-independent text prompt module;

[0011] The text prompt module generates text prompts, which are encoded and superimposed with image features to generate a heat map; the heat map is processed by a convolutional neural network in the prompt learner module, and is aligned and connected with learnable prompt parameters and output features of a frozen pre-trained model, and finally input into a multimodal large language model to complete multimodal reasoning;

[0012] The defect identification dialogue module sends dialogue data to the multimodal large language model, and the multimodal large language model returns the detection result to the user.

[0013] Preferably, the object-independent text prompt module includes a normal sample description template and a defect sample description template;

[0014] Among them, the text encoding results of the normal sample description template and the defect sample description template are fused through a feature addition operation to generate a joint text representation, which is then mapped into a low-dimensional vector through an embedding layer for alignment with image features.

[0015] Preferably, the prompt learner module includes a convolutional neural network and a learnable prompt learner;

[0016] The heat map is input into the convolutional neural network to extract spatial features, and is spliced ​​with the global feature vector and dynamic prompt vector of the query image in channel dimension to form a unified multimodal feature; the aligned features are reduced in dimension and normalized by the fully connected layer and input into the multimodal large language model.

[0017] According to a second aspect of an embodiment of the present invention, a zero-sample driven conversational industrial defect detection method is provided, comprising the following steps:

[0018] S1: Create an auxiliary data sample library based on the defect dataset in the industrial field and generate auxiliary training samples;

[0019] S2: Build a large multimodal language model in the training phase, freeze the network weights of the model, and only allow the prompt learner module and the object-independent text prompt module to participate in the training;

[0020] S3: Filter samples containing defects from the auxiliary data sample library as the defect detection auxiliary sample set for fine-tuning the multimodal large language model;

[0021] S4: Screening defect-free normal samples from the auxiliary data sample library as the defect-free detection auxiliary sample set for fine-tuning the multimodal large language model;

[0022] S5: Self-collected data covering different background conditions, industrial sample types and defect categories are used as auxiliary sample sets for zero-shot detection to verify the generalization ability of the multimodal large language model under target-free scene training data;

[0023] S6: Setting an object-independent text prompt module, which includes a normal sample description template and a defect sample description template; the text prompt module generates structured text prompt data;

[0024] S7: Input the query image to the query image encoding module, extract image features and vectorize them, and divide them into two branches: the first branch directly inputs the vectorized features into the multimodal large language model, and the second branch generates a heat map by aligning the vectorized features with the text prompt data;

[0025] S8: Input the heat map to the prompt learner module, which receives the spatial features extracted by the convolutional neural network, the global image feature vector output by the query image encoder, and the dynamic prompt vector generated by the prompt learner; realize multimodal feature fusion through channel dimension splicing, and input the multimodal large language model after normalization by the fully connected layer;

[0026] S9: The output features of the prompt learner module and the questions submitted by the user through the defect identification dialogue module are input into the multimodal large language model to generate a dialogue result including defect type, location and explanation.

[0027] Preferably, constructing a multimodal large language model in the training phase includes the following steps:

[0028] First, the output result of the text prompt module after text encoding and text embedding is superimposed with the auxiliary training sample after image encoding and image embedding to obtain a multimodal model;

[0029] Secondly, the multimodal model is processed by the prompt learner module and input into the large language inference model to obtain the multimodal large language model.

[0030] Preferably, in step S2, ViT-L / 14@336px is selected as the image encoder pre-training model and Vicuna-7b is selected as the large language pre-training model.

[0031] Preferably, setting the object-independent text prompt module includes the following sub-steps:

[0032] First, set the normal sample description template and the abnormal sample description template; second, define the geometric list of normal state, abnormal state and defect type; finally, set the object category and embed the state description to indicate the state of the object.

[0033] Preferably, the normal sample description template includes: prefix description+[object placeholder]+[normal status word]; the defect sample description template includes: prefix description+[object placeholder]+[defect type word]+[abnormal status word].

[0034] Preferably, in step S7, extracting image features includes the following sub-steps:

[0035] S71: Select the Vision Transformer deep learning model containing 4 layers of encoders as the base model for image visual feature extraction, extract the visual features of the image layer by layer in a bottom-up manner, and retain the output features of each layer of encoders;

[0036] S72: Align the visual features obtained by each layer of encoder with the text feature vector, and then merge the aligned visual features and text feature vector to obtain the final vectorized image features.

[0037] Preferably, in step S8, inputting the heat map into the prompt learner module comprises the following sub-steps:

[0038] The heat map first undergoes a convolution operation to extract local features in the image, then undergoes a nonlinear transformation through an activation function, then undergoes feature dimensionality reduction and feature mapping through a maximum pooling operation, and finally undergoes another convolution operation to further extract features and generate an embedded representation.

[0039] Compared with the prior art, the present invention achieves industrial defect detection without target scene training data by utilizing the prior knowledge of pre-trained models, object-independent prompt design, and dynamic feature fusion mechanism. The system uses frozen pre-trained models (such as ViT-L / 14@336px and Vicuna-7b) and only trains the object-independent text prompt module and prompt learner module, so that no re-training is required when reasoning with a single image. The present invention not only solves the problem that traditional methods rely on labeled data and have weak generalization capabilities, but also has high precision and strong generalization capabilities, and can support multi-type defect detection in complex industrial environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 A schematic diagram of the structure of a zero-sample driven conversational industrial defect detection system provided by the first embodiment of the present invention;

[0041] Figure 2 A simplified flowchart of a zero-sample driven conversational industrial defect detection method provided in the second embodiment of the present invention;

[0042] Figure 3 A schematic diagram of text encoding in an embodiment of the present invention;

[0043] Figure 4 A schematic diagram of multi-scale feature extraction in an embodiment of the present invention;

[0044] Figure 5 Schematic diagram of the principle of a prompt learner in an embodiment of the present invention. DETAILED DESCRIPTION

[0045] The technical content of the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] First embodiment

[0047] like Figure 1As shown, a zero-sample driven conversational industrial defect detection system provided by the first embodiment of the present invention at least includes: a query image encoding module, an object-independent text prompt module, a prompt learner module and a defect recognition dialogue module. Among them, the system first encodes and embeds the query image through the query image encoding module to generate two branches: the first branch directly inputs the visual features into the multimodal large language model in the defect recognition dialogue module; the second branch is fused with the output of the object-independent text prompt module. The text prompt module generates text prompts, which are superimposed with image features after text encoding to generate a heat map. Subsequently, the heat map is processed by the convolutional neural network in the prompt learner module, and is aligned and connected with the learnable prompt parameters and the output features of the frozen pre-trained model, and finally input into the multimodal large language model to complete multimodal reasoning. For example, the user can interact through the defect recognition dialogue module (such as asking questions such as "Is this a photo of a railroad track? Is there any abnormality in the picture?"), and the system returns the detection result (such as "there is a defect"). The entire conversational industrial defect detection system is based on a zero-sample driven design. It does not require target scene training data. It achieves high-precision defect location and classification by integrating visual features and text prompts. It also supports natural language interaction and is suitable for real-time detection needs in complex industrial environments.

[0048] In one embodiment of the present invention, the object-independent text prompt module includes a normal sample description template and a defect sample description template, and the two types of templates generate prompt information through structured text. Among them, the normal sample description template adopts the form of "prefix description + [object placeholder] + [normal status word]", where the status word set is {good, perfect, flawless}; the defect sample description template is extended to "prefix description + [object placeholder] + [defect type word] + [abnormal status word]", where the defect type can include {blowhole, break, crack}, etc., and the abnormal status word can include {broken, defect, flaw}, etc. The text encoding results of the two types of templates are fused through a feature addition operation to generate a joint text representation, which is then mapped to a low-dimensional vector through an embedding layer for alignment with image features.

[0049] In one embodiment of the present invention, the prompt learner module is composed of a convolutional neural network (CNN) and a learnable prompt learner, and its workflow is as follows: the heat map is input into the CNN to extract spatial features, and the channel dimension is spliced ​​with the global feature vector of the query image (from the image encoder) and the dynamic prompt vector (generated by the learnable parameters) to form a unified multimodal feature; the aligned features are reduced in dimension and normalized by the fully connected layer, and finally input into the multimodal large language model to drive its reasoning on the defect type, location and status. This module enhances the model's sensitivity to complex defects by fusing visual features with dynamic prompts.

[0050] In the factory production process, by using the interactive industrial defect detection system provided by the embodiment of the present invention, it is possible to identify whether a product has defects and the type of defects. By analyzing the detection results, the frequency of occurrence of various defects can be counted, thereby helping enterprises to take corresponding measures to avoid or reduce the occurrence of these defects. Even in the absence of product images, the present invention can be directly used to implement reasoning of a single image, providing convenience for industrial production.

[0051] Second embodiment

[0052] like Figure 2 As shown, based on the above-mentioned interactive industrial defect detection system, the second embodiment of the present invention provides a zero-sample driven interactive industrial defect detection method, which at least includes the following steps:

[0053] S1: Create an auxiliary data sample library based on the defect dataset in the industrial field and generate auxiliary training samples;

[0054] S2: Construct a multimodal large language model in the training phase, where ViT-L / 14@336px is selected as the image encoder pre-training model and Vicuna-7b is selected as the large language pre-training model. The network weights of the two models are frozen, and only the prompt learner module and the object-independent text prompt module are allowed to participate in the training;

[0055] S3: Filter samples containing defects from the auxiliary data sample library as the defect detection auxiliary sample set for fine-tuning the multimodal large language model;

[0056] S4: Screening defect-free normal samples from the auxiliary data sample library as the defect-free detection auxiliary sample set for fine-tuning the multimodal large language model;

[0057] S5: Self-collected data covering different background conditions, industrial sample types and defect categories are used as auxiliary sample sets for zero-shot detection to verify the generalization ability of the multimodal large language model under target-free scene training data;

[0058] S6: Setting an object-independent text prompt module, which includes two types of structures: a normal sample description template includes: prefix description + [object placeholder] + [normal state word]; a defect sample description template includes: prefix description + [object placeholder] + [defect type word] + [abnormal state word]; the text prompt module generates structured text prompt data;

[0059] S7: Input a single query image to the query image encoding module, extract image features and vectorize them through the multi-scale Vision Transformer model, and divide it into two branches: the first branch directly inputs the vectorized features into the multimodal large language model, and the second branch generates a heat map by aligning the vectorized features with the text prompt data;

[0060] S8: Input the heat map to the prompt learner module, which receives three input sources: the spatial features of the heat map extracted by the convolutional neural network (CNN); the global image feature vector output by the query image encoder; and the dynamic prompt vector generated by the learnable prompt learner;

[0061] The three are spliced ​​together in the channel dimension to achieve multimodal feature fusion, and then input into the multimodal large language model after normalization in the fully connected layer;

[0062] S9: The output features of the prompt learner module and the questions submitted by the user through the defect recognition dialogue module (such as "Are there cracks in the picture?") are jointly input into the multimodal large language model to generate a dialogue result including the defect type, location and explanation (such as "Two cracks were detected, located in the right edge area").

[0063] It should be noted that the zero-sample drive in the embodiment of the present invention means that no target scene training data is required in the inference stage, but the training stage still relies on general industrial defect datasets (such as AITEX, MVTec-AD, etc.). In one embodiment of the present invention, step S1 uses public datasets of multiple materials and multiple defect types in the industrial field to build an auxiliary sample library, which specifically includes the following seven types of datasets:

[0064] AITEX: Focuses on textile defect detection, including high-resolution images of various defects on the fabric surface, suitable for textile quality inspection scenarios.

[0065] RSDDs: Targets rail surface defects, including typical abnormalities such as rail cracks, rust, and wear, and serves rail transit equipment inspection.

[0066] Magnetic: Focuses on surface defects of ceramic tiles, including five common abnormalities: uneven surface, fracture, crack, wear and pores, suitable for quality monitoring in the building materials industry.

[0067] DAGM: Provides miscellaneous defect data under 10 types of optical texture backgrounds. The data is large in scale and covers comprehensive defect types, supporting anomaly detection research under complex texture backgrounds.

[0068] VisA: Contains 10 types of industrial component defects (such as printed circuit boards, capacitors, transistors, chips, etc.), and is also extended to daily necessities (such as surface defects of cashews and chewing gum), taking into account the diverse needs of both industrial and daily scenarios.

[0069] MVTec-AD: The current mainstream industrial detection benchmark dataset, covering comparative data of normal samples and abnormal samples, is widely used in algorithm performance testing and model optimization.

[0070] DTD-Synthetic: An industrial defect dataset generated through synthetic technology that simulates the abnormal morphology of surfaces of various materials to supplement training needs in scenarios where real data is scarce.

[0071] The above datasets cover industrial fields such as textiles, building materials, electronics, and transportation, and include 13 types of defects such as cracks, holes, and deformations. They take into account both real data and synthetic data to provide diverse sample support for model training.

[0072] In one embodiment of the present invention, in step S2, the ViT-L / 14@336px model is selected as the pre-training model of the image encoder, and the Vicuna-7b model is selected as the pre-training model of the multimodal large language model; during the training process, the network weights of the two models are frozen, and only the prompt learner module and the object-independent text prompt module are allowed to participate in the network training. Among them, the ViT-L / 14@336px model is a large variant based on the Vision Transformer (ViT) architecture, with the following characteristics: the architecture is a Large version, and the image block size is set to 14×14; in the pre-training stage, a high-resolution image of 336×336 pixels can be used as input; after completing the pre-training, a higher-resolution image will be further used for fine-tuning to improve the performance of the model. This model is widely used in multimodal tasks. For example, it is used as an image encoder in the CLIP model, and works together with the text encoder to complete tasks such as image-text matching and zero-sample classification. In image classification tasks, especially when faced with high-resolution images, it can effectively capture the global features of the image and perform well. In addition, through fine-tuning, the ViT-L / 14@336px model can also be applied to tasks such as target detection and semantic segmentation.

[0073] Here is a code example using the ViT-L / 14@336px model:

[0074]

[0075] The above code implements the function of extracting image features using a pre-trained vision model (ViT-L / 14@336px), thus demonstrating strong performance in multimodal tasks, especially in high-resolution image processing.

[0076] On the other hand, Vicuna-7b is an open source, high-performance multimodal large language model. The model has 7 billion parameters and powerful performance. Users can easily interact with the model through the command line interface to achieve multi-round human-computer dialogue, greatly improving the convenience and flexibility of use. As a conventional technology commonly known by those skilled in the art, it will not be described in detail here.

[0077] In one embodiment of the present invention, the method for constructing a multimodal large language model in the training phase in step S2 includes the following two sub-steps: first, the output result of the object-independent text prompt module that has undergone text encoding and text embedding is superimposed with the auxiliary training sample that has undergone image encoding and image embedding, thereby obtaining a multimodal model that integrates text and image information; second, the multimodal model is further processed by a prompt learner module and then input into the large language reasoning model, and finally a multimodal large language model with multimodal understanding and generation capabilities is obtained.

[0078] In one embodiment of the present invention, the defect types involved in step S3 include, but are not limited to: missing, scratches, cracks, dirt, holes, deformation, pits, breakage, burrs, delamination, impurities, rust and bubbles. These defect types cover various abnormal conditions commonly seen in industrial production. By using samples containing these defect types in the training phase, the model's ability to identify and detect different types of defects can be effectively improved.

[0079] In one embodiment of the present invention, the self-collected data mentioned in step S5 refers to samples collected under different background conditions and different types of industrial samples, including samples with different defect types and samples without defects. By using these self-collected data as auxiliary sample sets for zero-sample detection, the generalization ability and practical application effect of the multimodal large language model in the absence of target scene training data can be further verified and improved.

[0080] In one embodiment of the present invention, the method for setting an object-independent text prompt module in step S6 is as follows: First, set a normal sample description template and an abnormal sample description template. The structure of the normal sample description template is "prefix description + [object placeholder] + [normal state word]", and the structure of the abnormal sample description template is "prefix description + [object placeholder] + [defect type word] + [abnormal state word]". Secondly, define the geometric lists of normal state, abnormal state and defect type. The geometric list of normal state includes {good, perfect, flawless}, the geometric list of abnormal state includes {defect, broken, bad, flaw, damaged}, and the geometric list of defect type includes {blowhole, break, crack, deletion, scratch, smudge, deformation, pit, rag, hierarchy, impurity, corrosion, bubble}. Finally, set the object category to [object], and embed the state description thereafter to indicate the state of the object. If the object state is abnormal, further describe the specific defect type.

[0081] It should be noted that the above template design intentionally masks object categories and does not focus on specific sample types, but focuses on the state attributes and defect types of the samples. This design enables the model to focus on common defect features rather than specific objects, thereby improving cross-scenario generalization capabilities. Through contrastive learning, the text encoder focuses on state word differences, and the generated text embedding is fused with image features to drive defect location and classification. The embedded content generated by this text prompt module is particularly suitable for defect detection in the industrial field. It can quickly output defect types and locate defects, improve detection efficiency and accuracy, and at the same time has high versatility and flexibility, adapting to different types of industrial samples and defect detection needs.

[0082] like Figure 3 As shown, the embodiment of the present invention makes targeted improvements to the CLIP text encoder and constructs an object-independent text prompt module: the normal sample description template is represented by a fixed-length word vector sequence (V 1 …V e ) describes the general state of the object (such as "good", "flawless"), and the exception text template is expanded to (W 1 …W e) structure, embedding defect types (such as "crack", "deformation") and abnormal status words (such as "broken", "defect") after the object word. Compared with the original CLIP text encoder (which encodes object categories and attributes at the same time), the embodiment of the present invention deletes the object feature extraction branch and reconstructs the encoder into a two-way parallel structure - normal / abnormal templates share the same Transformer encoding layer. Through contrastive learning, the model is forced to focus on the semantic differences between status words and defect words. The final output text embedding (X) only contains status attributes and defect type information, which effectively improves the sensitivity to abnormal features in industrial scenarios and achieves the "object-independent, defect-driven" detection goal.

[0083] like Figure 4 As shown, in one embodiment of the present invention, the method for extracting image features in step S7 is a multi-scale bottom-up image feature extraction method, and the specific steps are as follows:

[0084] S71: Select base model

[0085] The Vision Transformer (ViT) deep learning model is selected as the base model for image visual feature extraction. The model contains 4 layers of encoders, which extract the visual features of the image layer by layer in a bottom-up manner and retain the output features of each layer of encoder for subsequent processing.

[0086] It should be noted that the Vision Transformer (ViT) deep learning model is a deep learning model that applies the Transformer architecture to computer vision tasks. The core idea of ​​ViT is to divide the image into fixed-size image blocks and treat these image blocks as sequences and input them into the Transformer model, similar to word sequences in natural language processing. In this way, ViT can effectively capture the global and local features in the image, providing strong support for subsequent image analysis and processing.

[0087] S72: Feature alignment and merging

[0088] The visual features obtained by each layer of encoder are aligned with the text feature vector, and then the aligned visual features and text feature vectors are merged. Specifically, the correlation between the visual features and the text features is measured by calculating the cosine similarity, and these similarity values ​​are added to the visual features to obtain a mixed vector that incorporates the text information. Finally, the mixed vectors obtained by each layer of encoder are merged to obtain the final vectorized image features.

[0089] like Figure 5As shown, in one embodiment of the present invention, inputting the heat map into the prompt learner module in step S8 includes the following sub-steps: the heat map first undergoes a convolution operation to extract local features in the image, then performs a nonlinear transformation through an activation function, then performs feature dimensionality reduction and feature mapping through a maximum pooling operation, and finally undergoes another convolution operation to further extract features and generate an embedded representation.

[0090] Next, a learnable hint learner is embedded into the hint learner module to provide additional information for industrial defect detection. The learnable hint learner automatically adjusts its parameters through the learning process to generate more discriminative hint embeddings, thereby helping the model to better identify and locate defects. In addition, auxiliary training samples that have undergone image encoding and image embedding are embedded into the hint learner module. The image features of these auxiliary training samples are combined with other information in the hint learner module to further enhance the model's ability to learn and understand defect features.

[0091] Through the above steps, the prompt learner module extracts the spatial information in the heat map using a convolutional network, and combines it with the semantic prompt embedding dynamically generated by the learnable prompt parameters, and splices it with the global image features to form a unified multimodal representation, thereby achieving effective fusion of multimodal features. This process not only provides rich feature information for the multimodal large language model, but also enables the model to adapt the general knowledge obtained in the pre-training stage to new scenarios by learning common defect samples on auxiliary datasets, so that it does not need to rely on the data of the target scenario in actual application.

[0092] The above embodiments are only examples, and the technical solutions of the various embodiments can be combined, all within the protection scope of the present invention.

[0093] It should be noted that the essence of the "zero-sample drive" implemented in the embodiment of the present invention is to utilize the prior knowledge of the pre-trained model, combined with the object-independent prompt design and dynamic feature fusion mechanism, so that the conversational industrial defect detection system can complete industrial defect detection without the need for target scene training data, solving the pain points of traditional methods that rely on labeled data and have weak generalization capabilities. Specifically, through the knowledge transfer of the pre-trained model, ViT-L / 14@336px (visual model) and Vicuna-7b (large language model) are used as pre-trained base models, and their network weights are frozen during training, relying only on their existing multimodal understanding capabilities, without the need to retrain for the target scene. These pre-trained models learn rich visual and semantic features through massive general data, which can be directly transferred to industrial defect detection tasks, avoiding dependence on target scene data.

[0094] Compared with the prior art, the present invention specifically targets 13 types of industrial defects, including missing, scratches, cracks, dirt, holes, deformation, pits, breakage, burrs, delamination, impurities, rust and bubbles, and proposes a conversational industrial defect detection method. During the training process, data of corresponding categories are selected from public data sets and self-collected data sets to form auxiliary data sets for training a conversational industrial defect detection system of a multimodal large language model. In the network structure of the system, the image encoder model, the text encoder model and the multimodal large language model all use official pre-trained models, such as the ViT-L / 14@336px model and the Vicuna-7b model, and these models are frozen during the training process. Only the network weights of the object-independent text prompt module and the prompt learner module can be trained and updated. By using the auxiliary data set for training, the model can only use the pre-trained network weights that have been pre-trained when reasoning with a single image, without the need to train the model again. Since the auxiliary dataset already contains comprehensive and rich industrial defect samples, the model can be used directly to reason about a single image in industrial testing without the need for additional sample sets, thus achieving zero-sample driven detection capabilities.

[0095] The zero-sample driven interactive industrial defect detection system and method provided by the present invention are described in detail above. For those skilled in the art, any obvious changes made to it without departing from the essence of the present invention will constitute an infringement of the patent right of the present invention and will bear corresponding legal liabilities.

Claims

1. A zero-sample driven conversational industrial defect detection system, characterized in that It includes a query image encoding module, an object-independent text prompt module, a prompt learner module, and a defect recognition dialogue module; The query image encoding module encodes and embeds features of the query image to generate two branches: the first branch inputs visual features into the multimodal large language model in the defect recognition dialogue module; the second branch integrates the output of the object-independent text prompt module; The text prompt module generates text prompts, which are encoded and superimposed with image features to generate a heat map; the heat map is processed by a convolutional neural network in the prompt learner module, and is aligned and connected with learnable prompt parameters and output features of a frozen pre-trained model, and finally input into a multimodal large language model to complete multimodal reasoning; The defect identification dialogue module sends dialogue data to the multimodal large language model, and the multimodal large language model returns the detection result to the user.

2. The interactive industrial defect detection system according to claim 1, characterized in that The object-independent text prompt module includes a normal sample description template and a defect sample description template; Among them, the text encoding results of the normal sample description template and the defect sample description template are fused through a feature addition operation to generate a joint text representation, which is then mapped into a low-dimensional vector through an embedding layer for alignment with image features.

3. The interactive industrial defect detection system according to claim 1, characterized in that The cue learner module includes a convolutional neural network and a learnable cue learner; The heat map is input into the convolutional neural network to extract spatial features, and is spliced ​​with the global feature vector and dynamic prompt vector of the query image in channel dimension to form a unified multimodal feature; the aligned features are reduced in dimension and normalized by the fully connected layer and input into the multimodal large language model.

4. A zero-sample driven conversational industrial defect detection method, implemented based on the conversational industrial defect detection system according to any one of claims 1 to 3, characterized in that The steps include: S1: Create an auxiliary data sample library based on the defect dataset in the industrial field and generate auxiliary training samples; S2: Build a large multimodal language model in the training phase, freeze the network weights of the model, and only allow the prompt learner module and the object-independent text prompt module to participate in the training; S3: Filter samples containing defects from the auxiliary data sample library as the defect detection auxiliary sample set for fine-tuning the multimodal large language model; S4: Screening defect-free normal samples from the auxiliary data sample library as the defect-free detection auxiliary sample set for fine-tuning the multimodal large language model; S5: Self-collected data covering different background conditions, industrial sample types and defect categories are used as auxiliary sample sets for zero-shot detection to verify the generalization ability of the multimodal large language model under target-free scene training data; S6: Setting an object-independent text prompt module, which includes a normal sample description template and a defect sample description template; the text prompt module generates structured text prompt data; S7: Input the query image to the query image encoding module, extract image features and vectorize them, and divide them into two branches: the first branch directly inputs the vectorized features into the multimodal large language model, and the second branch generates a heat map by aligning the vectorized features with the text prompt data; S8: Input the heat map to the prompt learner module, which receives the spatial features extracted by the convolutional neural network, the global image feature vector output by the query image encoder, and the dynamic prompt vector generated by the prompt learner; realize multimodal feature fusion through channel dimension splicing, and input the multimodal large language model after normalization by the fully connected layer; S9: The output features of the prompt learner module and the questions submitted by the user through the defect identification dialogue module are input into the multimodal large language model to generate a dialogue result including defect type, location and explanation.

5. The interactive industrial defect detection method according to claim 4, characterized in that Building a large multimodal language model in the training phase includes the following steps: First, the output result of the text prompt module after text encoding and text embedding is superimposed with the auxiliary training sample after image encoding and image embedding to obtain a multimodal model; Secondly, the multimodal model is processed by the prompt learner module and input into the large language inference model to obtain the multimodal large language model.

6. The interactive industrial defect detection method according to claim 5, characterized in that: In the step S2, ViT-L / 14@336px is selected as the image encoder pre-training model and Vicuna-7b is selected as the large language pre-training model.

7. The interactive industrial defect detection method according to claim 4, characterized in that Setting up the object-independent text prompt module includes the following sub-steps: First, set the normal sample description template and the abnormal sample description template; second, define the geometric list of normal state, abnormal state and defect type; finally, set the object category and embed the state description to indicate the state of the object.

8. The interactive industrial defect detection method according to claim 7, characterized in that: The normal sample description template includes: prefix description+[object placeholder]+[normal status word]; the defect sample description template includes: prefix description+[object placeholder]+[defect type word]+[abnormal status word].

9. The interactive industrial defect detection method according to claim 4, characterized in that In step S7, extracting image features includes the following sub-steps: S71: Select the Vision Transformer deep learning model containing 4 layers of encoders as the base model for image visual feature extraction, extract the visual features of the image layer by layer in a bottom-up manner, and retain the output features of each layer of encoders; S72: Align the visual features obtained by each layer of encoder with the text feature vector, and then merge the aligned visual features and text feature vector to obtain the final vectorized image features.

10. The interactive industrial defect detection method according to claim 4, characterized in that In step S8, inputting the heat map into the prompt learner module includes the following sub-steps: The heat map first undergoes a convolution operation to extract local features in the image, then undergoes a nonlinear transformation through an activation function, then undergoes feature dimensionality reduction and feature mapping through a maximum pooling operation, and finally undergoes another convolution operation to further extract features and generate an embedded representation.

Citation Information

Patent Citations

  • Anomaly detection methods and devices in industrial testing

    CN116777906B

Cited By

  • Small-sample defect identification method based on cross-modal text semantic driving

    CN120580702A

  • A few-shot defect identification method driven by cross-modal text semantics

    CN120580702B

  • Cross-model defect detection method

    CN121304540A

  • Steel surface defect detection method based on multi-modal characteristics

    CN122156054A

  • A steel surface defect detection method based on multi-modal features

    CN122156054B