Surface defect detection method and system based on multimodal large model
By applying the apparent defect detection method based on multimodal large models in an industrial production environment, combined with the defect detection mask and text description, the problems of low efficiency and low accuracy of manual evaluation in the prior art are solved, automated detection and evaluation are realized, and production efficiency and detection accuracy are improved.
Patent Information
- Application Number
- CN202510258419.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The prior art relies on manual evaluation in surface defect detection in industrial production environments, resulting in low efficiency, low accuracy and high cost, and lack of text description and detailed evaluation of defective parts.
The apparent defect detection method based on a multimodal large model is adopted, and the defect image is encoded and decoded using a large language backbone network to generate the final detection mask and text description.
It realizes automated defect detection and evaluation, improves production efficiency, reduces management decision-making costs, enhances the accuracy and credibility of defect detection, and provides detailed text descriptions and repair suggestions for defect parts.
Smart Images

Figure CN119762485B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and specifically to a method and system for detecting apparent defects based on a multimodal large model. Background Technology
[0002] In industrial production environments, surface defect detection is crucial for ensuring product quality and yield. However, in practical applications, simply detecting defects is far from sufficient. Factory managers want to monitor production in real time based on defect detection results to promptly correct errors and adjust production methods. Currently, however, the evaluation of defect detection results is mostly done by experienced engineers. This manual evaluation method is not only time-consuming and labor-intensive, but its accuracy and reliability are also affected by the engineers' fluctuating energy and work status, severely impacting factory production efficiency. Therefore, exploring a defect detection and automated evaluation mechanism based on artificial intelligence technology is extremely important.
[0003] In recent years, with the proposal and development of large language backbone networks, algorithms based on multimodal large models have been widely applied in fields such as image segmentation, image classification, object detection, and visual question answering, and their application effects have been proven to be quite remarkable by numerous research studies. Leveraging the powerful inference and zero-shot capabilities of large language backbone networks, multimodal large models can not only understand inputs from multiple modalities, such as text modalities, image modalities, and mask modalities, but also output the desired detection and evaluation results for multiple modalities through inference. Existing cutting-edge deep learning-related defect detection techniques can be broadly categorized into two types: 1) feature embedding-based methods; and 2) image reconstruction-based methods. The former relies on feature extractors pre-trained with a large amount of data to extract image features, and then uses custom criteria and nearest neighbor strategies to measure the distance between normal and abnormal data to obtain detection results. The latter relies on generative models to learn the data distribution of normal images, reconstructing the input abnormal image into a normal image, and then comparing the differences between the two to obtain detection results.
[0004] Both of these detection methods have significant drawbacks: 1) Feature embedding-based methods are limited by the quality of feature extraction. If the feature extractor is poorly trained or lacks domain adaptation, it will severely affect the defect detection performance. Furthermore, in high-dimensional space, sample sparsity can cause distance metrics to fail, affecting the identification of anomalies. It is worth noting that in actual industrial production environments, the defect area of products is often very small. This means that in high-dimensional space, feature embedding-based methods may fail to accurately identify anomalies, leading to misjudgments and allowing defective products to enter the market, resulting in incalculable economic and reputational losses. 2) Image reconstruction-based methods are limited by the quality of the selected generative model. If the generative model cannot effectively capture the distribution of normal data, the reconstruction error may be inaccurate, leading to poor defect detection performance. Moreover, because the data distribution of defective images cannot be observed during training, the reconstruction quality of defective regions during inference is relatively high, thus affecting the overall defect detection quality. More notably, current defect detection technology focuses more on identifying and outputting visual inspection results, lacking textual descriptions and detailed assessments of defect locations. Such descriptions and detailed assessments still require experienced engineers, and this lack of textual descriptions and detailed assessments affects the efficiency of decision-making and the speed of adjustments in production. Summary of the Invention
[0005] To overcome the above-mentioned shortcomings, this invention provides a method and system for detecting apparent defects based on a multimodal large model. By combining defect detection masks and textual descriptions of defects, this method greatly empowers the defect detection process in industrial production environments and improves production efficiency.
[0006] The apparent defect detection method based on a multimodal large model designed in this invention includes the following steps:
[0007] A labeled training dataset, which contains defect images and corresponding masks and text descriptions;
[0008] The defective image is encoded and an L×N token is assigned to the encoder, where L represents the number of visual scales and N represents the number of tokens assigned to each visual scale.
[0009] Align the encoded visual features to the language feature space;
[0010] Using a large language backbone network, the text description of the defect image, multi-scale token groups, and visual features aligned to the language feature space are input. After processing, language type tokens and visual type tokens are obtained, and the language type tokens are decoded into text descriptions and evaluations of the defect image.
[0011] Align visual type tokens to the visual feature space;
[0012] The encoded visual features and aligned visual type tokens are decoded to obtain the final detection mask.
[0013] A multimodal large model consisting of an encoder, language features aligned with visual features, a large language backbone network, visual type tokens aligned with language features, and a decoder is trained, and the trained multimodal large model is used for defect detection.
[0014] Furthermore, the textual descriptions in the training dataset include the type, shape, and size of the defect, the cause of the defect, and suggestions for repairing the defect.
[0015] Furthermore, GPT-4V was used to generate a labeled training dataset.
[0016] Furthermore, the alignment of visual features to language features and the alignment of language features to visual features both adopt the same MLP structure, specifically including a linear layer, a ReLU layer, a linear layer, and a Dropout layer.
[0017] Furthermore, the encoder adopts a pre-trained CLIP-ViT encoder; the large language backbone network adopts a pre-trained Vicuna-v1.5-13B. The encoder and the large language backbone network are frozen in a state during multimodal large model training, and their parameters are not updated.
[0018] Furthermore, the decoder has a total of L layers, each of which adopts a self-attention layer structure of multidimensional convolutional head transposed attention and gated convolutional feedforward network.
[0019] Furthermore, when training a large multimodal model, the following loss is used:
[0020] = +
[0021] in, This represents the autoregressive cross-entropy loss, ensuring the quality of text generation. This represents pixel-by-pixel binary cross-entropy loss, ensuring the quality of mask generation; and This is a hyperparameter.
[0022] Based on the same inventive concept, this invention also discloses a system for detecting apparent defects based on a multimodal large model, comprising:
[0023] Visual encoder The input defect image is encoded to extract visual features. The encoder is assigned L×N tokens, where L represents the number of visual scales and N represents the number of tokens assigned to each visual scale.
[0024] Visual-Language Mapper Align the encoded visual features to the language feature space;
[0025] The large language backbone network F receives text descriptions of defect images, multi-scale token groups, and visual features aligned to the language feature space as input. After processing, it obtains language type tokens and visual type tokens, and directly decodes the language type tokens into text descriptions and evaluations of the defect images.
[0026] Language-visual mapper Align visual type tokens to the visual feature space;
[0027] Visual decoder : Decode the encoded visual features and the aligned visual type token to output the final detection mask.
[0028] Based on the same inventive concept, this method also discloses an electronic device, comprising:
[0029] One or more processors;
[0030] Storage device for storing one or more programs;
[0031] When one or more programs are executed by the one or more processors, the one or more processors implement a method for detecting apparent defects based on a multimodal large model.
[0032] Based on the same inventive concept, this method also discloses a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements a method for detecting apparent defects based on a multimodal large model.
[0033] The advantages of this invention are:
[0034] 1. This patent innovatively introduces an automatic defect evaluation mechanism in defect detection tasks by utilizing a multimodal large model, which greatly optimizes the adjustment efficiency of production methods, reduces management decision-making costs, and improves production efficiency.
[0035] 2. This patent innovatively designs a multi-scale token group mechanism to decouple the visual features output by the encoder, optimize the generation quality of the mask, and improve the accuracy of defect detection.
[0036] 3. This patent innovatively designs a lightweight visual decoder, reducing the time complexity of the mask generation process from O(n log n). ) decreased to O( This greatly accelerates the training and inference process of the entire model and improves the efficiency of defect detection. Attached Figure Description
[0037] Figure 1 This is the overall architecture diagram of the present invention.
[0038] Figure 2 This is a detailed structural diagram of the visual decoder part of the present invention. Detailed Implementation
[0039] To facilitate understanding and implementation of this invention by those skilled in the art, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are for illustration and explanation only and are not intended to limit the invention.
[0040] Terminology Explanation:
[0041] GPT-4V: A multimodal large model developed by openAI based on the GPT-4 model for vision tasks.
[0042] CLIP-ViT Encoder: Composed of CLIP and VIT (Vision Transformer). CLIP (Contrastive Language–Image Pre-training), a model proposed by OpenAI, includes an image encoder, a text encoder, and a contrastive training method. ViT is a transformer architecture suitable for computer vision tasks. Similar to the Transformer for text processing, ViT uses a self-attention mechanism to capture local and global features of an image. CLIP-ViT can simultaneously understand and process images and text, making it suitable for tasks such as image retrieval and text description generation. Zero-shoot learning: CLIP-ViT can be directly applied to various vision and language tasks without retraining. This makes it more versatile when facing new tasks.
[0043] Vicuna-v1.5-13B: Vicuna is based on the transformer architecture and includes a multi-layered self-attention mechanism, enabling it to understand the context of the input text. This design makes it excel in dialogue generation and natural language processing tasks. Vicuna can generate natural and fluent conversational text, suitable for chatbots, customer service systems, etc. Text Understanding and Generation: It can perform a range of natural language processing tasks such as text summarization, question answering, and text generation. Control and Customization: It supports adjusting the response style according to user needs, such as requesting a more formal or more relaxed language style.
[0044] Token: In text processing, text is broken down into smaller units. These units can be words, sub-words, characters, or other text elements. In natural language processing, models typically cannot directly process raw text; therefore, it must first be broken down into a form suitable for model operation. These decomposed units are tokens.
[0045] The time complexity is In computer science, the time complexity of an algorithm is expressed as quadratic. : Represents the "asymptotic upper bound," used to describe the growth rate of algorithm performance, such as time or space complexity. It gives the upper limit of the algorithm's running time or space requirements as the input size increases infinitely. n: Represents the size or scale of the input data. It can be any measure; typically, in algorithm analysis, n represents the number of elements in the data to be processed. For example, in an array or list, n can represent the length of the array.
[0046] Gated-Dconv Feed-Forward Network (GDFN): Composed of multiple convolutional layers, it is used to extract global features and perform denoising. It enhances detail recovery by integrating features from multiple layers.
[0047] Multi-Dconv Head Transposed Attention (MDTA) consists of multiple convolutional layers and a self-attention mechanism. It calculates cross-channel cross-covariance to generate an attention map that implicitly encodes the global context. Functionally, it captures long-range dependencies, thereby improving feature representation and information delivery during image generation.
[0048] Example 1
[0049] This embodiment provides a method for detecting apparent defects based on a multimodal large model, as shown in the attached figure. Figure 1 As shown, its specific implementation includes the following steps:
[0050] Step S1: Label the training dataset to train the appearance defect detection model established in this method.
[0051] S1-1, Design a prompt to guide GPT-4V on how to generate data. In this embodiment, the prompt is: "Imagine you are an experienced defect detection engineer. I am now giving you a diagram. Please carefully examine this diagram and answer the following three questions based on my requirements: 1. Describe the type, shape, and size of the defect in this diagram. 2. Analyze the cause of this defect. 3. Give your suggestions for repairing this defect. Requirements: Each part of your answer should correspond to the relevant question, and your answers should be as detailed as possible." This prompt not only allows GPT-4V to deduce an analysis and description of the defect through reasoning but also provides reference suggestions for repairing the defect.
[0052] S1-2 calls the GPT-4V API to perform batch annotation of the data. The base dataset contains a large number of defect images and their masks. All defect images from the base dataset, along with the cue words from S1-1, are input into GPT-4V. GPT-4V will output a text description and evaluation of the defect image. An example of the text description and evaluation is shown below:
[0053] 1. Defect type, defect shape, and defect size
[0054] This image shows an asphalt pavement. Small cracks and uneven textures are visible, likely indicating minor defects. These cracks are not very prominent, are irregularly distributed, and are relatively narrow. A rough estimate suggests that these cracks are approximately 1-2 millimeters wide and several centimeters long.
[0055] 2. Analysis of the causes of defects
[0056] These fine crack defects may be caused by the following reasons:
[0057] Natural aging: Over time, road materials become increasingly fragile and prone to cracking due to natural factors such as weathering, ultraviolet radiation, and temperature changes.
[0058] Temperature changes: Alternating between hot and cold temperatures can cause asphalt to expand and contract, resulting in thermal expansion and contraction, which in turn can cause cracks.
[0059] Loading: Under prolonged vehicle traffic, road materials suffer fatigue damage, especially when heavy vehicles frequently pass by, making them more prone to cracking.
[0060] 3. Repair Recommendations
[0061] To repair these tiny cracks, the following measures are recommended:
[0062] Surface sealing treatment: A layer of asphalt sealant can be applied to the road surface to fill the cracks, prevent moisture from seeping in, and slow down the expansion of the cracks.
[0063] Preventive maintenance: It is recommended to regularly check the condition of the road surface and repair early cracks in a timely manner to prevent the cracks from expanding further.
[0064] Material improvement: If the road section has a high frequency of cracking, consider using a more crack-resistant asphalt material or adding a crack-resistant agent during repaving to improve the durability of the material.
[0065] After labeling each defect image, a dataset containing three types of data can be obtained: defect image, text description, and mask.
[0066] Step S2, design the mapper.
[0067] Design Visual-Language Mapper and language-visual mapper Both mappers use the same MLP structure, which consists of: linear layer → ReLU → linear layer → Dropout. For the vision-language mapper... In this process, the first linear transformation layer primarily maps the input visual features linearly, adjusting their dimension and distribution to suit subsequent processing. This can be understood as feature selection and preliminary transformation of visual features to align them with the text space. The ReLU activation function is introduced to introduce non-linearity, enabling the model to capture complex feature relationships. It truncates the input, retaining only positive values, thus improving the model's expressive power and allowing visual features to better align with the text space. The second linear layer further maps the activated features linearly to the target text space dimension. Through this layer, the model can transform the non-linearly transformed visual features into a representation more consistent with text features, effectively converting visual information into a text-understandable form. Dropout randomly discards a portion of the neuron outputs during training to prevent overfitting. Although it doesn't directly participate in the feature alignment process, this regularization method improves the model's generalization ability, making the generated aligned features more robust across different samples; for language-visual mappers... In addition, the functions of each layer are similar to those described above, and will not be repeated here;
[0068] The roles of the two mappers in the overall model are formally described using mathematical language, as follows:
[0069] = , indicating the input image All visual features output after being encoded by the visual encoder E, and These represent the visual features output by each layer of E. Specifically, These are the features output by the last layer of E. Used to align visual features into the language feature space:
[0070] = ( (1)
[0071] in Features aligned to the language feature space. And the language-visual mapper. Its function is to place the features in the language feature space The embedding form is aligned to the visual feature space that the decoder D can understand:
[0072] = ( (2)
[0073] in, for Features aligned to the visual feature space. (Note: The original text contains some formatting errors and inconsistencies. A more accurate translation ={ ,..., }, For the visual type tokens output by the large language backbone network, the first... l The layer utilizes the tokens fused from the linear layer.
[0074] Step S3: Design a multi-scale token group mechanism to enhance defect detection results.
[0075] In the process of visual encoder E extracting image features layer by layer, existing research has confirmed that shallower layers of E extract features containing more information on texture, geometry, edges, and corners, while deeper layers contain more high-dimensional semantic information. In the field of defect detection, to accurately detect the location and shape of defects, it is necessary to comprehensively consider both texture and geometric information and high-dimensional semantic information. Existing technologies only assign one token to the final output layer of E, tightly coupling information from different visual scales into a single token, thus conflating geometric information with high-dimensional semantic information. This invention proposes a multi-scale token grouping mechanism, which assigns not only one token to the final output layer of E, but also N tokens to each layer of E. E has a total of L layers, with L×N tokens representing visual features. The L×N tokens are divided into L groups of N tokens each, with each group having a different visual feature granularity. At the start of training, L×N tokens are randomly initialized to represent the visual features of a defective image at different scales. During the calculation of loss and backpropagation, the model learns the token features that carry visual information and adds them to the Large Language Backbone Network vocabulary (LLM vocabulary).
[0076] S3-2, using mathematical language, formally describes the multi-scale token group as follows:
[0077] like E( ), = Having L layers, for the first Visual features of layers All tokens allocated to it are ={ }. That is, the first There are N tokens in the layer. These N tokens are considered as a token group, denoted as . L token groups are further divided into ={ Finally, , and input text prompts Input into the large language backbone network:
[0078] =F( , , (3)
[0079] Output of large language backbone network In the process, the language type token is directly decoded into the final defect text description and evaluation, and the remaining visual type token is denoted as... = The definitions of N and L remain unchanged. There are L groups of s, and each group contains N tokens. After alignment, it is input into decoder D.
[0080] Step S4, Design the visual decoder
[0081] As attached Figure 2 As shown, decoder D is used to output the final defect detection mask. It consists of L attention blocks, each of which is responsible for processing... Each layer outputs features and the corresponding token group. Each attention block outputs an attention score map, which is then weighted and summed to obtain the final defect detection mask. The internal structure of each attention block is the same, using a multidimensional convolutional head transposed attention network (MDTA) + gated convolutional feedforward network (GDFN), abbreviated as MDTA+GDFN.
[0082] Since the images being processed originate from real-world production environments, their resolution may be high, and the system demands high processing speed and efficiency. Therefore, a lightweight decoder with a fast inference rate is required. Traditional methods often use SAM (Segment Anything Model) decoders. This invention employs a Transformer-like decoder architecture. In traditional Transformer architectures, the computational overhead primarily comes from the self-attention layer when processing images, resulting in a time complexity of O(n log n). Therefore, traditional self-attention layers are inefficient when processing high-resolution images, resulting in significant time consumption during training and inference. This invention uses an MDTA+GDFN self-attention layer structure, achieving a time complexity of O(log n) for calculating self-attention scores. This significantly reduces computational costs and improves defect detection efficiency. Specifically, the decoder D has a total of L layers, each consisting of MDTA + GDFN. MDTA calculates the self-attention score with linear complexity, while GDFN ensures that useful information can be further propagated and eliminates useless information.
[0083] Describe the decoder formally using mathematical language. Its function is as follows:
[0084] Output visual tokens in the large language backbone network = After that, regarding the first All outputs of the group { A linear layer is used to merge all tokens into a single token.
[0085] = Linear( (4)
[0086] The merged token will be used with a language-visual encoder Aligning to the visual feature space yields ={ After this, it will be fed into the decoder D along with the visual features encoded by the visual encoder to obtain the mask:
[0087] mask =D( , (5)
[0088] Specifically, the operations performed inside D are as follows:
[0089] = (6)
[0090] = (7)
[0091] in This is an element-wise multiplication operation. For sigmoid operation, For the decoder's first An attention block, consisting of MDTA and GDFN. The attention values calculated for each attention block are then used to calculate the final detection mask through a weighted summation operation.
[0092] mask = (8)
[0093] in They are learned as weighting factors during training.
[0094] Step S5: Train a multimodal large model consisting of an encoder, language features aligned with visual features, a large language backbone network, visual features aligned with language features, and a decoder.
[0095] To ensure the generated data conforms to the format of visual question answers and to fine-tune the instructions, the dialogue format during training is fixed as: "User: <Image> Please detect and analyze the image I provided, and provide a defect detection mask and corresponding detailed evaluation for this image. Multimodal Large Model: {Text}<Mask>". After fine-tuning the instructions using this format, the large language backbone network demonstrates strong instruction compliance ability and solves various downstream tasks through zero-shot learning. The loss is calculated by combining the {Text} with the text descriptions in the dataset, denoted as... The loss is calculated between the <mask> and the ground truth mask in the dataset from step S1, denoted as . These two loss functions are used to optimize parameters to train the model. The entire combined network model is trained end-to-end, and the trainable parts of the model are gradually optimized during backpropagation of gradients. In this embodiment, LoRA technology is used to fine-tune the pre-trained large language backbone network. The trainable part during training is the vision-language mapper. LoRA parameters in the large language backbone network F, language-visual mapper Visual decoder Parameters and weighting factors of the central attention block .
[0096] The definition of the loss function is formally explained in mathematical language as follows:
[0097] To ensure the quality of text generation, autoregressive cross-entropy loss is used as... :
[0098] = - (9)
[0099] Where Y={ } represents the actual text label sequence in the dataset. This represents the conditional probability given by the large language backbone network at time step t, which is the probability of predicting the label at the current time step based on the known real text label sequence Y. The predicted probability at each time step is evaluated and accumulated by summing the corresponding probabilities of the true labels. This method allows the model to learn how to generate new outputs using previous output information during training. To ensure the quality of mask generation, The definition is as follows:
[0100] = (10)
[0101] in This represents the pixel-wise binary cross-entropy loss, where H and W represent the height and width of the mask, respectively. This represents the binary value of the i-th pixel in the predicted mask image. This represents the ground truth mask in the dataset. The total loss is defined as:
[0102] = + (11)
[0103] in and This is a hyperparameter.
[0104] The input data for a multimodal large model is the defect image to be detected. and text prompts . After being converted into an embedding, the defect image is directly input into the large language backbone network. After passing through encoder E( Visual features are obtained by extracting features. = Then, visual features The last layer of features Then pass it to the visual-language mapper In the middle, through ( After aligning the features from the visual feature space to the text feature space, we obtain... The data is then input into the large language backbone network. Simultaneously, defect images... After tokenization, the tokens that best describe the visual features at each scale are extracted from the vocabulary of the large language backbone network model for each layer of features. Then, the L×N tokens are converted into embeddings and input into the large language backbone network.
[0105] After inference, the large language backbone network will output a sequence of tokens. }. A subsequence at a certain position { } refers to tokens related to visual features, that is = This is used to generate visual features in subsequent steps. For this part of the tokens, each token group is fused through a linear layer to obtain { }back,{ After language-visual mapper Transformed into a feature space that decoder D can understand, i.e. = ( For text-related tokens, { The text is decoded into a text description. Decoder D receives the aligned features. With the output of the last layer of encoder E After being input, it will output a defect detection mask.
[0106] Example 2
[0107] Based on the same inventive concept, this invention also discloses a system for detecting apparent defects based on a multimodal large model, comprising:
[0108] Visual encoder The input defect image is encoded, visual features are extracted, and the encoder assigns L×N tokens, where L represents the number of visual scales and N represents the number of tokens assigned to each visual scale.
[0109] Visual-Language Mapper Align the encoded visual features to the language feature space;
[0110] The large language backbone network F receives text descriptions of defect images, multi-scale token groups, and visual features aligned to the language feature space as input. After processing, it obtains language type tokens and visual type tokens, and decodes the language type tokens into text descriptions and evaluations of defect images.
[0111] Language-visual mapper Align visual type tokens to the visual feature space;
[0112] Visual decoder : Decode the encoded visual features and the aligned visual type token to output the final detection mask.
[0113] Since the system described in Embodiment 2 of the present invention is the system used to implement the appearance defect detection method based on multimodal large model in Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the electronic device based on the method described in Embodiment 1 of the present invention, and therefore will not be described again here.
[0114] Example 3
[0115] Based on the same inventive concept, the present invention also provides an electronic device, including one or more processors; a storage device for storing one or more programs; and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in Embodiment 1.
[0116] Since the device described in Embodiment 3 of this invention is the electronic device used in implementing the appearance defect detection method based on a multimodal large model in Embodiment 1 of this invention, those skilled in the art can understand the specific structure and variations of this electronic device based on the method described in Embodiment 1 of this invention, and therefore will not be repeated here. All electronic devices used in any method of this invention fall within the scope of protection of this invention.
[0117] Example 4
[0118] Based on the same inventive concept, the present invention also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method described in Embodiment 1.
[0119] Since the device described in Embodiment 4 of this invention is a computer-readable medium used to implement the apparent defect detection method based on a multimodal large model in Embodiment 1 of this invention, those skilled in the art can understand the specific structure and variations of this electronic device based on the method described in Embodiment 1 of this invention, and therefore will not be repeated here. All electronic devices used in any method of this invention fall within the scope of protection of this invention.
[0120] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.
Claims
1. The surface defect detection method based on multi-modal large model is characterized by: The following steps are involved: Annotating a training data set, wherein the training data set includes a defect image, a mask corresponding to the defect image, and a text description; Encode the defect image, extract visual features, and assign L×N tokens to the encoder, where L represents the number of visual scales and N represents the number of tokens assigned to each visual scale; Align the encoded visual features to the language feature space; Using a large language backbone network, during training, the text description of the defect image, the multi-scale token group, and the visual features aligned to the language feature space are input. After processing, the language type token and the visual type token are obtained, and the language type token is decoded into the text description and evaluation of the defect image; Align visual type tokens to visual feature space; Decode the encoded visual features and aligned visual type tokens to obtain the final detection mask image; The multimodal large model consisting of an encoder, language feature alignment with visual features, a large language backbone network, visual feature alignment with language features, and a decoder is trained, and defect detection is performed using the trained multimodal large model; when the multimodal large model is trained, the following loss is used: = + in, represents the autoregressive cross entropy loss, and is a hyperparameter; = - Y={ } is the real text label sequence in the dataset, To predict the probability of the current label based on the known true text label sequence Y; = represents the pixel-by-pixel binary cross entropy loss, H and W represent the height and width of the mask respectively. represents the binary value of the i-th pixel in the predicted mask image, Represents the ground truth masks in the dataset.
2. The surface defect detection method based on multimodal large model according to claim 1 is characterized in that: The text description in the training data set includes the type, shape, size of the defect, the cause of the defect and the repair suggestion for the defect.
3. The surface defect detection method based on multimodal large model according to claim 1 is characterized in that: GPT-4V is used to generate annotated training datasets.
4. The surface defect detection method based on multimodal large model according to claim 1 is characterized in that: The alignment of the visual features to the language features and the alignment of the visual type token to the visual features all adopt the same MLP structure, specifically including a linear layer, a ReLU layer, a linear layer and a Dropout layer.
5. The surface defect detection method based on multimodal large model according to claim 1 is characterized in that: The encoder uses a pre-trained CLIP-ViT encoder; The large language backbone network adopts the pre-trained Vicuna-v1.5-13B. The encoder and the large language backbone network are in a frozen state during the multimodal large model training, and the parameters are not updated.
6. The surface defect detection method based on multimodal large model according to claim 1 is characterized in that: The decoder has a total of L layers, each of which adopts a self-attention layer structure of a multi-dimensional convolutional head transposed attention and a gated convolutional feedforward network.
7. A system for implementing the surface defect detection method based on a multimodal large model as described in any one of claims 1 to 6, characterized in that: include: Vision Encoder : Encode the input defect image and extract visual features. The encoder is assigned L×N tokens, where L represents the number of visual scales and N represents the number of tokens assigned to each visual scale; Vision-Language Mapper : Align the encoded visual features to the language feature space; Large language backbone network F: receives the defect image text description, multi-scale token group and visual features aligned to the language feature space as input, obtains language type tokens and visual type tokens after processing, and decodes the language type tokens into text description and evaluation of the defect image; Language-Visual Mapper : Align visual type tokens to visual feature space; Visual Decoder : Decode the encoded visual features and the aligned visual type tokens, and output the final detection mask.
8. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When one or more programs are executed by the one or more processors, the one or more processors implement the surface defect detection method as described in any one of claims 1-6.
9. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the surface defect detection method as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Controllable defect image generation method and equipment based on visual semantic fusion
CN118097318A
Welding radiograph defect detection method and device, computer equipment and medium
CN118314123A