An unsupervised cylinder head inner cavity surface defect detection method for MiniGPT-4 large visual language model

By combining the MiniGPT-4 large-scale visual language model with multi-scale local neighborhood aggregation and mutual scoring mechanisms, the problems of model complexity and local feature sensitivity in unsupervised detection in industrial scenarios are solved, achieving efficient and accurate detection of defects on the inner surface of automotive cylinder heads without the need for annotation and thresholds.

CN119942192BActive Publication Date: 2025-10-24SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510003523.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-10-24
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

Existing unsupervised anomaly detection methods suffer from complex model design and insufficient deployment capabilities in industrial scenarios. They also lack sensitivity to local features and fail to fully utilize the implicit normal prior information of unlabeled images, thus limiting the improvement of detection accuracy.

Method used

By employing the MiniGPT-4 large-scale visual language model combined with a multi-scale local neighborhood aggregation method, a mutual scoring mechanism, a cue learner, and a geometric information learner, detailed abnormal region attributes and geometric information are generated, enabling unsupervised detection of surface defects in the inner cavity of automotive cylinder heads.

Benefits of technology

It improves detection efficiency and accuracy, reduces reliance on defective samples, and enables multi-round interactive detection without manual annotation and threshold settings, making it suitable for quality control in complex manufacturing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942192B_ABST
    Figure CN119942192B_ABST
Patent Text Reader

Abstract

The application discloses a MiniGPT-4-oriented unsupervised automobile cylinder cover inner cavity surface defect detection method, which comprises the following steps: constructing an automobile cylinder cover inner cavity surface parameterized imaging model and collecting normal images; constructing an image dataset, generating abnormal images from the normal images to form the dataset; collecting image features of the dataset, extracting the features through a multi-scale local neighborhood aggregation method and a mutual scoring mechanism; establishing an unsupervised automobile cylinder cover inner cavity surface defect detection model, which comprises an image encoder, an image decoder, a picture-text matching module, a prompt learner, a geometric information learner and a large language model; inputting the images into the image encoder, the image decoder, the picture-text matching module, the prompt learner, the geometric information learner and the large language model to determine a detection image mask mask image and defect attributes and geometric information. The application has great practical engineering application value for product quality control in a complex manufacturing scene and for improving defect detection efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of nondestructive testing, and particularly relates to an unsupervised automobile cylinder head inner cavity surface defect detection method for a MiniGPT-4 large visual language model. BACKGROUND

[0002] Industrial anomaly detection (IAD) is an important part of product quality control in modern manufacturing, and is widely used in part processing, product assembly and surface defect detection. Existing unsupervised anomaly detection methods include image reconstruction, generative model and feature embedding paradigm, although they perform well in some scenarios, but there are problems such as complex model design, insufficient deployment capability, limited description of abnormal region attributes, etc. In addition, the implicit normal prior information of unlabeled images in industrial scenarios is not fully utilized, which limits the improvement of detection accuracy.

[0003] In recent years, large visual language models have shown excellent performance in general visual tasks through alignment of visual features and text features. Their semantic understanding, zero-shot learning and multi-round interaction capabilities provide new possibilities for unsupervised industrial anomaly detection. However, these models have limited prior knowledge in the industrial field and are not sensitive enough to local features, and still need to be optimized for industrial scenarios.

[0004] Therefore, the present application proposes an unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4, which combines a multi-scale local neighborhood aggregation method, a mutual scoring mechanism, a prompt learner and a geometric information learner, uses implicit normal prior information to improve detection performance, and generates detailed abnormal region attributes and geometric information descriptions to realize efficient interactive detection. The present application has great practical engineering application value for product quality control in complex manufacturing scenarios and improving defect detection efficiency and accuracy.

[0005] In the prior art of the automobile cylinder head inner cavity surface defect detection method, the following several patents and documents are compared:

[0006] Patent CN 118470014 A discloses an industrial anomaly detection method and system, which comprises the following steps: first, acquiring an image to be anomaly detected; based on the acquired image, an anomaly detection model pre-trained is used to detect and locate the anomaly; the anomaly detection model comprises a backbone network, a pooling layer, a cascade flow, a twin flow and a constant flow; the cascade flow comprises a plurality of flow blocks arranged in sequence. The present application is different from the above, and only normal samples need to be input in the training stage, without manual labeling, solving the problems of few defect samples and difficult standards in industrial scenarios.

[0007] Patent CN 118568650 A discloses an industrial anomaly detection method and system based on fine-grained text prompt feature engineering, including the following steps: extracting text features, image block features and image features of industrial images; using image features to optimize and update text prompts to obtain fine-grained text prompt features; comparing and analyzing the similarity of image block features and fine-grained text prompt features, adding and fusing the abnormal result images generated by comparison to obtain the final abnormal detection result; optimizing the parameters of the model to minimize the loss function, and using the trained model to test the abnormal detection of the test set. The present application is different from the above, which is not based on fine-grained text features, but generates natural language prompts based on industrial domain knowledge through a prompt learner, and combines a large language model to perform multi-round interactive detection on prompt embedding, achieving more intuitive and detailed abnormal description.

[0008] Liu Yongjiang et al. [1] In "Pixel-level unsupervised industrial anomaly detection based on multi-scale memory bank" published in Computer Application in March 2024, a pre-trained feature extraction network is used to extract features from normal samples in the training set, and three scale positive sample feature memory banks are constructed. When training the segmentation network, the difference features between the pseudo abnormal sample features and the nearest positive sample features in the memory bank are calculated to further guide the segmentation network to learn how to locate abnormal pixels. However, this method lacks sufficient description of defect attribute information, and a threshold still needs to be manually set to distinguish between normal and abnormal samples.[1]Liu Yongjiang, Chen Bin. Pixel-level unsupervised industrial anomaly detection based on multi-scale memory bank [J / OL]. Computer Application, 1-9. SUMMARY

[0009] To solve the above technical problems, the present application proposes a MiniGPT-4 oriented unsupervised automobile cylinder cover inner cavity surface defect detection method.

[0010] The purpose of the present application is achieved by the following technical solutions:

[0011] A MiniGPT-4 oriented unsupervised automobile cylinder cover inner cavity surface defect detection method, comprising:

[0012] A parameterized imaging model of the automobile cylinder cover inner cavity surface is constructed, and the automobile cylinder cover inner cavity surface image is collected;

[0013] B Construct an image dataset, generate simulated abnormal images from normal images, and jointly form an image dataset from normal images and abnormal images;

[0014] C Collect image-level features of the dataset, extract features through a multi-scale local neighborhood aggregation method and a mutual scoring mechanism;

[0015] D Establish an unsupervised automobile cylinder cover inner cavity surface defect detection model;

[0016] E Determine the detection image mask and defect attribute, geometric information.

[0017] Compared with the prior art, one or more embodiments of the present application can have the following advantages:

[0018] The present application can apply a pre-trained model on a large data set to a data set of the inner surface of a cylinder head of a vehicle, improve the image feature extraction accuracy, only use normal sample images in the training stage, solve the problem of few defect samples in an industrial scene, and complete industrial anomaly detection without human annotation of defects; At the same time, without manually setting a threshold value, a user can realize multi-round interactive anomaly detection according to a model answer and further inquiry. The present application has great practical engineering application value for reducing the work intensity of the inner surface defect detection of the cylinder head of the vehicle, guaranteeing quality control in a complex manufacturing scene, and improving the defect detection efficiency and accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 is an unsupervised automobile cylinder head inner surface defect detection method flow chart for MiniGPT-4;

[0020] Figure 2 is a principle diagram of the unsupervised automobile cylinder head inner surface defect detection method for MiniGPT-4;

[0021] Figure 3 is an abnormal sample generated by applying a Cut-paste+Poisson Editing method and a Grount Truth comparison diagram;

[0022] Figure 4 is an abnormal image feature map extracted by a multi-scale local neighborhood aggregation method and a mutual scoring mechanism and a corresponding mask after thresholding;

[0023] Figure 5a 、 5b is a normal image and defect image detection result of the model on the MVTec-AD and VisA public data set. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below in combination with embodiments and drawings.

[0025] As shown in Figure 1 , it is an unsupervised automobile cylinder head inner surface defect detection method flow chart for MiniGPT-4, which includes:

[0026] Step 10, constructing an automobile cylinder head inner surface parameterized imaging model, collecting automobile cylinder head inner surface images;

[0027] Step 20 constructs an image data set, generates simulated abnormal images from normal images, and the normal images and abnormal images together constitute the image data set;

[0028] Step 30 collects image-level features of the dataset and extracts features using a multi-scale local neighborhood aggregation method and a mutual scoring mechanism;

[0029] Step 40: establishing an unsupervised automobile cylinder head inner cavity surface defect detection model;

[0030] Step 50 determines the detection image mask and defect attributes and geometric information.

[0031] The above step 10 specifically includes:

[0032] Based on the inspection requirements and process flow, parametric imaging models (PIM) of the inner surface of the automobile cylinder head are constructed, and the optimal parameters are determined.

[0033] Based on defect detection accuracy Number of parts n, expected detection time T(s) of the system, moving speed v(m / s), camera sampling frame rate F frames / s, camera focal length f, camera imaging field size H(mm)×W(mm), size S of CCD chip in the i∈(W,H) direction i , the pixel resolution P of the camera in the i direction i (pixel), the distance between the camera and the inner surface of the car cylinder head L O (mm) Construction of PIM on the inner surface of automobile cylinder head:

[0034]

[0035] Combining equations (1) and (2) we can calculate L O :

[0036]

[0037] The above step 20 specifically includes:

[0038] Acquire normal image of the inner surface of automobile cylinder head according to the optimal imaging parameters I normal , construct an image dataset, and the normal image I normal Generate simulated abnormal image I anomaly , normal images and abnormal images together constitute the image dataset D G ={I normal ,I anomaly};

[0039] The simulated abnormal image I anomalyGenerated by Cut Paste+Possion Editing method:

[0040]

[0041] wherein, represents a random cutting region pasting operation, Poisson(·) represents Poisson editing to achieve natural fusion. Cut-Paste randomly cuts a rectangular region Patch from a normal image, the size and position of which are determined by the following formula:

[0042]

[0043] wherein, W and H are the width and height of the image respectively, (x, y) is the left upper corner coordinate of the cutting region, w, h are the width and height of the cutting region. The cut Patch is randomly pasted into the target position (x', y'), to obtain the preliminary generated abnormal image I paste :

[0044] I paste (x':x'+w, y':y'+h) = Patch, x', y' ∈ [0, W-H] (6)

[0045] Poisson Editing fuses the boundary of the pasted region in I paste , solves the edge discontinuity problem caused by direct pasting. Poisson Editing is implemented by the following formula:

[0046]

[0047] wherein, Δ represents the Laplacian operator, Ω is the pasted region, is the boundary of the pasted region. Poisson Editing adjusts the pixel value by solving the Laplace equation to make the pasted region and the background naturally fused, to generate the final abnormal image I anomaly .

[0048] The above step 30 specifically comprises:

[0049] The VIT-L-14-336 model is used as a feature extractor, and the input image I ∈ D G is divided into blocks, and patch-level features F are extracted through a multi-scale local neighborhood aggregation method and a mutual scoring mechanism; wherein the multi-scale local neighborhood aggregation method is that the image dataset D G = {I normal ,I anomaly} is extracted by the feature extractor ViT-L-14-336 to obtain image patch tokens Where l∈{1,2…,L} represents the different levels of the feature extractor ViT, Q and C are the number of patches and the number of image channels respectively. The layer normalization processing is performed on the features of each scale:

[0050]

[0051] Among them, r = {1, 3, 5} is the characteristic aggregation degree, μ i,r , σ i,r are the mean and standard deviation of the feature map of the i-th sample at scale r, and ε is a small constant to avoid the denominator being zero.

[0052] Then aggregated by adaptive average pooling method The r×r neighborhood features, after aggregation, the feature output is The shape is restored to Q×C, and after each polymerization The aggregated feature patch token of each image is Where q∈[1,Q]. Multi-scale local neighborhood aggregation method outputs aggregated features to the mutual rating mechanism.

[0053] The mutual rating mechanism includes:

[0054] ① Feature similarity calculation: Let the image to be detected I i The patch token is Combine it with D G In addition to I i All images of I j patch token Perform similarity score Calculation, the expression is as follows:

[0055]

[0056] If I i of with I j Any The more similar, the The lower the score, the more likely it is to be in the normal range.

[0057] ② Similarity score enhancement: There may be a small amount of feature differences in normal areas. In order to amplify the score difference between normal and abnormal areas, the interval average method can be used to enhance the similarity score. There are K images with the lowest scores of 30%, and the kth image is represented as Then define the anomaly score of the region for:

[0058]

[0059] ③ Multi-scale score fusion: In order to integrate multi-scale feature information, different aggregation degrees can be combined The similarity scores at different levels L are averaged and fused, and the overall anomaly score u i,q It is expressed as follows:

[0060]

[0061] Definition D G The overall anomaly score set is Convert U to shape Then upsample to the original image size to obtain image I i Feature Results Then D G The result of all image features is Φ=[φ1,φ2…,φ G ] T .

[0062] The above step 40 specifically includes:

[0063] The unsupervised automotive cylinder head defect detection model includes:

[0064] Image Encoder: ImageBind-Huge of MiniGPT-4 is used as ImageEncoder to extract global and local features of the input image;

[0065] Image Decoder: The Image Decoder uses a four-stage lightweight, feature-matching structure to extract corresponding intermediate patch-level features at each stage. The input image and preset text are encoded by the ImageEncoder and TextEncoder respectively, and then input into the Image Decoder, which then outputs the anomaly detection results. The Image Encoder and Text Encoder are the same, both using MiniGPT-4's ImageBind-Huge.

[0066] Graphic Text Matching Module (GTMM): Aligns image features with text features. The intermediate patch-level features extracted in the i-th stage are recorded as However, these features have not been image-text aligned and cannot be compared with text features. We can introduce an additional linear layer to project the intermediate features onto Align it with text features representing normal and abnormal semantics, and the abnormal positioning results Upsample means up-sampling, then:

[0067]

[0068] Prompt Learner (PL): the image decoder output is integrated into (n1+n2) prompt embedding vectors and input into the LLM together with the image features, fully utilizing the fine-grained semantic information in the image and keeping the semantic matching between the image decoder and the LLM;

[0069] Geometry Information Learner (GIL): GIL realizes accurate recognition and extraction of abnormal geometric information in the input image, with the following steps:

[0070] ① Image standardization: the input abnormal positioning result I is converted into a grayscale image I gray , and then standardized into a 224x224 pixel size I norm to ensure consistency of different size image analysis specifications;

[0071] ② Region segmentation: I norm is processed by an adaptive threshold algorithm to obtain I binary ;

[0072] ③ Geometric information extraction: contour recognition is performed on I binary by Canny edge detection to obtain an image edge function δ(E(x,y), edge differential segment ds. Let the image abnormal area be R i , then its abnormal area A anomaly , perimeter P anomaly , and center coordinates CC anomaly (x,y) are calculated as follows:

[0073]

[0074] Large Language Model (LLM): MiniGPT-4's Vicuna-7B is used as the LLM, and the prompt embedding is input into the LLM to generate the detection result R through the following query:

[0075] R = M LLM (Q, E prompt ) (14)

[0076] The above step 50 specifically includes:

[0077] ​Image I is input into the image encoder, image decoder, image-text matching module, prompt learner, geometric information learner, and large language model. The model training process selects Cross-Entropy Loss, Focal Loss, and Dice Loss as the loss functions to train Image Decoder and PL:

[0078] Cross-Entropy Loss is commonly used to train language models and is a loss function that measures the difference between the predicted probability distribution and the true distribution in classification problems:

[0079]

[0080] where n is the number of image feature patch tokens, y i is the true label of each patch token, p i is the predicted probability of each patch token.

[0081] Focal Loss introduces a modulation factor γ based on Cross-Entropy Loss, which reduces the loss weight of easy-to-classify samples and increases the loss weight of difficult-to-classify samples, so that the model pays more attention to difficult-to-classify samples during training:

[0082]

[0083] where N P is the total number of image pixels, is the predicted value of each pixel, α t is the positive and negative sample proportion weight balance factor, and the modulation factor γ is set to 2 during training.

[0084] Dice Loss is mainly used to evaluate the similarity of binary images:

[0085]

[0086] where N P is the total number of image pixels, z i , are the true value and predicted value of each pixel, respectively.

[0087] By combining equations (15) to (17), the overall loss function of the model can be obtained:

[0088] L total = L ce + L focal + L dice (18)

[0089]

Example 1

[0090] Figure 3 is a simulated abnormal sample generated by the Cut-paste+Poisson Editing method and a Ground Truth comparison chart. The simulated abnormal sample is generated by the Cut-paste+Poisson Editing method, and the generated image has continuous and natural characteristics;

[0091] Figure 4 is the abnormal image feature map extracted by the multi-scale local neighborhood aggregation method and the mutual scoring mechanism, and the corresponding mask after thresholding, which uses the implicit normal prior information in the unlabeled test image to assist in detecting abnormalities and improve the detection accuracy of the model;

[0092] Figure 5a 、 5b is the detection result of the model on the MVTec-AD and VisA public data sets, which can achieve accurate abnormal detection without manual threshold setting, and accurately describe the abnormal area attributes and geometric information.

[0093] Although the embodiments of the present application are as described above, the content described is only for the purpose of facilitating the understanding of the present application, and is not intended to limit the present application. Any person skilled in the art to which the present application belongs can make any modification and change in the form and details without departing from the spirit and scope of the present application. The patent protection scope of the present application shall be subject to the scope defined in the appended claims.

Claims

1. A MiniGPT-4-oriented unsupervised cylinder head inner cavity surface defect detection method, characterized in that, The method comprises: A. Constructing an automobile cylinder cover inner cavity surface parameterized imaging model, collecting automobile cylinder cover inner cavity surface images; B. Constructing an image data set, generating simulated abnormal images from normal images, and jointly forming an image data set from normal images and abnormal images; C. Collecting data set image level features, extracting features through multi-scale local neighborhood aggregation method and mutual scoring mechanism; D. Establishing an unsupervised automobile cylinder cover inner cavity surface defect detection model; E. Determining the detection image mask mask image and defect attribute, geometric information; In step D, the unsupervised automobile cylinder cover defect detection model comprises: Image encoder: MiniGPT-4 ImageBind-Huge is used as Image Encoder to extract global and local features of input images; Image decoder: Image Decoder adopts a four-stage lightweight and feature matching structure to extract corresponding intermediate patch-level features in each stage; the input image and the preset text are respectively encoded by Image Encoder and Text Encoder and then input into Image Decoder, and Image Decoder outputs the abnormal detection result; wherein Image Encoder and Text Encoder are the same, both are MiniGPT-4 ImageBind-Huge; Figure-text matching module: align image features with text features, and the intermediate patch-level features extracted in the i-th stage are denoted as The features are not aligned with image-text, and the intermediate features are projected to align them with text features representing normal and abnormal semantics, and the abnormal positioning results Upsample means up-sampling, and there is: Prompt learner: decode image encoder output with independently trained base vectors integrated into (n1+n2) prompt embedding vectors with image features jointly input into LLM, utilizing fine-grained semantic information in images, and keeping image decoder and LLM semantic matching; Geometric information outputter: to realize accurate identification and extraction of abnormal geometric information in the input image, the steps are as follows: ① Image standardization: the input abnormal positioning result R is converted into a gray image I gray , and then normalized to 224x224 pixels norm , to ensure the consistency of different size image analysis specifications; (ii) region segmentation: I norm I is processed by using adaptive threshold algorithm binary ; ③ Geometric information extraction: through Canny edge detection on I binary contour recognition, get image edge function δ(E(x,y), edge differential segment ds; let image abnormal area be R i , its abnormal area A anomaly , perimeter P anomaly , center coordinates CC anomaly (x,y) calculation formula: Large language model: MiniGPT-4's Vicuna-7B is used as the LLM, and the prompt is embedded into the input LLM to generate the detection result R through the following query text : R text = M LLM (Q, E prompt ) (14).

2. The MiniGPT-4 oriented unsupervised cylinder head inner cavity surface defect detection method of claim 1, wherein, In step A, According to the detection requirements and process flow, an automobile cylinder cover inner cavity surface parameterized imaging model PIM is constructed; According to defect detection accuracy The number of parts n, the system expected detection time T (s), the part moving speed v (m / s), the camera sampling frame rate F (frame / s), the camera focal length f, the camera imaging field of view size H (mm) x W (mm), the CCD chip size S in i ∈ (W, H) direction i , the pixel resolution of the camera in i direction P i (pixel), the distance L O (mm) between the camera and the surface of the inner cavity of the automobile cylinder head to construct the surface PIM of the inner cavity of the automobile cylinder head: Simultaneous equations (1), (2) can be calculated to obtain L O :

3. The MiniGPT-4 oriented unsupervised cylinder head inner cavity surface defect detection method of claim 1, wherein, The step B comprises: collecting a normal image I of the inner surface of the automobile cylinder head; normal , construct an image dataset, and the normal image I normal Generate simulated abnormal image I anomaly , normal images and abnormal images together constitute the image dataset D G ={I normal ,I anomaly }; The simulated abnormal image I anomaly By Cut Paste+Possion Editing method from normal image I normal Generated: where, represents the random cropping and pasting operation, Poisson(·) represents the Poisson editing to achieve natural fusion; Cut-Paste randomly crops a rectangular region Patch from the normal image, the size and position of which are determined by the following formula: where W and H are the width and height of the image respectively, (x, y) is the top-left coordinate of the cropped region, and w, h are the width and height of the cropped region; the cropped Patch is randomly pasted into the target position (x', y') to obtain the initially generated abnormal image I paste : I paste (x' : x' + w, y' : y' + h) = Patch, x', y' ∈ [0, W - H] (6) Poisson Editing is to I paste paste the boundary of the region to solve the edge discontinuity caused by direct paste; Poisson Editing is achieved by the following formula: where Δ denotes the Laplacian operator, Ω is the pasting region, is the boundary of the pasting region; Poisson Editing adjusts the pixel values by solving the Laplace equation to make the pasting region and the background naturally blend, and generates the final abnormal image I anomaly .

4. The MiniGPT-4 oriented unsupervised cylinder head inner cavity surface defect detection method of claim 1, wherein, In step C, the VIT-L-14-336 model is used as a feature extractor to extract patch-level features F of the input image I G The patch-level features F are extracted by a multi-scale local neighborhood aggregation method and a mutual scoring mechanism. Wherein A multiscale local neighborhood aggregation method for image dataset D G = {I normal , I anomaly Image patch tokens are extracted by a feature extractor ViT-L-14-336 Wherein l∈{1,2…,L} represents different levels of feature extractor ViT, Q and C are the number of patches and the number of image channels respectively; Layer normalization processing is performed on each scale feature to calculate: where r = {1, 3, 5} is the characteristic aggregation degree, are the feature map mean and standard deviation of the ith sample at the lth layer and scale r, respectively, and ε is a small constant to avoid division by zero. Then aggregate the r x r neighborhood features by adaptive average pooling method The output of the aggregated features is Reduce the shape to Q x C, and each aggregated The patch token of each image aggregated feature is Where q ∈ [1, Q]; the multi-scale local neighborhood aggregation method outputs aggregated features to the mutual scoring mechanism.

5. The MiniGPT-4 oriented unsupervised cylinder head inner cavity surface defect detection method of claim 4, wherein, In step C, the mutual scoring mechanism Comprises: ① Feature similarity computation: Let the patch token of the image I i to be detected be Compute the similarity score with the patch token of all images I G in D i except I j , denoted as The expression is as follows: ​ If I i of with I j any the more similar, the score is lower, the more normal region belongs to; ② Similarity score enhancement: For the possibility of a small feature difference in the normal region, in order to amplify the score difference between the normal region and the abnormal region, the interval average method is used to enhance the similarity score; let patch token The 30% lowest score images have K, and the kth image is represented as Then define the abnormal score of the region : ③Multi-scale score fusion: To integrate multi-scale feature information, the anomaly scores under different feature aggregation degrees r and different levels L are aggregated Average fusion is performed to obtain the overall anomaly score u i,q is expressed as follows: Definition D G The set of overall anomaly scores is U = [u i,1 ,…,u i,q ] T U is converted to shape After up-sampling to the original image size, the image I i Feature results D G The feature results for all images in D are Φ = [φ1,φ2…,φ G ] T .

6. The MiniGPT-4 oriented unsupervised cylinder head inner cavity surface defect detection method of claim 1, wherein, In step E, the image I is input into the image encoder, image decoder, image-text matching module, prompt learner, geometric information learner, and large language model, and the model training process selects Cross-Entropy Loss, Focal Loss, and Dice Loss three loss functions to train the image decoder and the prompt learner: Cross-Entropy Loss is used to train the language model, which is a loss function for measuring the difference between the predicted probability distribution and the true distribution in the classification problem: wherein n is the number of image feature patch tokens, y i is the true label for each patch token, p i is the predicted probability for each patch token; Focal Loss introduces a modulation factor γ based on Cross-Entropy Loss, reduces the loss weight of easy-to-classify samples, and increases the loss weight of difficult-to-classify samples, so that the model pays more attention to difficult-to-classify samples during training: Wherein, N P is the total number of image pixels, is the predicted value of each pixel point, α t is the positive and negative sample proportion weight balancing factor, and the modulation factor γ is set to 2 during the training process. Dice Loss is used to evaluate the similarity of binary images: where N P is the total number of image pixels, z i , are the true value and predicted value of each pixel, respectively. By simultaneously solving equations (15)-(17), the overall loss function of the model can be obtained: L total = L ce + L focal + L dice (18).

7. The MiniGPT-4 oriented unsupervised cylinder head inner cavity surface defect detection method of claim 1, wherein, The method adopts MiniGPT-4 ImageBind-Huge as an image encoder, Vicuna-7B as an LLM, and connects the visual language representations of the two through a linear layer to align them; the image resolution is set to 224x224, features are extracted from the 8th, 16th, 24th and 32nd layers of the ImageBind-Huge image encoder respectively and transmitted to the image decoder; the training round is 50, the learning rate is set to 1e-3, and the batch size is 16; in addition, a linear warm-up and periodic cosine learning rate decay strategy is adopted to optimize the training process.

8. The MiniGPT-4 oriented unsupervised cylinder head inner cavity surface defect detection method of claim 1, wherein, The method supports the user to make multi-round interactive query through the LLM, and the output abnormal area geometric information includes an abnormal area A anomaly , a perimeter P anomaly , and a center coordinate CC anomaly (x, y).

9. The MiniGPT-4 oriented unsupervised cylinder head inner cavity surface defect detection method of claim 1, wherein, The method is suitable for various industrial defect detection scenarios, including part machining, product assembly and surface defect detection.

Citation Information

Patent Citations

  • Undecimated wavelet and Gumbel distribution-based fabric defect detection method

    CN108399614A

  • Automobile camshaft mounting hole inner wall defect detection method with strong generalization capability

    CN118014931A