Unsupervised automobile cylinder cover inner cavity surface defect detection method facing MiniGPT-4 large visual language model
Through the unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4, combined with multi-scale local neighborhood aggregation method and mutual scoring mechanism and other technical means, the problems of limited complexity and accuracy of unsupervised abnormal detection methods in industrial scenarios are solved, and efficient and accurate defect detection is achieved.
Patent Information
- Application Number
- CN202510003523.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-02
AI Technical Summary
The existing unsupervised anomaly detection methods have problems in industrial scenarios such as complex model design, insufficient deployment capabilities, limited description of abnormal area attributes, and lack of use of implicit normal prior information, resulting in limited detection accuracy.
A method for detecting surface defects of the inner cavity of the cylinder head for MiniGPT-4 is proposed, combining multi-scale local neighborhood aggregation method, mutual scoring mechanism, prompt learner and geometric information learner, and using implicit normal prior information to improve detection performance, and generate detailed descriptions of abnormal area attributes and geometric information.
It realizes efficient interactive detection in complex manufacturing scenarios, improves defect detection efficiency and accuracy, reduces detection work intensity, and does not require manual labeling or manual threshold setting.
Smart Images

Figure CN119942192A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of nondestructive testing technology, and in particular to an unsupervised automobile cylinder head inner cavity surface defect detection method oriented to the MiniGPT-4 large-scale visual language model. Background Art
[0002] Industrial Anomaly Detection (IAD) is an important part of product quality control in modern manufacturing and is widely used in parts processing, product assembly and surface defect detection. Existing unsupervised anomaly detection methods include image reconstruction, generative models and feature embedding paradigms. Although they perform well in some scenarios, they have problems such as complex model design, insufficient deployment capabilities and limited description of abnormal area attributes. In addition, the implicit normal prior information of unlabeled images in industrial scenarios is not fully utilized, which limits the improvement of detection accuracy.
[0003] In recent years, large-scale visual language models have performed well in general visual tasks by aligning visual features with text features. Their semantic understanding, zero-shot learning, and multi-round interaction capabilities provide new possibilities for unsupervised industrial anomaly detection. However, these models have limited prior knowledge in the industrial field and are not sensitive enough to local features, so they still need to be optimized for industrial scenarios.
[0004] To this end, the present invention proposes an unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4, which combines a multi-scale local neighborhood aggregation method, a mutual scoring mechanism, a prompt learner and a geometric information learner, uses implicit normal prior information to improve detection performance, and generates detailed abnormal area attributes and geometric information descriptions to achieve efficient interactive detection. The present invention has significant practical engineering application value for product quality control in complex manufacturing scenarios and improving defect detection efficiency and accuracy.
[0005] In the prior art of the automobile cylinder head inner cavity surface defect detection method, there are several comparative patents and documents as follows:
[0006] Patent CN 118470014 A discloses an industrial anomaly detection method and system, including the following steps: first, obtain an image to be detected for anomaly; based on the acquired image, detect and locate the anomaly based on a pre-trained anomaly detection model; the anomaly detection model includes: a backbone network, a pooling layer, a cascade flow, a twin flow, and a constant flow; the cascade flow contains a number of flow blocks arranged in sequence. The present invention is different from the above. In the training stage, only normal samples need to be input, and no manual labeling is required, which solves the problem of few defect samples and difficult standards in industrial scenes.
[0007] Patent CN 118568650 A discloses an industrial anomaly detection method and system based on fine-grained text prompt feature engineering, including the following steps: extracting text features, image block features and image features of industrial images; optimizing and updating text prompts using image features to obtain fine-grained text prompt features; performing similarity comparison analysis on image block features and fine-grained text prompt features, adding and fusing the abnormal result graphs generated by the comparison to obtain the final abnormality detection results; optimizing the parameters of the model to minimize the loss function, and using the trained model to perform abnormality detection on the test set. The present invention is different from the above, and is not based on fine-grained text features, but generates natural language prompts based on industrial domain knowledge through a prompt learner, and performs multiple rounds of interactive detection on prompt embedding in combination with a large language model to achieve a more intuitive and detailed abnormal description.
[0008] Liu Yongjiang et al. [1] The article "Pixel-level unsupervised industrial anomaly detection based on multi-scale memory bank" was published in the March 2024 issue of "Computer Applications". This article extracts features from normal samples in the training set through a pre-trained feature extraction network and constructs a positive sample feature memory bank at three scales. When training the segmentation network, the difference features are calculated by simulating the features of pseudo-abnormal samples and the features of the closest positive samples in the memory bank, which further guides the segmentation network to learn how to locate abnormal pixels. However, this method does not adequately describe the defect attribute information, and it is still necessary to manually set the threshold to distinguish between normal and abnormal samples. [1] Liu Yongjiang, Chen Bin. Pixel-level unsupervised industrial anomaly detection based on multi-scale memory bank [J / OL]. Computer Applications, 1-9. Summary of the invention
[0009] In order to solve the above technical problems, the present invention proposes an unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4.
[0010] The purpose of the present invention is achieved through the following technical solutions:
[0011] An unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4, comprising:
[0012] A. Construct a parametric imaging model of the inner surface of the automobile cylinder head and collect the inner surface image of the automobile cylinder head;
[0013] B. constructing an image data set, generating simulated abnormal images from normal images, and the normal images and abnormal images together constitute an image data set;
[0014] C collects image-level features of the dataset and extracts features through multi-scale local neighborhood aggregation method and mutual scoring mechanism;
[0015] D. Establish an unsupervised automobile cylinder head inner cavity surface defect detection model;
[0016] E. Determine the detection image mask, defect attributes, and geometric information.
[0017] Compared with the prior art, one or more embodiments of the present invention may have the following advantages:
[0018] The present invention can apply the pre-trained model on a large data set to the data set of the inner cavity surface of the automobile cylinder head, improve the accuracy of image feature extraction, and only need to use normal sample images in the training stage to solve the problem of few defect samples in industrial scenes, and can complete industrial anomaly detection without manual defect labeling; at the same time, there is no need to manually set the threshold, and the user can ask again according to the model answer to realize multiple rounds of interactive anomaly detection. The use of the present invention has great practical engineering application value in reducing the workload of defect detection on the inner cavity surface of automobile cylinder heads, ensuring quality control in complex manufacturing scenes, and improving defect detection efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a flow chart of the unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4;
[0020] Figure 2 It is the schematic diagram of the unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4;
[0021] Figure 3 It uses the Cut-paste+Poisson Editing method to generate abnormal samples and Grount Truth comparison charts;
[0022] Figure 4 It is the abnormal image feature map extracted by the multi-scale local neighborhood aggregation method and the mutual scoring mechanism and the corresponding mask after thresholding;
[0023] Figure 5a , 5b These are the normal image and defect image detection results of the model on the MVTec-AD and VisA public datasets. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below in conjunction with embodiments and drawings.
[0025] like Figure 1 As shown, it is a flow chart of the unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4, including:
[0026] Step 10: construct a parametric imaging model of the inner surface of the automobile cylinder head, and collect an image of the inner surface of the automobile cylinder head;
[0027] Step 20 constructs an image data set, generates simulated abnormal images from normal images, and the normal images and abnormal images together constitute the image data set;
[0028] Step 30 collects image-level features of the dataset and extracts features through a multi-scale local neighborhood aggregation method and a mutual scoring mechanism;
[0029] Step 40: establishing an unsupervised automobile cylinder head inner cavity surface defect detection model;
[0030] Step 50 determines the detection image mask and defect attributes and geometric information.
[0031] The above step 10 specifically includes:
[0032] According to the detection requirements and process flow, a parametric imaging model (PIM) of the inner surface of the automobile cylinder head is constructed, and the optimal parameters are determined.
[0033] Based on defect detection accuracy Number of parts n, expected detection time T(s), part moving speed v(m / s), camera sampling frame rate F frames / s, camera focal length f, camera imaging field size H(mm)×W(mm), CCD chip size S in the i∈(W,H) direction i , the pixel resolution P of the camera in the i direction i (pixel), the distance between the camera and the inner surface of the cylinder head of the car is L O (mm) Construction of PIM on the inner surface of automobile cylinder head:
[0034]
[0035] Combining equations (1) and (2) we can calculate L O :
[0036]
[0037] The above step 20 specifically includes:
[0038] Acquire normal image of the inner surface of automobile cylinder head according to the optimal imaging parameters I normal , construct an image dataset, and consist of normal images I normal Generate simulated abnormal image I anomaly , normal images and abnormal images together constitute the image dataset D G = {I normal ,I anomaly};
[0039] The simulated abnormal image I anomalyGenerated by Cut Paste+Possion Editing method:
[0040]
[0041] in, Represents the random cropping and pasting operation, and Poisson(·) represents Poisson editing to achieve natural fusion. Cut-Paste randomly crops a rectangular area Patch from the normal image, and its size and position are determined by the following formula:
[0042]
[0043] Where W and H are the width and height of the image, (x, y) is the coordinate of the upper left corner of the cropped area, and w and h are the width and height of the cropped area. The cropped patch is randomly pasted to the target position (x', y') to obtain the initial abnormal image I paste :
[0044] I paste (x′:x′+w,y′:y′+h)=Patch, x′,y′∈[0,WH] (6)
[0045] Poisson Editing Fusion I paste The Poisson Editing operation is performed on the border of the pasted area to solve the edge discontinuity problem caused by direct pasting. Poisson Editing is implemented by the following formula:
[0046]
[0047] Among them, Δ represents the Laplace operator, Ω is the pasting area, The PoissonEditing solves the Laplace equation and adjusts the pixel value to make the pasted area blend naturally with the background to generate the final abnormal image I anomaly .
[0048] The above step 30 specifically includes:
[0049] The VIT-L-14-336 model is used as the feature extractor, and the input image I∈D G The image is divided into blocks and the image patch-level features F are extracted through the multi-scale local neighborhood aggregation method and the mutual scoring mechanism; the multi-scale local neighborhood aggregation method is an image dataset D G = {I normal ,I anomaly}Extract image patch tokens through feature extractor ViT-L-14-336 Where l∈{1,2…,L} represents different levels of the feature extractor ViT, Q and C are the number of patches and the number of image channels respectively. The layer normalization processing is performed on each scale feature:
[0050]
[0051] Among them, r = {1, 3, 5} is the characteristic aggregation degree, μ i,r , σ i,r are the mean and standard deviation of the feature map of the ith sample at scale r, and ε is a small constant to avoid the denominator being zero.
[0052] Then aggregated by adaptive average pooling method The r×r neighborhood features, after aggregation, the feature output is The shape is restored to Q×C, and each polymerization The aggregated feature patch token of each image is Where q∈[1,Q]. Multi-scale local neighborhood aggregation method outputs aggregated features to a mutual rating mechanism.
[0053] The mutual rating mechanism includes:
[0054] ① Feature similarity calculation: Let the image to be detected I i The patch token is Combine it with D G In addition to I i All images of I j Patch token Similarity score Calculation, the expression is as follows:
[0055]
[0056] If I i of with I j Any The more similar, the The lower the score, the more likely it is to be in the normal zone.
[0057] ② Similarity score enhancement: There may be a small number of feature differences in normal areas. In order to amplify the score difference between normal areas and abnormal areas, the interval average method can be used to enhance the similarity score. There are K images with the lowest 30% scores, and the kth image is represented as Then define the anomaly score of the region for:
[0058]
[0059] ③ Multi-scale score fusion: In order to integrate multi-scale feature information, different aggregation degrees can be The similarity scores at different levels L are averaged and fused, and the overall anomaly score u i,q It is expressed as follows:
[0060]
[0061] Definition D G The overall anomaly score set is Convert U to shape Then upsample to the original image size to obtain image I i Feature Results Then D G The result of all image features is Φ=[φ1,φ2…,φ G ] T .
[0062] The above step 40 specifically includes:
[0063] The unsupervised automobile cylinder head defect detection model includes:
[0064] Image Encoder: ImageBind-Huge of MiniGPT-4 is used as ImageEncoder to extract global and local features of the input image;
[0065] Image Decoder: Image Decoder adopts a four-stage lightweight, feature matching structure to extract corresponding intermediate patch-level features at each stage. The input image and preset text are encoded by ImageEncoder and Text Encoder respectively and then input into Image Decoder, which then outputs the anomaly detection result. Image Encoder and Text Encoder are the same, both are ImageBind-Huge of MiniGPT-4;
[0066] Graphic Text Matching Module (GTMM): Align image features with text features. The intermediate patch-level features extracted in the i-th stage are recorded as However, these features have not been image-text aligned and cannot be compared with text features. We can introduce an additional linear layer to project the intermediate features to Align it with the text features representing normal and abnormal semantics, and the abnormal location results Upsample means upsampling, then:
[0067]
[0068] Prompt Learner (PL): Output of the image decoder With independently trained basis vectors Integrate into (n1+n2) hint embedding vectors The image features are input into LLM together to fully utilize the fine-grained semantic information in the image and keep the image decoder and LLM semantics matching;
[0069] Geometry Information Learner (GIL): GIL implements accurate anomaly identification and extraction of anomaly geometry information in the input image. The steps are as follows:
[0070] ① Image standardization: Convert the input anomaly location result I into a grayscale image I gray , and then to the standardized I norm , ensuring the consistency of analysis specifications for images of different sizes;
[0071] ②Region segmentation: I norm The adaptive threshold algorithm is used to process I binary ;
[0072] ③Geometric information extraction: I binary Perform contour recognition and obtain the image edge function δ(E(x,y), edge differential line segment ds. Let the abnormal area of the image be R i , then its abnormal area A anomaly , perimeter P anomaly , center coordinate CC anomaly (x,y) calculation formula:
[0073]
[0074] Large Language Model (LLM): MiniGPT-4’s Vicuna-7B is used as the LLM. The prompt embedding is input into the LLM. The detection result R is generated by the following query:
[0075] R=M LLM (Q,E prompt ) (14)
[0076] The above step 50 specifically includes:
[0077] The image I is input into the image encoder, image decoder, image-text matching module, prompt learner, geometric information learner, and large language model. The model training process uses three loss functions: Cross-Entropy Loss, Focal Loss, and Dice Loss to train Image Decoder and PL:
[0078] Cross-Entropy Loss is often used to train language models. It is a loss function that measures the difference between the predicted probability distribution and the true distribution in classification problems:
[0079]
[0080] Among them, n is the number of image feature patch tokens, y i The true label of each patch token, p i is the predicted probability for each patch token;
[0081] Focal Loss introduces a modulation factor γ based on Cross-Entropy Loss. By reducing the loss weight of easy-to-classify samples and increasing the loss weight of difficult-to-classify samples, the model pays more attention to difficult-to-classify samples during training:
[0082]
[0083] Among them, N P is the total number of image pixels, Predict the value for each pixel, α t is the weight balance factor of the positive sample and negative sample ratio, and the modulation factor γ is set to 2 during training.
[0084] Dice Loss is mainly used to evaluate the similarity of binary images:
[0085]
[0086] Among them, N P is the total number of image pixels, z i , are the true value and predicted value of each pixel respectively.
[0087] By combining equations (15) to (17), we can obtain the overall loss function of the model:
[0088] L total =L ce +L focal +L dice (18)
[0089] [Example 1]
[0090] Figure 3 The simulated abnormal samples and Grount Truth comparison images are generated by the Cut-paste+Poisson Editing method. The simulated abnormal samples are generated by the Cut-paste+Poisson Editing method, and the generated images have continuous and natural characteristics;
[0091] Figure 4 The method uses a multi-scale local neighborhood aggregation method and a mutual scoring mechanism to extract abnormal image feature maps and the corresponding masks after thresholding, and uses implicit normal prior information in unlabeled test images to assist in detecting abnormalities and improve the detection accuracy of the model.
[0092] Figure 5a , 5b It is the detection result of the model on the MVTec-AD and VisA public datasets. It can achieve accurate anomaly detection without manually setting thresholds and accurately describe the attributes and geometric information of the abnormal area.
[0093] Although the embodiments disclosed in the present invention are as above, the above contents are only embodiments adopted for facilitating the understanding of the present invention and are not intended to limit the present invention. Any technician in the technical field to which the present invention belongs can make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed in the present invention, but the patent protection scope of the present invention shall still be subject to the scope defined in the attached claims.
Claims
1. An unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4, characterized in that: The method comprises: A. Construct a parametric imaging model of the inner surface of the automobile cylinder head and collect the inner surface image of the automobile cylinder head; B. constructing an image data set, generating simulated abnormal images from normal images, and the normal images and abnormal images together constitute an image data set; C collects image-level features of the dataset and extracts features through multi-scale local neighborhood aggregation method and mutual scoring mechanism; D. Establish an unsupervised automobile cylinder head inner cavity surface defect detection model; E. Determine the detection image mask, defect attributes, and geometric information.
2. The unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4 as claimed in claim 1 is characterized in that: In the step A, Construct the parametric imaging model PIM of the inner surface of the automobile cylinder head according to the detection requirements and process flow; Based on defect detection accuracy Number of parts n, expected detection time T(s), part moving speed v(m / s), camera sampling frame rate F frames / s, camera focal length f, camera imaging field size H(mm)×W(mm), CCD chip size S in the i∈(W,H) direction i , the pixel resolution P of the camera in the i direction i (pixel), the distance between the camera and the inner surface of the cylinder head of the car is L O (mm) Construction of PIM on the inner surface of automobile cylinder head: Combining equations (1) and (2) we can calculate L O :
3. The unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4 as claimed in claim 1 is characterized in that: The step B comprises: collecting a normal image I of the inner surface of the automobile cylinder head normal , construct an image dataset, and consist of normal images I normal Generate simulated abnormal image I anomaly , normal images and abnormal images together constitute the image dataset D G = {I normal ,I anomaly }; The simulated abnormal image I anomaly By using the Cut Paste + Position Editing method, the normal image I normal generate: In the formula, Indicates the random cropping and pasting operation, Poisson(·) indicates Poisson editing to achieve natural fusion; Cut-Paste randomly crops a rectangular area Patch from the normal image, and its size and position are determined by the following formula: Where W and H are the width and height of the image, (x, y) is the coordinate of the upper left corner of the cropped area, and w and h are the width and height of the cropped area. The cropped patch is randomly pasted to the target position (x', y') to obtain the initially generated abnormal image I paste : I paste (x′:x′+w,y′:y′+h)=Patch,x′,y′∈[0,W-H] (6) Poisson Editing Fusion I paste The Poisson fusion operation is performed on the boundary of the pasted area to solve the edge discontinuity problem caused by direct pasting; Poisson Editing is implemented by the following formula: Among them, Δ represents the Laplace operator, Ω is the pasting area, The Poisson Editing solves the Laplace equation and adjusts the pixel value to make the pasted area blend naturally with the background to generate the final abnormal image I anomaly .
4. The unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4 as claimed in claim 1 is characterized in that: In step C, the VIT-L-14-336 model is used as a feature extractor to extract the input image I∈D G Divide the image into blocks and extract the image patch-level features F through a multi-scale local neighborhood aggregation method and a mutual scoring mechanism; in Multi-scale local neighborhood aggregation method for image dataset D G = {I normal ,I anomaly }Extract image patch tokens through feature extractor ViT-L-14-336 Where l∈{1,2…,L} represents different levels of the feature extractor ViT, Q and C are the number of patches and the number of image channels respectively; the layer normalization processing is performed on the features of each scale: Among them, r = {1, 3, 5} is the characteristic aggregation degree, μ i,r , σ i,r are the mean and standard deviation of the feature map of the i-th sample at scale r, and ε is a small constant to avoid the denominator being zero; Then aggregated by adaptive average pooling method The r×r neighborhood features, after aggregation, the feature output is The shape is restored to Q×C, and each polymerization The aggregated feature patch token of each image is Where q∈[1,Q]; the multi-scale local neighborhood aggregation method outputs the aggregated features to a mutual rating mechanism.
5. The unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4 as claimed in claim 4, characterized in that: In step C, the mutual scoring mechanism include: ① Feature similarity calculation: Let the image to be detected I i The patch token is Combine it with D G In addition to I i All images of I j Patch token Similarity score Calculation, the expression is as follows: If I i of with I j Any The more similar, the The lower the score, the more it belongs to the normal area; ② Similarity score enhancement: In order to amplify the score difference between normal and abnormal regions, the interval average method is used to enhance the similarity score. There are K images with the lowest scores of 30%, and the kth image is denoted as I k , then the anomaly score of the defined region for: ③ Multi-scale score fusion: In order to integrate multi-scale feature information, different aggregation degrees are combined The similarity scores at different levels L are averaged and fused, and the overall anomaly score u i,q It is expressed as follows: Definition D G The overall anomaly score set is Convert U to shape Then upsample to the original image size to obtain image I i Feature Results Then D G The result of all image features is Φ=[φ1,φ2…,φ G ] T .
6. The unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4 as claimed in claim 1 is characterized in that: In step D, the unsupervised automobile cylinder head defect detection model includes: Image Encoder: ImageBind-Huge of MiniGPT-4 is used as Image Encoder to extract global and local features of the input image; Image Decoder: Image Decoder adopts a four-stage lightweight, feature matching structure to extract corresponding intermediate patch-level features at each stage; the input image and preset text are encoded by Image Encoder and Text Encoder respectively and then input into Image Decoder, which then outputs the anomaly detection result; Image Encoder and Text Encoder are the same, both are ImageBind-Huge of MiniGPT-4; Image-text matching module: align image features with text features. The intermediate patch-level features extracted in the i-th stage are recorded as i=1,...,4, the feature is not image-text aligned, and the intermediate feature is projected to Align it with the text features representing normal and abnormal semantics, and the abnormal location results Upsample means upsampling, then: Tip learner: Output the image decoder with independently trained basis vectors Integrate into (n1+n2) hint embedding vectors It is fed into LLM together with image features to utilize fine-grained semantic information in the image and keep the image decoder semantically matched with LLM. Geometric information output: It realizes accurate identification of anomalies in the input image and extraction of abnormal geometric information. The steps are as follows: ① Image standardization: Convert the input anomaly location result I into a grayscale image I gray , and then to the standardized I norm , ensuring the consistency of analysis specifications for images of different sizes; ②Region segmentation: I norm The adaptive threshold algorithm is used to process I binary ; ③Geometric information extraction: Through Canny edge detection binary Perform contour recognition and obtain the image edge function δ(E(x,y), edge differential line segment ds; let the abnormal area of the image be R i , then its abnormal area A anomaly , perimeter P anomaly , center coordinate CC anomaly (x,y) calculation formula: Large language model: MiniGPT-4’s Vicuna-7B is used as the LLM. The prompt embedding is input into the LLM, and the detection result R is generated through the following query: R=M LLM (Q,E prompt ) (14)。 7. The unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4 as claimed in claim 1 is characterized in that: In the step E, the image I is input into the image encoder, image decoder, image-text matching module, prompt learner, geometric information learner, and large language model. The model training process uses three loss functions: Cross-Entropy Loss, Focal Loss, and Dice Loss to train Image Decoder and PL: The Cross-Entropy Loss is used to train the language model and is a loss function that measures the difference between the predicted probability distribution and the true distribution in the classification problem: Among them, n is the number of image feature patch tokens, y i The true label of each patch token, p i is the predicted probability for each patch token; The Focal Loss introduces a modulation factor γ based on the Cross-Entropy Loss. By reducing the loss weight of easy-to-classify samples and increasing the loss weight of difficult-to-classify samples, the model pays more attention to difficult-to-classify samples during training: Among them, N P is the total number of image pixels, Predict the value for each pixel, α t is the weight balance factor of the positive sample and negative sample ratio, and the modulation factor γ is set to 2 during training; Dice Loss is mainly used to evaluate the similarity of binary images: Among them, N P is the total number of image pixels, z i , They are the true value and predicted value of each pixel respectively; By combining equations (15) to (17), we can obtain the overall loss function of the model: L total =L ce +L focal +L dice (18) 8. The unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4 as claimed in claim 1 is characterized in that: The method adopts ImageBind-Huge of MiniGPT-4 as Image Encoder and Vicuna-7B as LLM, and aligns the visual language representation of the two through linear layer connection; the image resolution is set to 224×224, and features are extracted from the 8th, 16th, 24th, and 32nd layers of the ImageBind-Huge image encoder and passed to the image decoder; the training rounds are 50, the learning rate is set to 1e-3, and the batch size is 16; in addition, linear warm-up and periodic cosine learning rate decay strategies are adopted to optimize the training process.
9. The unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4 as claimed in claim 1, characterized in that: The method supports users to conduct multiple rounds of interactive queries through LLM, and outputs geometric information of abnormal areas including abnormal area A anomaly , perimeter P anomaly , center coordinate CC anomaly (x,y).
10. The unsupervised automobile cylinder head inner cavity surface defect detection method for MiniGPT-4 as claimed in claim 1, characterized in that: The method is applicable to a variety of industrial defect detection scenarios, including but not limited to parts processing, product assembly and surface defect detection.
Citation Information
Patent Citations
Undecimated wavelet and Gumbel distribution-based fabric defect detection method
CN108399614A
Automobile camshaft mounting hole inner wall defect detection method with strong generalization capability
CN118014931A
System and Method for Finding Dents on an Automobile using a Booth
US20180187409A1
Cited By
Defect detection method and device for inner cavity of valve body and electronic equipment
CN121259461A
Zero-sample multi-mode industrial defect segmentation method based on test sample relation calculation
CN121640051A
GIL equipment defect identification method based on X-ray image enhancement and semantic segmentation
CN122493059A