An industrial anomaly segmentation method, device, equipment and storage medium

By preprocessing and fine-tuning of industrial defect data sets and fine-tuning of multimodal pretrained models, combining image reconstruction and foreground segmentation, the problems of insufficient data sets and noise in the existing technology are solved, and the detection accuracy and model generalization capabilities are improved.

CN119888741BActive Publication Date: 2025-07-25NODING INTELLIGENCE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510387174.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-25
Estimated Expiration
2045-03-31

Smart Images

  • Figure CN119888741B_ABST
    Figure CN119888741B_ABST
Patent Text Reader

Abstract

The present application discloses an industrial anomaly segmentation method, device, equipment and storage medium, which relates to the field of industrial anomaly detection, and includes: obtaining an industrial defect data set, performing data preprocessing on an initial industrial defect image to obtain a target industrial defect data set; fine-tuning an initial multi-modal pre-trained model based on the target industrial defect data set to obtain a fine-tuned industrial multi-modal pre-trained model; using the initial multi-modal pre-trained model to perform image reconstruction and foreground segmentation processing on a current industrial defect image to obtain first and second processed industrial defect images; inputting the first processed industrial defect image into the fine-tuned industrial multi-modal pre-trained model to obtain a target reconstructed industrial defect image, and performing dot multiplication on the target reconstructed industrial defect image and the second processed industrial defect image to obtain a target industrial defect image, so as to complete the industrial anomaly segmentation of the industrial defect image. The accuracy of industrial anomaly detection and the generalization ability of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial anomaly detection, and particularly to an industrial anomaly segmentation method, device, equipment and storage medium. Background Art

[0002] At present, industrial anomaly detection technology is widely used in many current industrial production processes, such as detecting surface defects of mobile phone shell glass, detecting defects of computer panels, detecting surface defects of railway rails, etc. However, for existing industrial anomaly detection, there is currently no widely recognized and unified large-scale dataset for all researchers to use. Existing datasets are usually designed for specific research goals and application environments. Commonly used datasets currently usually have problems such as few samples, class imbalance, and obvious image noise.

[0003] With the rapid development of deep learning technology, in the field of industrial defect detection, researchers have proposed various methods to address problems such as sample imbalance, data scarcity, and a large amount of image noise. However, despite this, further work is still needed to improve the perfection and practical applicability of these technologies, especially in aspects such as precise detection of small targets, processing strategies for imbalanced datasets, and improvement of model generalization ability.

[0004] In summary, how to improve the accuracy of industrial anomaly detection and the generalization ability of the model is an urgent problem to be solved currently. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide an industrial anomaly segmentation method, device, equipment and storage medium, which can improve the accuracy of industrial anomaly detection and the generalization ability of the model. The specific solutions are as follows:

[0006] In the first aspect, the present application provides an industrial anomaly segmentation method, including:

[0007] Obtain an industrial defect dataset, and perform data preprocessing on the initial industrial defect images in the industrial defect dataset to obtain a target industrial defect dataset;

[0008] Fine-tune an initial multi-modal pre-trained model based on the target industrial defect dataset to obtain a fine-tuned industrial multi-modal pre-trained model;

[0009] Use the initial multi-modal pre-trained model to perform image reconstruction and foreground segmentation processing on the current industrial defect images in the target industrial defect dataset to obtain a first processed industrial defect image corresponding to the image reconstruction and a second processed industrial defect image corresponding to the foreground segmentation;

[0010] Input the first processed industrial defect image corresponding to the image reconstruction into the fine-tuned industrial multi-modal pre-trained model to obtain the target reconstructed industrial defect image. Multiply the target reconstructed industrial defect image with the second processed industrial defect image to obtain the target industrial defect image, so as to complete the industrial anomaly segmentation of the industrial defect image.

[0011] Optionally, the data preprocessing of the initial industrial defect images in the industrial defect dataset to obtain the target industrial defect dataset includes:

[0012] Adjust the initial industrial defect image to a preset standard size to obtain the size-adjusted industrial defect image;

[0013] Convert the size-adjusted industrial defect image into a corresponding grayscale image;

[0014] Use a preset image enhancement algorithm to perform preset image processing on the grayscale image to obtain the target industrial defect dataset including the current industrial defect image after image enhancement;

[0015] Wherein, the preset image processing includes enhancing the local contrast of the image and / or eliminating image noise.

[0016] Optionally, the fine-tuning of the initial multi-modal pre-trained model based on the target industrial defect dataset includes:

[0017] Input the current industrial defect image and the corresponding first contaminated phrase into the initial multi-modal pre-trained model to determine the cosine similarity between the current industrial defect image and the corresponding first contaminated phrase. Based on the cosine similarity, obtain the cross-entropy loss values of the current industrial defect image and the corresponding first contaminated phrase respectively, and fine-tune the initial multi-modal pre-trained model based on the average loss value corresponding to the cross-entropy loss value;

[0018] Wherein, the first contaminated phrase is a phrase obtained by contaminating the target keyword phrase corresponding to the current industrial defect image with non-target keyword phrases.

[0019] Optionally, the use of the initial multi-modal pre-trained model to perform image reconstruction and foreground segmentation processing on the current industrial defect images in the target industrial defect dataset includes:

[0020] Use the initial multi-modal pre-trained model to perform image reconstruction and foreground segmentation processing on the current industrial defect images in the target industrial defect dataset respectively based on a preset attention mechanism.

[0021] Optionally, the use of the initial multi-modal pre-trained model to perform image reconstruction and foreground segmentation processing on the current industrial defect image in the target industrial defect dataset includes:

[0022] Using the initial multi-modal pre-trained model and the target keyword group to perform image reconstruction and foreground segmentation processing on the current industrial defect image in the target industrial defect dataset respectively.

[0023] Optionally, the input of the first processed industrial defect image corresponding to the image reconstruction into the fine-tuned industrial multi-modal pre-trained model includes:

[0024] Inputting the first processed industrial defect image corresponding to the image reconstruction and the corresponding second contaminated phrase group into the fine-tuned industrial multi-modal pre-trained model to obtain the target reconstructed industrial defect image;

[0025] Wherein, the second contaminated phrase group is a phrase group obtained by contaminating the target keyword group corresponding to the first processed industrial defect image with a non-target keyword group.

[0026] Optionally, after the input of the first processed industrial defect image corresponding to the image reconstruction into the fine-tuned industrial multi-modal pre-trained model, it further includes:

[0027] Obtaining a fine-tuned industrial multi-modal pre-trained model of a preset target object corresponding to the first processed industrial defect image, so as to perform industrial anomaly segmentation on the industrial defect image to be detected corresponding to the preset target object through the fine-tuned industrial multi-modal pre-trained model of the preset target object.

[0028] In a second aspect, the present application provides an industrial anomaly segmentation device, including:

[0029] A data preprocessing module, configured to obtain an industrial defect dataset, and perform data preprocessing on the initial industrial defect image in the industrial defect dataset to obtain a target industrial defect dataset;

[0030] A model fine-tuning module, configured to fine-tune an initial multi-modal pre-trained model based on the target industrial defect dataset to obtain a fine-tuned industrial multi-modal pre-trained model;

[0031] An image acquisition module, configured to use the initial multi-modal pre-trained model to perform image reconstruction and foreground segmentation processing on the current industrial defect image in the target industrial defect dataset to obtain the first processed industrial defect image corresponding to the image reconstruction and the second processed industrial defect image corresponding to the foreground segmentation;

[0032] The industrial anomaly segmentation completion module is configured to input the first processed industrial defect image corresponding to the image reconstruction into the fine-tuned industrial multi-modal pre-trained model to obtain a target reconstructed industrial defect image, and perform a dot product of the target reconstructed industrial defect image and the second processed industrial defect image to obtain a target industrial defect image, so as to complete the industrial anomaly segmentation of the industrial defect image.

[0033] In a third aspect, the present application provides an electronic device, including:

[0034] A memory for storing a computer program;

[0035] A processor for executing the computer program to implement the industrial anomaly segmentation method as described above.

[0036] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the industrial anomaly segmentation method as described above is implemented.

[0037] In summary, the present application first obtains an industrial defect dataset, preprocesses the initial industrial defect images in the industrial defect dataset to obtain a target industrial defect dataset; fine-tunes an initial multi-modal pre-trained model based on the target industrial defect dataset to obtain a fine-tuned industrial multi-modal pre-trained model; uses the initial multi-modal pre-trained model to perform image reconstruction and foreground segmentation on the current industrial defect images in the target industrial defect dataset to obtain a first post-processed industrial defect image corresponding to the image reconstruction and a second post-processed industrial defect image corresponding to the foreground segmentation; inputs the first post-processed industrial defect image corresponding to the image reconstruction into the fine-tuned industrial multi-modal pre-trained model to obtain a target reconstructed industrial defect image, and multiplies the target reconstructed industrial defect image by the second post-processed industrial defect image to obtain a target industrial defect image, so as to complete the industrial anomaly segmentation of the industrial defect image. As can be seen from the above, the present application obtains an industrial defect dataset and preprocesses the initial industrial defect images therein, fine-tunes the initial multi-modal pre-trained model based on the preprocessed current industrial defect images to obtain a fine-tuned industrial multi-modal pre-trained model, uses the initial multi-modal pre-trained model to perform image reconstruction and foreground segmentation on the current industrial defect images to obtain corresponding first and second post-processed industrial defect images, then inputs the first post-processed industrial defect image into the fine-tuned industrial multi-modal pre-trained model to obtain a target reconstructed industrial defect image, multiplies it by the second post-processed industrial defect image to obtain a target industrial defect image, and completes the industrial anomaly segmentation of the industrial defect image. In this way, the present application can effectively segment the abnormal part of the industrial anomaly image by combining text and image information through multi-modal transfer learning, and at the same time, the fine-tuned industrial multi-modal pre-trained model for the industrial dataset can more specifically perform multi-modal segmentation of industrial anomaly data. By using multi-task and foreground segmentation methods, the influence of background noise on industrial anomaly detection can be better reduced, so that better and more accurate segmentation results can be obtained, and the running memory and time cost can also be reduced through the multi-task method. At the same time, the generalization ability of the model itself can be better improved through multi-modal transfer learning. Description of the Drawings

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0039] Figure 1 It is a flowchart of an industrial anomaly segmentation method disclosed by the present invention;

[0040] Figure 2 Schematic diagram of the fine-tuning process of an initial multi-modal pre-training model disclosed by the present invention;

[0041] Figure 3 Schematic diagram of a specific industrial anomaly segmentation disclosed by the present invention; wherein, Figure 3 (a) in it is the schematic diagram of industrial anomaly segmentation when detecting the surface defects of cashew nuts disclosed by the present invention, Figure 3 (b) in it is the schematic diagram of industrial anomaly segmentation when detecting the defects of circuit boards disclosed by the present invention;

[0042] Figure 4 Schematic diagram of the structure of an industrial anomaly segmentation device disclosed by the present invention;

[0043] Figure 5 Schematic diagram of the structure of an electronic device disclosed by the present invention. Detailed implementation manners

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0045] Currently, industrial anomaly detection technology is widely used in many current industrial production processes, such as detecting the surface defects of the glass of mobile phone casings, detecting the defects of computer panels, detecting the surface defects of railway rails, etc. However, for the existing industrial anomaly detection, there is currently no widely recognized and unified large-scale dataset for all researchers to use. The existing datasets are usually designed for specific research goals and application environments. The commonly used datasets currently usually have problems such as few samples, class imbalance, and obvious image noise. With the rapid development of deep learning technology, in the field of industrial defect detection, researchers have proposed various methods to deal with problems such as sample imbalance, data scarcity, and a lot of image noise. However, despite this, further work is still needed to improve the perfection and practical applicability of these technologies, especially in aspects such as the precise detection of small targets, the processing strategies for unbalanced datasets, and the improvement of the model generalization ability. To solve the above technical problems, the present application discloses an industrial anomaly segmentation method, device, equipment, and storage medium, which can improve the accuracy of industrial anomaly detection and the model generalization ability.

[0046] See Figure 1 As shown, the embodiments of the present invention disclose an industrial anomaly segmentation method, which may include:

[0047] Step S11: Obtain an industrial defect dataset, and perform data preprocessing on the initial industrial defect images in the industrial defect dataset to obtain a target industrial defect dataset.

[0048] In this embodiment, first, an industrial defect dataset is obtained. In a specific implementation manner, a suitable industrial defect dataset can be found from a public dataset platform. These dataset platforms often collect and organize a large number of datasets in different fields and of different types, which can meet certain research and application requirements. In a specific implementation manner, a large number of industrial defect image data accumulated by the enterprise itself during the production process can be relied on. By deploying image acquisition devices on the production line, continuously acquire images of products, including images of normal products and products with various defects, and organize the obtained images into an industrial defect dataset.

[0049] Further, after obtaining the industrial defect dataset, since there may be some differences in the industrial defect images in terms of lighting conditions, image resolution, contrast, and color saturation, it is necessary to perform data preprocessing on the initial industrial defect images. First, adjust the initial industrial defect images to a preset standard size to obtain the size-adjusted industrial defect images; convert the size-adjusted industrial defect images into corresponding grayscale images; use a preset image enhancement algorithm to perform preset image processing on the grayscale images to obtain a target industrial defect dataset containing the current industrial defect images after image enhancement; wherein, the preset image processing includes enhancing the local contrast of the image and / or eliminating image noise. Specifically, the image preprocessing stage includes four processes: image size adjustment, image contrast enhancement, noise reduction, and data enhancement. First, the industrial defect images of different sizes are resized and cropped, and the industrial defect images are adjusted to a standard size of (512×512) pixels. After that, since the resolution of some industrial defect images in the dataset is poor and the abnormal features are difficult to distinguish, in order to improve the quality of the original images, it is necessary to convert the input industrial defect images into grayscale images, and then use an image enhancement algorithm to perform preset image processing on the grayscale images. An image enhancement algorithm using the CLAHE (Contrast Limited Adaptive Histogram Equalization) method can be used to enhance the local contrast and eliminate noise of the input images. Among them, the small areas in the CLAHE images are mainly processed. Then, according to the brightness and intensity of the images, the two parameters of clip limit and block size are adjusted to adjust the quality of the enhanced images. Generally, the higher the value of the CL (clip limit, brightness) parameter, the higher the brightness of the input image. Increasing the value of the BS (block size, contrast) parameter can improve the contrast level, and finally obtain a target industrial defect dataset containing the current industrial defect images after image enhancement.

[0050] Step S12: Fine-tune the initial multi-modal pre-training model based on the target industrial defect dataset to obtain a fine-tuned industrial multi-modal pre-training model.

[0051] In this embodiment, since the traditional initial multi-modal pre-training model lacks knowledge in the field of industrial anomaly detection, after obtaining the target industrial defect dataset, it is necessary to fine-tune the traditional initial multi-modal pre-training model to obtain a fine-tuned industrial multi-modal pre-training model that can process industrial anomaly datasets. First, input the current industrial defect image and the corresponding first contaminated phrase into the initial multi-modal pre-training model to determine the cosine similarity between the current industrial defect image and the corresponding first contaminated phrase, obtain the cross-entropy loss values of the current industrial defect image and the corresponding first contaminated phrase respectively based on the cosine similarity, and fine-tune the initial multi-modal pre-training model based on the average loss value corresponding to the cross-entropy loss value; where the first contaminated phrase is a phrase obtained by contaminating the target keyword phrase corresponding to the current industrial defect image with non-target keyword phrases. Specifically, as Figure 2 shown, first, input the current industrial defect image and the corresponding first contaminated phrase into the traditional initial multi-modal pre-training model together. The first contaminated phrase here is a phrase obtained by contaminating the target keyword phrase originally corresponding to the current industrial defect image with non-target keyword phrases. The so-called non-target keyword phrases refer to a set of words that have little semantic association with the core features of industrial defect images. For example, this is a picture with a [defect] in [material]; this is an industrial image with a [defect] in [material] on it. When the traditional initial multi-modal pre-training model receives the current industrial defect image and the first contaminated phrase, it will determine the cosine similarity between the two according to its own algorithm mechanism. Based on the obtained cosine similarity, further obtain the cross-entropy loss value of the current industrial defect image and the cross-entropy loss value of the corresponding first contaminated phrase respectively. By calculating the cross-entropy loss value, the deviation degree between the predicted similarity and the expected similarity of the model when processing the current industrial defect image and the first contaminated phrase can be evaluated. Finally, take the average value of the cross-entropy loss values of the current industrial defect image and the first contaminated phrase as the overall loss value to obtain the fine-tuned industrial multi-modal pre-training model.

[0052] Step S13: Use the initial multi-modal pre-training model to perform image reconstruction and foreground segmentation processing on the current industrial defect image in the target industrial defect dataset, so as to obtain a first processed industrial defect image corresponding to the image reconstruction and a second processed industrial defect image corresponding to the foreground segmentation.

[0053] In this embodiment, use the initial multi-modal pre-training model to perform image reconstruction and foreground segmentation processing on the current industrial defect image. In the image reconstruction task, the initial multi-modal pre-training model restores or enhances the detailed information of the current industrial defect image according to the patterns and feature representations learned from a large amount of data, such asFigure 3 As shown in part (a) of Figure 3 As shown in part (b) of

[0054] In addition, in this embodiment, the initial multi-modal pre-trained model is used to perform image reconstruction and foreground segmentation processing on the current industrial defect image in the target industrial defect dataset based on a preset attention mechanism. Specifically, while performing image reconstruction and foreground segmentation processing on the current industrial defect image, a preset attention mechanism is used. During the image reconstruction and foreground segmentation processing, the image reconstruction and foreground segmentation tasks share a common feature block. With the help of this shared feature block, the information obtained by the foreground segmentation task can provide strong assistance for the image reconstruction task, thereby optimizing the effect of image reconstruction. Finally, through such a cascaded operation, the first processed industrial defect image (corresponding to the result of the image reconstruction task) and the second processed industrial defect image (corresponding to the result of the foreground segmentation task) are output respectively. This multi-task parallel processing method based on the attention mechanism not only realizes the cooperation and mutual assistance between tasks, but also effectively improves the processing efficiency while significantly reducing the number of parameters required by the model and optimizing the overall performance.

[0055] Meanwhile, the initial multi-modal pre-trained model and the target keyword group are used to perform image reconstruction and foreground segmentation processing on the current industrial defect image in the target industrial defect dataset respectively. Specifically, with the help of the initial multi-modal pre-trained model combined with the target keyword group, image reconstruction and foreground segmentation processing are carried out on the current industrial defect image to achieve accurate analysis and positioning of the defect. For example, for an industrial defect image, the corresponding target keyword group contains words such as [material], and the initial multi-modal pre-trained model will focus on relevant features in the industrial defect image according to these semantic clues.

[0056] Step S14: Input the first processed industrial defect image corresponding to the image reconstruction into the fine-tuned industrial multi-modal pre-trained model to obtain the target reconstructed industrial defect image, and perform a dot product of the target reconstructed industrial defect image and the second processed industrial defect image to obtain the target industrial defect image, so as to complete the industrial anomaly segmentation of the industrial defect image.

[0057] In this embodiment, after obtaining the first processed industrial defect image corresponding to the image reconstruction and the second processed industrial defect image corresponding to the foreground segmentation, the first processed industrial defect image corresponding to the image reconstruction and the corresponding second contaminated phrase are input into the fine-tuned industrial multi-modal pre-trained model to obtain the target reconstructed industrial defect image; wherein, the second contaminated phrase is a phrase obtained by contaminating the target keyword phrase corresponding to the first processed industrial defect image with a non-target keyword phrase. Specifically, if the first processed industrial defect image and the second processed industrial defect image are successfully obtained, the first processed industrial defect image generated by the image reconstruction is used as an important input part of the visual information. At the same time, the corresponding second contaminated phrase is introduced. Subsequently, the first processed industrial defect image and the second contaminated phrase are jointly input into the fine-tuned industrial multi-modal pre-trained model. Based on its powerful multi-modal fusion and analysis architecture, feature extraction will be performed on the input first processed industrial defect image and the second contaminated phrase. Then the target reconstructed industrial defect image is output. Finally, a dot product of the target reconstructed industrial defect image and the second processed industrial defect image is performed to obtain the final required target industrial defect image.

[0058] In addition, while obtaining the target reconstructed post-industrial defect image, a fine-tuned industrial multi-modal pre-training model of a preset target object corresponding to the first processed industrial defect image is obtained, so as to perform industrial anomaly detection on the industrial defect image to be detected corresponding to the preset target object through the fine-tuned industrial multi-modal pre-training model of the preset target object. Specifically, when obtaining the target reconstructed post-industrial defect image, a fine-tuned industrial multi-modal pre-training model of a preset target object corresponding to the first processed industrial defect image can also be obtained. The fine-tuned industrial multi-modal pre-training model is a model for industrial anomaly detection over a large range of industrial anomalies, and the initial multi-modal pre-training model is fine-tuned for the preset target object to obtain a model for industrial anomaly detection related to the preset target object. The preset target object can be a specific product or component in industrial production, such as a mechanical part, etc. The initial multi-modal pre-training model is fine-tuned with a large amount of relevant industrial defect image data, focusing on learning the unique feature patterns of the preset target object and the association of modal information. Once a deviation in the features in the image from the normal pattern is found, it can accurately identify possible industrial anomaly situations, such as dimensional deviations of components, tiny cracks on the surface, etc. Compared with the fine-tuned industrial multi-modal pre-training model for general industrial anomalies, the fine-tuned industrial multi-modal pre-training model of the preset target object can significantly improve the accuracy and efficiency of detection, timely discover potential industrial defects, provide strong support for the quality control and optimization of industrial production, and effectively avoid various losses and risks caused by product defects.

[0059] As can be seen from the above, the embodiment of the present application obtains an industrial defect data set and performs data preprocessing on the initial industrial defect images therein. Based on the current industrial defect images after preprocessing, the initial multi-modal pre-training model is fine-tuned to obtain a fine-tuned industrial multi-modal pre-training model. The initial multi-modal pre-training model is used to perform image reconstruction and foreground segmentation processing on the current industrial defect images to obtain the corresponding first and second processed industrial defect images. Then, the first processed industrial defect image is input into the fine-tuned industrial multi-modal pre-training model to obtain the target reconstructed post-industrial defect image, and it is dot-multiplied with the second processed industrial defect image to obtain the target industrial defect image, completing the industrial anomaly segmentation of the industrial defect image. In this way, through the use of multi-modal transfer learning, the present application can effectively combine text and image information to effectively segment the abnormal part of the industrial anomaly image. At the same time, adding a fine-tuned industrial multi-modal pre-training model for the industrial data set can be more targeted for multi-modal segmentation of industrial anomaly data. By using multi-task and foreground segmentation methods, the influence of background noise on anomaly detection can be better reduced, so as to obtain better and more accurate segmentation results. And through the multi-task method, the running memory and time cost can also be reduced. At the same time, through multi-modal transfer learning, the generalization ability of the model itself can also be better improved.

[0060] Based on the previous embodiment, the present application discloses an industrial anomaly segmentation method, which can improve the accuracy of industrial anomaly detection and the generalization ability of the model. Next, a detailed description of the industrial anomaly segmentation method will be given.

[0061] The present application first obtains an industrial defect dataset. After obtaining the industrial defect dataset, data preprocessing operations such as image size adjustment, image contrast enhancement, noise reduction, and data augmentation are performed on the initial industrial defect images in the industrial defect dataset to obtain a target industrial defect dataset including the current industrial defect images after image enhancement. Then, the target industrial defect dataset is input into a traditional initial multi-modal pre-trained model, and the traditional initial multi-modal pre-trained model is fine-tuned to obtain a fine-tuned industrial multi-modal pre-trained model related to industrial anomaly detection. Next, an image reconstruction task is performed on the current industrial defect image in the initial multi-modal pre-trained model to obtain a corresponding first processed industrial defect image, and at the same time, a foreground segmentation task is performed on the current industrial defect image in the initial multi-modal pre-trained model to obtain a corresponding second processed industrial defect image. After obtaining the first processed industrial defect image corresponding to the image reconstruction and the second processed industrial defect image corresponding to the foreground segmentation, the first processed industrial defect image corresponding to the image reconstruction and the corresponding second contaminated phrase group are input into the fine-tuned industrial multi-modal pre-trained model to obtain the target reconstructed industrial defect image. Finally, the target reconstructed industrial defect image and the second processed industrial defect image are multiplied point by point to obtain the target industrial defect image, so as to complete the industrial anomaly segmentation of the industrial defect image. In this way, it can play an important role in the quality control and defect detection links of industrial production and is beneficial to improving the overall quality and stability of products.

[0062] See Figure 4 As shown, an embodiment of the present invention discloses an industrial anomaly segmentation device, which may include:

[0063] A data preprocessing module 11, configured to obtain an industrial defect dataset and perform data preprocessing on the initial industrial defect images in the industrial defect dataset to obtain a target industrial defect dataset;

[0064] A model fine-tuning module 12, configured to fine-tune an initial multi-modal pre-trained model based on the target industrial defect dataset to obtain a fine-tuned industrial multi-modal pre-trained model;

[0065] An image acquisition module 13, configured to perform image reconstruction and foreground segmentation processing on the current industrial defect images in the target industrial defect dataset by using the initial multi-modal pre-trained model to obtain the first processed industrial defect image corresponding to the image reconstruction and the second processed industrial defect image corresponding to the foreground segmentation;

[0066] The industrial anomaly segmentation completion module 14 is configured to input the first processed industrial defect image corresponding to the image reconstruction into the fine-tuned industrial multi-modal pre-trained model to obtain a target reconstructed industrial defect image, and perform a dot product of the target reconstructed industrial defect image and the second processed industrial defect image to obtain a target industrial defect image, so as to complete the industrial anomaly segmentation of the industrial defect image.

[0067] As can be seen from the above, in this application, an industrial defect data set is obtained and data preprocessing is performed on the initial industrial defect images therein. Based on the current industrial defect images after preprocessing, an initial multi-modal pre-trained model is fine-tuned to obtain a fine-tuned industrial multi-modal pre-trained model. The initial multi-modal pre-trained model is used to perform image reconstruction and foreground segmentation processing on the current industrial defect images to obtain corresponding first and second processed industrial defect images. Then, the first processed industrial defect image is input into the fine-tuned industrial multi-modal pre-trained model to obtain a target reconstructed industrial defect image, and a dot product is performed on it and the second processed industrial defect image to obtain a target industrial defect image, thereby completing the industrial anomaly segmentation of the industrial defect image. In this way, by using multi-modal transfer learning, this application can effectively combine text and image information to effectively segment the abnormal parts of industrial abnormal images. At the same time, the fine-tuned industrial multi-modal pre-trained model for industrial data sets can perform more targeted multi-modal segmentation of industrial abnormal data. By using multi-task and foreground segmentation methods, the influence of background noise on anomaly detection can be better reduced, so as to obtain better and more accurate segmentation results. Through multi-modal transfer learning, the generalization ability of the model itself can also be better improved.

[0068] In a specific implementation manner, the data preprocessing module 11 may specifically include:

[0069] An industrial defect image acquisition unit after size adjustment is configured to adjust the initial industrial defect image to a preset standard size to obtain an industrial defect image after size adjustment;

[0070] An industrial defect image conversion unit after size adjustment is configured to convert the industrial defect image after size adjustment into a corresponding grayscale image;

[0071] A current industrial defect image acquisition unit is configured to perform preset image processing on the grayscale image by using a preset image enhancement algorithm to obtain a target industrial defect data set including the current industrial defect image after image enhancement; wherein, the preset image processing includes enhancing the local contrast of the image and / or eliminating image noise.

[0072] In a specific implementation manner, the model fine-tuning module 12 may specifically include:

[0073] An initial multi-modal pre-training model fine-tuning unit is used to input the current industrial defect image and the corresponding first contaminated phrase into the initial multi-modal pre-training model to determine the cosine similarity between the current industrial defect image and the corresponding first contaminated phrase, obtain the cross-entropy loss values of the current industrial defect image and the corresponding first contaminated phrase respectively based on the cosine similarity, and fine-tune the initial multi-modal pre-training model based on the average loss value corresponding to the cross-entropy loss value; wherein, the first contaminated phrase is a phrase obtained by contaminating the target keyword phrase corresponding to the current industrial defect image with non-target keyword phrases.

[0074] In a specific implementation manner, the image acquisition module 13 may specifically include:

[0075] A first current industrial defect image processing unit is used to perform image reconstruction and foreground segmentation processing on the current industrial defect images in the target industrial defect dataset respectively based on the initial multi-modal pre-training model and a preset attention mechanism.

[0076] In a specific implementation manner, the image acquisition module 13 may specifically include:

[0077] A second current industrial defect image processing unit is used to perform image reconstruction and foreground segmentation processing on the current industrial defect images in the target industrial defect dataset respectively by using the initial multi-modal pre-training model and the target keyword phrase.

[0078] In a specific implementation manner, the industrial anomaly segmentation completion module 14 may specifically include:

[0079] A target reconstructed industrial defect image acquisition unit is used to input the first processed industrial defect image corresponding to the image reconstruction and the corresponding second contaminated phrase into the fine-tuned industrial multi-modal pre-training model to obtain the target reconstructed industrial defect image; wherein, the second contaminated phrase is a phrase obtained by contaminating the target keyword phrase corresponding to the first processed industrial defect image with non-target keyword phrases.

[0080] In a specific implementation manner, the industrial anomaly segmentation device may further include:

[0081] A fine-tuned industrial multi-modal pre-training model acquisition module for a preset target object is used to obtain a fine-tuned industrial multi-modal pre-training model for the preset target object corresponding to the first processed industrial defect image, so as to perform industrial anomaly segmentation on the to-be-detected industrial defect image corresponding to the preset target object through the fine-tuned industrial multi-modal pre-training model for the preset target object.

[0082] Furthermore, an embodiment of the present application also discloses an electronic device. Figure 5 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be considered as any limitation on the scope of use of the present application.

[0083] Figure 5 It is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the industrial anomaly segmentation method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0084] In this embodiment, the power supply 23 is used to provide operating voltages for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitation is made here.

[0085] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be short-term storage or permanent storage.

[0086] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it may be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the industrial anomaly segmentation method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.

[0087] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the industrial anomaly segmentation method disclosed above is implemented. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details are not repeated here.

[0088] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0089] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0090] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0091] Finally, it should also be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0092] The technical solutions provided in this application have been introduced in detail above. Specific examples are used herein to elaborate on the principles and implementation manners of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. An industrial anomaly segmentation method, characterized in that, Including: Obtain an industrial defect dataset, and perform data preprocessing on the initial industrial defect images in the industrial defect dataset to obtain a target industrial defect dataset; Fine-tune an initial multi-modal pre-trained model based on the target industrial defect dataset to obtain a fine-tuned industrial multi-modal pre-trained model; Use the initial multi-modal pre-trained model to perform image reconstruction and foreground segmentation on the current industrial defect images in the target industrial defect dataset to obtain a first processed industrial defect image corresponding to the image reconstruction and a second processed industrial defect image corresponding to the foreground segmentation; Input the first processed industrial defect image corresponding to the image reconstruction into the fine-tuned industrial multi-modal pre-trained model to obtain a target reconstructed industrial defect image, and perform dot multiplication on the target reconstructed industrial defect image and the second processed industrial defect image to obtain a target industrial defect image, so as to complete the industrial anomaly segmentation of the industrial defect image; Among them, the fine-tuning of the initial multi-modal pre-trained model based on the target industrial defect dataset includes: Input the current industrial defect image and the corresponding first contaminated phrase into the initial multi-modal pre-trained model to determine the cosine similarity between the current industrial defect image and the corresponding first contaminated phrase, obtain the cross-entropy loss values of the current industrial defect image and the corresponding first contaminated phrase respectively based on the cosine similarity, and fine-tune the initial multi-modal pre-trained model based on the average loss value corresponding to the cross-entropy loss value; Among them, the first contaminated phrase is a phrase obtained by contaminating the target keyword phrase corresponding to the current industrial defect image with a non-target keyword phrase; Among them, the input of the first processed industrial defect image corresponding to the image reconstruction into the fine-tuned industrial multi-modal pre-trained model includes: Input the first processed industrial defect image corresponding to the image reconstruction and the corresponding second contaminated phrase into the fine-tuned industrial multi-modal pre-trained model to obtain the target reconstructed industrial defect image; Among them, the second contaminated phrase is a phrase obtained by contaminating the target keyword phrase corresponding to the first processed industrial defect image with a non-target keyword phrase.

2. The industrial anomaly segmentation method according to claim 1, wherein The data preprocessing of the initial industrial defect images in the industrial defect dataset to obtain a target industrial defect dataset includes: Adjust the initial industrial defect image to a preset standard size to obtain a size-adjusted industrial defect image; Convert the size-adjusted industrial defect image into a corresponding grayscale image; Use a preset image enhancement algorithm to perform preset image processing on the grayscale image to obtain a target industrial defect dataset including the current industrial defect image after image enhancement; Among them, the preset image processing includes enhancing the local contrast of the image and / or eliminating image noise.

3. The industrial anomaly segmentation method according to claim 1, characterized in that The use of the initial multi-modal pre-trained model to perform image reconstruction and foreground segmentation on the current industrial defect images in the target industrial defect dataset includes: Using the initial multi-modal pre-training model, perform image reconstruction and foreground segmentation processing on the current industrial defect images in the target industrial defect dataset respectively based on a preset attention mechanism.

4. The industrial anomaly segmentation method according to claim 1, wherein, The performing image reconstruction and foreground segmentation processing on the current industrial defect images in the target industrial defect dataset by using the initial multi-modal pre-training model includes: Using the initial multi-modal pre-training model and a target keyword group to perform image reconstruction and foreground segmentation processing on the current industrial defect images in the target industrial defect dataset respectively.

5. The industrial anomaly segmentation method according to claim 1, characterized in that After inputting the first processed industrial defect image corresponding to the image reconstruction into the fine-tuned industrial multi-modal pre-training model, it further includes: Obtaining a fine-tuned industrial multi-modal pre-training model of a preset target object corresponding to the first processed industrial defect image, so as to perform industrial anomaly segmentation on the industrial defect image to be detected corresponding to the preset target object through the fine-tuned industrial multi-modal pre-training model of the preset target object.

6. An industrial anomaly segmentation device, characterized in that, It includes: A data preprocessing module, configured to obtain an industrial defect dataset, and perform data preprocessing on the initial industrial defect images in the industrial defect dataset to obtain a target industrial defect dataset; A model fine-tuning module, configured to fine-tune an initial multi-modal pre-training model based on the target industrial defect dataset to obtain a fine-tuned industrial multi-modal pre-training model; An image acquisition module, configured to use the initial multi-modal pre-training model to perform image reconstruction and foreground segmentation processing on the current industrial defect images in the target industrial defect dataset to obtain a first processed industrial defect image corresponding to the image reconstruction and a second processed industrial defect image corresponding to the foreground segmentation; An industrial anomaly segmentation completion module, configured to input the first processed industrial defect image corresponding to the image reconstruction into the fine-tuned industrial multi-modal pre-training model to obtain a target reconstructed industrial defect image, and perform dot multiplication on the target reconstructed industrial defect image and the second processed industrial defect image to obtain a target industrial defect image, so as to complete the industrial anomaly segmentation of the industrial defect image; Among them, the model fine-tuning module includes: Inputting the current industrial defect image and a corresponding first contaminated phrase into the initial multi-modal pre-training model to determine the cosine similarity between the current industrial defect image and the corresponding first contaminated phrase, obtaining the cross-entropy loss values of the current industrial defect image and the corresponding first contaminated phrase respectively based on the cosine similarity, and fine-tuning the initial multi-modal pre-training model based on the average loss value corresponding to the cross-entropy loss value; Among them, the first contaminated phrase is a phrase obtained by contaminating the target keyword group corresponding to the current industrial defect image with a non-target keyword group; Among them, the industrial anomaly segmentation completion module includes: Inputting the first processed industrial defect image corresponding to the image reconstruction and a corresponding second contaminated phrase into the fine-tuned industrial multi-modal pre-training model to obtain the target reconstructed industrial defect image; Among them, the second pollution phrase is a phrase obtained by polluting the target keyword phrase corresponding to the first processed industrial defect image with a non-target keyword phrase.

7. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for executing the computer program to implement the industrial anomaly segmentation method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by a processor, the industrial anomaly segmentation method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Metal processing multi-mode defect on-line detection method and metal processing multi-mode defect on-line detection system

    CN118297926A