A defect detection method and system based on language prompts and collaborative teaching

By introducing technical means based on language prompts and collaborative teaching in the defect detection method, the problems of insufficient defect detection accuracy and insufficient sample data in the existing technology are solved, and more efficient and universal defect detection effects are achieved.

CN119338834BActive Publication Date: 2025-06-13JIANGXI NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411907721.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-06-13
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

In the prior art, industrial product defect detection relies on manual quality inspection, which is costly and limited in efficiency. The image detection method of computer vision technology requires a large amount of labeled sample data, resulting in insufficient detection accuracy. Especially in industrial production, there are various types of defects and it is difficult to obtain sample data.

Method used

Defect detection methods based on language prompting and collaborative teaching are adopted to expand sample data through data enhancement, learnable language prompt text is introduced, and more accurate local detailed information is obtained using the local feature attention enhancement mechanism, and two independent defect detection modules are prompted and updated to improve the accuracy and generality of detection.

Benefits of technology

It improves the accuracy and general applicability of defect detection methods, can more effectively adapt to the complex and changeable detection situation in industrial production, and reduces the dependence on large amounts of labeled sample data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119338834B_ABST
    Figure CN119338834B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing, and proposes a defect detection method and system based on language prompts and collaborative teaching. By data augmentation, the normal sample image data and defect sample image data are expanded, avoiding the problem of insufficient sample data in actual production. Then, by introducing learnable language prompt texts, the detection ability for different defects of different types of products is enhanced. Through the local feature attention enhancement mechanism, more accurate local detail information is obtained to enhance the overall detection accuracy. Also, through two independent defect detection result predictions, a more comprehensive defect detection result is obtained. At the same time, the two defect detection modules prompt and update each other, so that the overall detection is more in line with the actual application situation, improving the versatility of the detection method. The present invention improves the accuracy and versatility of the defect detection method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a defect detection method and system based on language prompts and collaborative teaching. Background Art

[0002] With the rapid development of industrial production, the requirements for industrial manufacturing have gradually increased. In particular, there are widespread problems of defects in industrial products in industrial production, and industrial product defects may cause various safety accidents. Therefore, detecting whether there are defects in industrial products in industrial production is an important part of industrial production.

[0003] In the prior art, it is often necessary for professional quality inspectors to perform defect detection and subjective judgment on the surface of industrial products manually. This not only has a high cost, but also the detection efficiency is relatively limited. Therefore, computer vision technology has become a key detection means, and the surface defects of industrial products are detected by the image anomaly detection method of deep learning. However, the existing image detection methods often require a large amount of labeled sample data. However, due to the large variety of product defects in industrial production, it is difficult to obtain a huge amount of labeled sample data, resulting in insufficient detection accuracy. At the same time, the detection situation in actual production applications is complex and changeable, resulting in the prediction results not conforming to the actual production situation.

[0004] Therefore, how to design a defect detection method to adapt to the changes in actual production applications and improve the accuracy of defect detection has become an urgent problem to be solved. Summary of the Invention

[0005] Based on this, a defect detection method and system based on language prompts and collaborative teaching provided by the present invention expand the normal sample image data and defect sample image data through data augmentation, avoiding the problem of insufficient sample data in actual production. Then, by introducing learnable language prompt texts, the detection ability for different defects of different types of products is enhanced. Through the local feature attention enhancement mechanism, more accurate local detail information is obtained to enhance the overall detection accuracy. Also, through two independent defect detection result predictions, a more comprehensive defect detection result is obtained. At the same time, the two defect detection modules prompt and update each other, making the overall detection more in line with the actual application situation and improving the versatility of the detection method. The present invention improves the accuracy and versatility of the defect detection method.

[0006] A defect detection method based on language prompts and collaborative teaching proposed by the present invention includes:

[0007] Real-time collecting the image data of the target to be measured and performing data augmentation processing, where the data augmentation processing is used to expand and obtain the normal sample image data and defect sample image data of the target to be measured;

[0008] Obtain a language prompt text according to the category of the target image data to be measured, and then obtain an enhanced feature of the language prompt text according to the language prompt. The language prompt text is in a learnable state and is used to guide feature localization and feature extraction;

[0009] Perform attention enhancement processing on the target image data to be measured according to the local feature attention enhancement algorithm to obtain local feature attention enhancement features. The attention enhancement processing includes multi-modal feature alignment according to the enhanced features of the language prompt text;

[0010] Obtain a language prompt defect detection result and a basic image defect detection result respectively according to the local feature attention enhancement features, so as to calculate a final anomaly score and perform defect detection update.

[0011] In summary, according to the above defect detection method based on language prompts and collaborative teaching, through data augmentation, the normal sample image data and defect sample image data are expanded, avoiding the problem of insufficient sample data in actual production. Then, by introducing learnable language prompt texts, the detection ability for different defects of different types of products is enhanced. Through the local feature attention enhancement mechanism, more and more accurate local detail information is obtained to enhance the overall detection accuracy. Also, through two independent defect detection result predictions, a more comprehensive defect detection result is obtained. At the same time, the two defect detection modules prompt and update each other, making the overall detection more in line with the actual application situation and improving the versatility of the detection method. The present invention improves the accuracy and versatility of the defect detection method. Specifically, the target image data to be measured is collected in real time and data augmentation processing is performed. The data augmentation processing is used to expand and obtain the normal sample image data and defect sample image data of the target to be measured, avoiding the problem of deviation in detection accuracy caused by insufficient sample data in actual use. A language prompt text is obtained according to the category of the target image data to be measured, and then an enhanced feature of the language prompt text is obtained according to the language prompt. The language prompt text is in a learnable state and is used to guide feature localization and feature extraction. The target image data to be measured is subjected to attention enhancement processing according to the local feature attention enhancement algorithm to obtain local feature attention enhancement features. The attention enhancement processing includes multi-modal feature alignment according to the enhanced features of the language prompt text, further enhancing the attention to local detail information and avoiding the influence of the complex detection situation of various products and defect types in the actual production process, further improving the overall detection accuracy. A language prompt defect detection result and a basic image defect detection result are obtained respectively according to the local feature attention enhancement features, so as to calculate a final anomaly score and perform defect detection update, further improving the versatility of the detection method and making it more in line with the actual application scenario. The present invention improves the accuracy and versatility of the defect detection method.

[0012] Further, the step of collecting the image data of the target to be measured in real time and performing data enhancement processing specifically includes:

[0013] After the data enhancement module collects the image data of the target to be measured in real time, it performs data enhancement;

[0014] According to the random affine transformation and sharpening, multiple normal sample image data are obtained, and then defects are simulated according to the Perlin noise and fitted into the normal sample image data to obtain multiple defective sample image data;

[0015] The specific algorithm of the data enhancement is as follows:

[0016] ,

[0017] Among them, represents the defective sample image data, represents the binary image obtained by the random Perlin noise through a random threshold, represents the inverse matrix of, represents the input image, represents the random texture image, represents the random transparency parameter, represents the element-wise multiplication.

[0018] Further, the step of obtaining the language prompt text according to the category of the image data of the target to be measured and then obtaining the enhanced features of the language prompt text according to the language prompt specifically includes:

[0019] The language prompt learning module obtains the language prompt text according to the category of the image data of the target to be measured. The language prompt text includes the normal language prompt text and the abnormal language prompt text. The specific algorithm for obtaining the language prompt text is as follows:

[0020] ,

[0021] ,

[0022] Among them, represents the normal language prompt text, represents the abnormal language prompt text, represents the learnable word embedding of the normal language prompt text, represents the learnable word embedding of the abnormal language prompt text, represents the category of the image data of the target to be measured, represents the defect guidance, represents the defect name;

[0023] Feed forward based on the above language hint text and generate corresponding enhanced features of the language hint text. The specific algorithm for the feed forward step is as follows:

[0024]

[0025] Among them, represents the input hint text of the hint layer, represents the hint layer, represents the hint layer ordinal number, represents the maximum learnable hint layer number, represents the learnable hint text.

[0026] Furthermore, the step of performing attention enhancement processing on the to-be-tested target image data according to the local feature attention enhancement algorithm to obtain local feature attention enhancement features specifically includes:

[0027] The image encoding module performs attention enhancement processing on the to-be-tested target image data according to the local feature attention enhancement algorithm. The image encoding module includes multiple local feature attention enhancement mechanism layers;

[0028] Extract local features. The specific algorithm for extracting local features is as follows:

[0029]

[0030] Among them, represents the original output feature, represents the local feature attention enhancement feature, represents the image category feature, represents the last local feature, respectively represent the query matrix, key matrix, and value matrix, represents attention projection, represents the local feature attention projection, represents the attention function.

[0031] Furthermore, the step of respectively obtaining the language hint defect detection result and the basic image defect detection result according to the local feature attention enhancement feature specifically includes:

[0032] The image encoding module selects different local feature attention enhancement mechanism layers according to the preset sequence ordinal number to obtain multiple intermediate image block-level features, and aligns the intermediate image block-level features with the language hint text to obtain multi-modal alignment features;

[0033] Obtain the language prompt defect detection result according to the language prompt text enhancement feature and the multimodal alignment feature. The specific algorithm for obtaining the language prompt defect detection result is as follows:

[0034] ,

[0035] wherein, represents the language prompt defect detection result, represents the upsampling operation, represents the cosine similarity calculation operation, represents the multimodal alignment feature, and respectively represent the normal language prompt text enhancement feature and the abnormal language prompt text enhancement feature, represents the local feature attention enhancement mechanism layer ordinal number, represents the total number of selected local feature attention enhancement mechanism layers, represents the exponential function;

[0036] The defect detection module includes a first defect detection sub-module and a second defect detection sub-module. The first defect detection sub-module and the second defect detection sub-module respectively obtain the basic image defect detection result according to the multimodal alignment feature;

[0037] The specific algorithm for obtaining the basic image defect detection result is as follows:

[0038] ,

[0039] ,

[0040] wherein, and respectively represent the basic image defect detection results obtained by the first defect detection sub-module and the second defect detection sub-module, and respectively represent the first defect detection sub-module and the second defect detection sub-module.

[0041] Furthermore, the step of calculating the final abnormal score specifically includes:

[0042] Calculate the final abnormal score according to the language prompt defect detection result and the basic image defect detection result. The specific algorithm for calculating the final abnormal score is as follows:

[0043] ,

[0044] wherein, represents the final abnormal score, and respectively represent the basic image defect detection results obtained by the first defect detection sub-module and the second defect detection sub-module, represent the language prompt defect detection result;

[0045] Then perform loss calculation, and the specific algorithm of the loss calculation is as follows:

[0046] ,

[0047] ,

[0048] ,

[0049] Among them, represents the focal loss, represents the dice loss, represents the total loss, represents the total number of pixels, represents the pixel ordinal number, represents the predicted probability, represents the adjustable parameter of the weight, represents the output of the decoder, represents the true label, and represents the loss balance coefficient.

[0050] Furthermore, the steps of performing defect detection update specifically include:

[0051] Perform threshold filtering according to the language prompt defect detection result to obtain the normal area of the image, and update the first defect detection sub-module and the second defect detection sub-module of the defect detection module according to the normal area of the image;

[0052] The specific algorithm of the update is as follows:

[0053] ,

[0054] ,

[0055] ,

[0056] Among them, represents the pixel area coordinate, represents the pixel area marker, represents the language prompt defect detection result, represents the threshold learned from the augmented defect sample image data, and respectively represent the first defect detection sub-module and the second defect detection sub-module, represents the learning rate, Indicates the calculated gradient, Indicates the total loss;

[0057] Conduct collaborative teaching on the first defect detection sub-module and the second defect detection sub-module of the defect detection module;

[0058] Calculate the loss based on the high-confidence image regions in the basic image defect detection results obtained by the first defect detection sub-module to update the second defect detection sub-module, and calculate the loss based on the high-confidence image regions in the basic image defect detection results obtained by the second defect detection sub-module to update the first defect detection sub-module. The high-confidence image regions used for updating in the first defect detection sub-module and the second defect detection sub-module are not updated;

[0059] The specific algorithm for the collaborative teaching is as follows:

[0060] ,

[0061]

[0062]

[0063]

[0064] Among them, and respectively represent the confidence levels of the basic image defect detection results of the first defect detection sub-module and the second defect detection sub-module, represents the defect classification threshold, and respectively represent the basic image defect detection results obtained by the first defect detection sub-module and the second defect detection sub-module, and represent the first defect detection sub-module and the second defect detection sub-module after collaborative teaching update.

[0065] A defect detection system based on language prompts and collaborative teaching proposed by the present invention includes:

[0066] A data augmentation module for real-time collecting data of the target image to be measured and performing data augmentation processing, where the data augmentation processing is used to expand and obtain normal sample image data and defect sample image data of the target image to be measured;

[0067] A language prompt learning module for obtaining language prompt texts according to the data category of the target image to be measured, and then obtaining language prompt text enhanced features according to the language prompts. The language prompt texts are in a learnable state, and the language prompt texts are used to guide feature localization and feature extraction;

[0068] An image encoding module, configured to perform attention enhancement processing on the to-be-detected target image data according to the local feature attention enhancement algorithm to obtain local feature attention enhancement features, where the attention enhancement processing includes performing multi-modal feature alignment according to the language prompt text enhancement features;

[0069] A defect detection module, configured to respectively obtain a language prompt defect detection result and a basic image defect detection result according to the local feature attention enhancement features, so as to calculate a final anomaly score and perform defect detection update.

[0070] The present invention also provides a storage medium, where the storage medium stores one or more programs, and when the programs are executed by a processor, the defect detection method based on language prompts and collaborative teaching as described above is implemented.

[0071] The present invention also provides a computer device, where the computer device includes a memory and a processor, and:

[0072] The memory is used to store a computer program;

[0073] When the processor is used to execute the computer program stored in the memory, the defect detection method based on language prompts and collaborative teaching as described above is implemented. Description of the Drawings

[0074] Figure 1 It is a flowchart of the defect detection method based on language prompts and collaborative teaching proposed in the first embodiment of the present invention;

[0075] Figure 2 It is a flowchart of the defect detection method based on language prompts and collaborative teaching proposed in the second embodiment of the present invention;

[0076] Figure 3 It is a schematic structural diagram of the defect detection system based on language prompts and collaborative teaching proposed in the third embodiment of the present invention.

[0077] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. Specific Embodiments

[0078] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0079] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.

[0080] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0081] Please refer to Figure 1 , which shows a flowchart of a defect detection method based on language prompts and collaborative teaching proposed in the first embodiment of the present invention. This defect detection method based on language prompts and collaborative teaching includes steps S01 to S04, where:

[0082] Step S01: Real-time collect the image data of the target to be measured and perform data enhancement processing;

[0083] It should be noted that in this embodiment, the data enhancement processing is used to expand and obtain the normal sample image data and defect sample image data of the target to be measured. After the data enhancement module real-time collects the image data of the target to be measured, data enhancement is performed;

[0084] According to random affine transformation and sharpening, multiple normal sample image data are obtained, and then defects are simulated according to Perlin noise and fitted into the normal sample image data to obtain multiple defect sample image data;

[0085] The specific algorithm of the data enhancement is as follows:

[0086] ,

[0087] where, represents the defect sample image data, represents the binary image obtained by the random threshold of the randomly generated Perlin noise, represents the inverse matrix of, represents the input image, represents the random texture image, represents the random transparency parameter, represents element-wise multiplication.

[0088] Step S02: Obtain a language prompt text according to the data category of the target image to be measured, and then obtain the enhanced features of the language prompt text according to the language prompt;

[0089] It should be noted that in this embodiment, the language prompt text is in a learnable state. The language prompt text is used to guide feature localization and feature extraction. The language prompt learning module obtains the language prompt text according to the data category of the target image to be measured. The language prompt text includes a normal language prompt text and an abnormal language prompt text. The specific algorithm for obtaining the language prompt text is as follows:

[0090] ,

[0091] ,

[0092] Among them, represents the normal language prompt text, represents the abnormal language prompt text, represents the learnable word embedding of the normal language prompt text, represents the learnable word embedding of the abnormal language prompt text, represents the data category of the target image to be measured, represents defect guidance, represents the defect name;

[0093] Feed forward according to the language prompt text and generate the corresponding enhanced features of the language prompt text. The specific algorithm for the feed forward step is as follows:

[0094]

[0095] Among them, represents the input prompt text of the prompt layer, represents the prompt layer, represents the prompt layer ordinal number, represents the maximum learnable prompt layer number, represents the learnable prompt text.

[0096] Step S03: Perform attention enhancement processing on the target image data to be measured according to the local feature attention enhancement algorithm to obtain the local feature attention enhancement features;

[0097] It should be noted that in this embodiment, the attention enhancement processing includes performing multi-modal feature alignment according to the enhanced features of the language prompt text. The image encoding module performs attention enhancement processing on the target image data to be measured according to the local feature attention enhancement algorithm. The image encoding module includes multiple local feature attention enhancement mechanism layers;

[0098] Extract local features. The specific algorithm for extracting local features is as follows:

[0099]

[0100] Among them, represents the original output feature, represents the local feature attention enhancement feature, represents the image category feature, represents the last local feature, represent the query matrix, key matrix, and value matrix respectively, represents attention projection, represents the local feature attention projection, represents the attention function.

[0101] Step S04: Obtain the language prompt defect detection result and the basic image defect detection result respectively according to the local feature attention enhancement feature, calculate the final anomaly score, and perform defect detection update;

[0102] It should be noted that in this embodiment, the image encoding module selects different local feature attention enhancement mechanism layers according to the preset sequence number to obtain multiple intermediate image block-level features, aligns the intermediate image block-level features with the language prompt text, and obtains the multi-modal alignment features;

[0103] Obtain the language prompt defect detection result according to the language prompt text enhancement feature and the multi-modal alignment feature. The specific algorithm for obtaining the language prompt defect detection result is as follows:

[0104] ,

[0105] Among them, represents the language prompt defect detection result, represents the upsampling operation, represents the cosine similarity calculation operation, represents the multi-modal alignment feature, and represent the normal language prompt text enhancement feature and the abnormal language prompt text enhancement feature respectively, represents the local feature attention enhancement mechanism layer sequence number, represents the total number of selected local feature attention enhancement mechanism layers, represents the exponential function;

[0106] The defect detection module includes a first defect detection sub-module and a second defect detection sub-module. The first defect detection sub-module and the second defect detection sub-module respectively obtain the basic image defect detection result according to the multi-modal alignment feature;

[0107] The specific algorithm for obtaining the basic image defect detection result is as follows:

[0108] ,

[0109] ,

[0110] Among them, and respectively represent the basic image defect detection results obtained by the first defect detection sub-module and the second defect detection sub-module, and respectively represent the first defect detection sub-module and the second defect detection sub-module;

[0111] Calculate the final anomaly score based on the language prompt defect detection result and the basic image defect detection result. The specific algorithm for calculating the final anomaly score is as follows:

[0112] ,

[0113] Among them, represents the final anomaly score, and respectively represent the basic image defect detection results obtained by the first defect detection sub-module and the second defect detection sub-module, represents the language prompt defect detection result;

[0114] Then perform loss calculation. The specific algorithm for loss calculation is as follows:

[0115] ,

[0116] ,

[0117] ,

[0118] Among them, represents the focal loss, represents the dice loss, represents the total loss, represents the total number of pixels, represents the pixel ordinal number, represents the predicted probability, represents the adjustable parameter of the weight, represents the output of the decoder, represents the true label, and represent the loss balance coefficient;

[0119] Threshold filtering is performed on the defect detection results according to the language prompt to obtain the normal area of the image, and the first defect detection sub-module and the second defect detection sub-module of the defect detection module are updated according to the normal area of the image;

[0120] The specific algorithm for the update is as follows:

[0121] ,

[0122] ,

[0123] ,

[0124] Among them, represents the pixel area coordinates, represents the pixel area mark, represents the defect detection result according to the language prompt, represents the threshold learned from the augmented defect sample image data, and represent the first defect detection sub-module and the second defect detection sub-module respectively, represents the learning rate, represents the calculated gradient, represents the total loss;

[0125] Cooperative teaching is carried out on the first defect detection sub-module and the second defect detection sub-module of the defect detection module;

[0126] Loss calculation is performed on the high-confidence image areas in the basic image defect detection results obtained by the first defect detection sub-module to update the second defect detection sub-module, and loss calculation is performed on the high-confidence image areas in the basic image defect detection results obtained by the second defect detection sub-module to update the first defect detection sub-module. The high-confidence image areas used for update in the first defect detection sub-module and the second defect detection sub-module are not updated;

[0127] The specific algorithm for the cooperative teaching is as follows:

[0128] ,

[0129]

[0130]

[0131]

[0132] Among them, and respectively represent the confidence levels of the basic image defect detection results of the first defect detection sub-module and the second defect detection sub-module, represents the defect classification threshold, and respectively represent the basic image defect detection results obtained by the first defect detection sub-module and the second defect detection sub-module, and represent the first defect detection sub-module and the second defect detection sub-module after collaborative teaching update.

[0133] In summary, according to the above defect detection method based on language prompts and collaborative teaching, through data augmentation, the normal sample image data and defect sample image data are expanded, avoiding the problem of insufficient sample data in actual production. Then, by introducing learnable language prompt texts, the detection ability for different defects of different types of products is enhanced. Through the local feature attention enhancement mechanism, more and more accurate local detail information is obtained to enhance the overall detection accuracy. Also, through two independent defect detection result predictions, more comprehensive defect detection results are obtained. At the same time, the two defect detection modules prompt and update each other to make the overall detection more in line with the actual application situation, improving the versatility of the detection method. The present invention improves the accuracy and versatility of the defect detection method. Specifically, the image data of the target to be measured is collected in real time and subjected to data augmentation processing, which is used to expand and obtain the normal sample image data and defect sample image data of the target to be measured, avoiding the problem of deviation in detection accuracy caused by insufficient sample data in actual use. The language prompt text is obtained according to the category of the image data of the target to be measured, and then the language prompt text enhancement feature is obtained according to the language prompt. The language prompt text is in a learnable state and is used to guide feature localization and feature extraction. The attention enhancement processing is performed on the image data of the target to be measured according to the local feature attention enhancement algorithm to obtain the local feature attention enhancement feature. The attention enhancement processing includes multi-modal feature alignment according to the language prompt text enhancement feature, further enhancing the attention to local detail information and avoiding the influence of the complex detection situation of various products and defect types in the actual production process, further improving the overall detection accuracy. According to the local feature attention enhancement feature, the language prompt defect detection result and the basic image defect detection result are respectively obtained to calculate the final anomaly score and perform defect detection update, further improving the versatility of the detection method to be more in line with the actual application scenario. The present invention improves the accuracy and versatility of the defect detection method.

[0134] Please refer to Figure 2 , which shows the flowchart of the polyp image segmentation method based on the detail restoration network proposed in the second embodiment of the present invention. This polyp image segmentation method based on the detail restoration network includes steps S11 to S16, where:

[0135] Step S11: After the data enhancement module collects the image data of the target to be measured in real time, it performs data enhancement. Multiple normal sample image data are obtained according to random affine transformation and sharpening, and then defects are simulated according to Perlin noise and pasted into the normal sample image data to obtain multiple defective sample image data;

[0136] Step S12: The language prompt learning module obtains the language prompt text according to the category of the image data of the target to be measured, and generates the corresponding enhanced feature of the language prompt text through feedforward;

[0137] Step S13: The image encoding module performs attention enhancement processing on the image data of the target to be measured according to the local feature attention enhancement algorithm, and extracts local features;

[0138] Step S14: The image encoding module selects different local feature attention enhancement mechanism layers according to the preset sequence number to obtain multiple intermediate image block-level features, aligns the intermediate image block-level features with the language prompt text to obtain multi-modal alignment features, obtains the language prompt defect detection result according to the enhanced feature of the language prompt text and the multi-modal alignment feature, and obtains the basic image defect detection result according to the multi-modal alignment feature;

[0139] It should be noted that the image encoding module in this embodiment includes 24 local feature attention enhancement mechanism layers, and the preset sequence number is [6, 12, 18, 24], so as to correspondingly select four intermediate image block-level features of the [6, 12, 18, 24] layers.

[0140] Step S15: Calculate the final anomaly score according to the language prompt defect detection result and the basic image defect detection result, and then calculate the loss;

[0141] Step S16: Perform threshold filtering according to the language prompt defect detection result to obtain the normal area of the image, and update the first defect detection sub-module and the second defect detection sub-module of the defect detection module according to the normal area of the image, and perform collaborative teaching on the defect detection module;

[0142] It should be noted that the defect detection method based on language prompt and collaborative teaching of the present invention is compared with the existing defect detection model, and the comparison is carried out under four evaluation indexes: average precision, region overlap rate, area under the pixel-level curve, and area under the image-level curve. The dataset for the comparison experiment uses the MVTec AD defect detection dataset, and the specific comparison results are shown in Table 1 below:

[0143] Table 1

[0144]

[0145] According to the above table, compared with the existing defect detection model, the defect detection method of the present invention is in a comprehensive lead in the four evaluation indicators of average accuracy, area overlap rate, area under the pixel-level curve, and area under the image-level curve, and has made significant progress.

[0146] In summary, according to the above-mentioned defect detection method based on language prompts and collaborative teaching, through data enhancement, the normal sample image data and the defective sample image data are expanded, so as to avoid the problem of insufficient sample data in actual production, and then by introducing learnable language prompt text, the detection ability of different defects of different types of products is enhanced, and more and more accurate local detail information is obtained through the local feature attention enhancement mechanism to enhance the overall detection accuracy, and through two independent defect detection result predictions, more comprehensive defect detection results are obtained. At the same time, the two defect detection modules prompt and update each other, so that the overall detection is more in line with the actual application situation, thereby improving the versatility of the detection method. The present invention improves the accuracy and versatility of the defect detection method. Specifically, the image data of the target to be tested is collected in real time and data enhancement processing is performed. The data enhancement processing is used to expand and obtain normal sample image data and defect sample image data of the target to be tested, thereby avoiding the problem of deviation in detection accuracy caused by insufficient sample data in actual use. A language prompt text is obtained according to the category of the image data of the target to be tested, and then a language prompt text enhancement feature is obtained according to the language prompt. The language prompt text is in a learnable state and is used to guide feature positioning and feature extraction. The image data of the target to be tested is subjected to attention enhancement processing according to a local feature attention enhancement algorithm to obtain a local feature attention enhancement feature. The attention enhancement processing includes multimodal feature alignment according to the language prompt text enhancement feature, further enhancing the attention to local detail information, avoiding the influence of complex detection conditions with a wide variety of products and defects in the actual production process, and further improving the accuracy of the overall detection. According to the local feature attention enhancement feature, the language prompt defect detection result and the basic image defect detection result are respectively obtained to calculate the final abnormality score, and the defect detection update is performed, thereby further improving the versatility of the detection method to be more in line with the actual application scenario. The present invention improves the accuracy and versatility of the defect detection method.

[0147] See also Figure 3 , which is a schematic diagram of the structure of a defect detection system based on language prompts and collaborative teaching proposed in the third embodiment of the present invention, the system includes:

[0148] The data enhancement module 10 is used to collect the image data of the target to be tested in real time and perform data enhancement processing, wherein the data enhancement processing is used to expand and obtain the normal sample image data and defect sample image data of the target to be tested;

[0149] The language prompt learning module 20 is used to obtain a language prompt text according to the category of the target image data to be measured, and then obtain a language prompt text enhanced feature according to the language prompt. The language prompt text is in a learnable state, and the language prompt text is used to guide feature localization and feature extraction;

[0150] The image encoding module 30 is used to perform attention enhancement processing on the target image data to be measured according to the local feature attention enhancement algorithm to obtain a local feature attention enhanced feature. The attention enhancement processing includes multi-modal feature alignment according to the language prompt text enhanced feature;

[0151] The defect detection module 40 is used to obtain a language prompt defect detection result and a basic image defect detection result respectively according to the local feature attention enhanced feature, so as to calculate a final anomaly score and perform defect detection update.

[0152] The present invention also proposes a computer storage medium, on which one or more programs are stored. When the program is executed by a processor, the above-mentioned defect detection method based on language prompt and collaborative teaching is implemented.

[0153] The present invention also proposes a computer device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored on the memory to implement the above-mentioned defect detection method based on language prompt and collaborative teaching.

[0154] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus or device), or in combination with these instruction execution systems, apparatus or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus or device.

[0155] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0156] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, the multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0157] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0158] The above-described embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.

Claims

1. A defect detection method based on language prompts and collaborative teaching, characterized in that: include: Collecting the image data of the target to be tested in real time and performing data enhancement processing, wherein the data enhancement processing is used to expand and obtain the normal sample image data and defect sample image data of the target to be tested; Acquire a language prompt text according to the category of the target image data to be tested, and then acquire a language prompt text enhanced feature according to the language prompt, wherein the language prompt text is in a learnable state and is used to guide feature positioning and feature extraction; Performing attention enhancement processing on the target image data to be tested according to a local feature attention enhancement algorithm to obtain a local feature attention enhancement feature, wherein the attention enhancement processing includes performing multimodal feature alignment according to the language prompt text enhancement feature; According to the local feature attention enhancement feature, the language prompt defect detection result and the basic image defect detection result are respectively obtained to calculate the final anomaly score and perform defect detection update.

2. The defect detection method based on language prompts and collaborative teaching according to claim 1 is characterized in that: The step of real-time acquisition of the target image data to be measured and performing data enhancement processing specifically includes: The data enhancement module collects the image data of the target to be tested in real time and then performs data enhancement; Acquire a plurality of normal sample image data according to random affine transformation and sharpening, and then simulate defects according to Perlin noise to fit the defects to the normal sample image data to acquire a plurality of defect sample image data; The specific algorithm of data enhancement is as follows: , in, represents defect sample image data, Represents a binary image obtained by randomly generating Perlin noise through a random threshold. express The inverse matrix of represents the input image, represents a random texture image, represents the random transparency parameter, Represents element-wise multiplication.

3. The defect detection method based on language prompts and collaborative teaching according to claim 1 is characterized in that: The step of obtaining a language prompt text according to the category of the target image data to be tested, and then obtaining a language prompt text enhancement feature according to the language prompt specifically includes: The language prompt learning module obtains the language prompt text according to the category of the target image data to be tested. The language prompt text includes normal language prompt text and abnormal language prompt text. The specific algorithm for obtaining the language prompt text is as follows: , , in, Indicates normal language prompt text, Indicates abnormal language prompt text. Learnable word embeddings representing normal language prompt text, Learnable word embeddings to represent abnormal language hint text, Indicates the category of the target image data to be tested, Defective boot, Indicates the defect name; According to the language prompt text, the corresponding language prompt text enhancement feature is generated by feeding forward, and the specific algorithm of the feedforward step is as follows: in, Represents the input hint text of the hint layer. Represents the prompt layer, Indicates the prompt layer number, represents the maximum number of learnable hint layers, Indicates learnable hint text.

4. The defect detection method based on language prompts and collaborative teaching according to claim 1 is characterized in that: The step of performing attention enhancement processing on the target image data to be tested according to the local feature attention enhancement algorithm to obtain the local feature attention enhancement feature specifically includes: The image encoding module performs attention enhancement processing on the target image data to be measured according to the local feature attention enhancement algorithm, and the image encoding module includes multiple local feature attention enhancement mechanism layers; Extract local features. The specific algorithm for extracting local features is as follows: in, represents the original output features, represents the local feature attention enhancement feature, represents the image category feature, represents the last local feature, denote the query matrix, key matrix and value matrix respectively, express Attention projection, represents the local feature attention projection, represents the attention function.

5. The defect detection method based on language prompts and collaborative teaching according to claim 1 is characterized in that: The step of respectively acquiring the language prompt defect detection result and the basic image defect detection result according to the local feature attention enhancement feature specifically includes: The image encoding module selects different local feature attention enhancement mechanism layers according to a preset sequence number to obtain multiple intermediate image block-level features, performs multimodal feature alignment on the intermediate image block-level features and the language prompt text, and obtains multimodal alignment features; The language prompt defect detection result is obtained according to the language prompt text enhancement feature and the multimodal alignment feature. The specific algorithm for obtaining the language prompt defect detection result is as follows: , in, Indicates the language prompt defect detection results, represents the upsampling operation, Represents the cosine similarity calculation operation, represents the multimodal alignment feature, and They represent normal language prompt text enhancement features and abnormal language prompt text enhancement features respectively. represents the number of layers of the local feature attention enhancement mechanism, represents the total number of selected local feature attention enhancement mechanism layers, represents the exponential function; The defect detection module includes a first defect detection submodule and a second defect detection submodule, wherein the first defect detection submodule and the second defect detection submodule respectively obtain basic image defect detection results according to multimodal alignment features; The specific algorithm for obtaining the basic image defect detection result is as follows: , , in, and represent the basic image defect detection results obtained by the first defect detection submodule and the second defect detection submodule respectively, and They represent the first defect detection submodule and the second defect detection submodule respectively.

6. The defect detection method based on language prompts and collaborative teaching according to claim 1, characterized in that: The step of calculating the final anomaly score specifically includes: The final abnormality score is calculated based on the language prompt defect detection result and the basic image defect detection result. The specific algorithm for calculating the final abnormality score is as follows: , in, represents the final anomaly score, and represent the basic image defect detection results obtained by the first defect detection submodule and the second defect detection submodule respectively, Indicates language prompt defect detection results; Then the loss calculation is performed, and the specific algorithm of the loss calculation is as follows: , , , in, represents focal loss, represents dice loss, represents the total loss, Represents the total number of pixels, Represents the pixel ordinal number, represents the predicted probability, represents the adjustable parameter of weight, represents the output of the decoder, represents the true label, and Represents the loss balance coefficient.

7. The defect detection method based on language prompts and collaborative teaching according to claim 1 is characterized in that: The step of performing defect detection and updating specifically includes: Performing threshold filtering according to the language prompt defect detection result to obtain a normal image area, so as to update a first defect detection submodule and a second defect detection submodule of the defect detection module according to the normal image area; The specific algorithm of the update is as follows: , , , in, represents the pixel region coordinates, Represents pixel region labeling, Indicates the language prompt defect detection results, represents the threshold learned from the expanded defect sample image data, and represent the first defect detection submodule and the second defect detection submodule respectively, represents the learning rate, represents the calculation of gradient, represents the total loss; Performing collaborative teaching on a first defect detection submodule and a second defect detection submodule of the defect detection module; Performing loss calculation according to the high-confidence image area in the basic image defect detection result obtained by the first defect detection submodule to update the second defect detection submodule, and performing loss calculation according to the high-confidence image area in the basic image defect detection result obtained by the second defect detection submodule to update the first defect detection submodule, and the high-confidence image area used for updating in the first defect detection submodule and the second defect detection submodule is not updated; The specific algorithm of the collaborative teaching is as follows: , in, and Respectively represent the confidence of the basic image defect detection results of the first defect detection submodule and the second defect detection submodule, represents the defect classification threshold, and represent the basic image defect detection results obtained by the first defect detection submodule and the second defect detection submodule respectively, and Represents the first defect detection submodule and the second defect detection submodule after collaborative teaching is updated.

8. A defect detection system based on language prompts and collaborative teaching, characterized in that: include: A data enhancement module is used to collect the image data of the target to be tested in real time and perform data enhancement processing, wherein the data enhancement processing is used to expand and obtain the normal sample image data and defect sample image data of the target to be tested; A language prompt learning module is used to obtain a language prompt text according to the category of the target image data to be tested, and then obtain a language prompt text enhancement feature according to the language prompt, wherein the language prompt text is in a learnable state and is used to guide feature positioning and feature extraction; An image encoding module, used for performing attention enhancement processing on the target image data to be tested according to a local feature attention enhancement algorithm to obtain a local feature attention enhancement feature, wherein the attention enhancement processing includes performing multimodal feature alignment according to the language prompt text enhancement feature; The defect detection module is used to obtain the language prompt defect detection result and the basic image defect detection result respectively according to the local feature attention enhancement feature to calculate the final anomaly score and perform defect detection update.

9. A storage medium, characterized in that: The storage medium stores one or more programs, which, when executed by a processor, implement the defect detection method based on language prompts and collaborative teaching as described in any one of claims 1 to 7.

10. A computer device, characterized in that: The computer device comprises a memory and a processor, wherein: The memory is used to store computer programs; When the processor is used to execute the computer program stored in the memory, the defect detection method based on language prompts and collaborative teaching described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Industrial defect detection method and device based on pre-training model and storage medium

    CN116468725A

  • Visual inspection multitask learning method based on multimodal prompt cooperation

    CN118918447A