Defect detection method and device based on visual perception, equipment and medium

By extracting defect features and global features through a multi-encoder network structure and comparing them with reference features, this method addresses the shortcomings of existing visual perception defect detection methods in terms of generalization and reliability. It achieves the ability to identify known defects and unknown anomalies, and solves the problems of environmental adaptability and reliability.

CN120765661BActive Publication Date: 2025-11-18SUZHOU HUI YING OPTICAL TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511292181.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-11-18
Estimated Expiration
2045-09-11

AI Technical Summary

Technical Problem

Existing visual perception-based defect detection methods are insufficient in terms of generalization and reliability, and are difficult to effectively identify unknown defects and adapt to changes in environmental factors.

Method used

A technical solution is adopted, in which defect features are extracted by a first encoder and combined with a fully connected layer to generate defect detection results; a global feature vector is obtained by a second encoder; a mask vector is generated by a third encoder to remove defect-related components; and the difference region is determined by comparing the reference feature vector with the component feature vector.

Benefits of technology

It improves the generalizability and reliability of defect detection, effectively identifies known defects and unknown anomalies, reduces false positives, and enhances adaptability to environmental factors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765661B_ABST
    Figure CN120765661B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of defect detection method, device, equipment and medium based on visual perception, the method extracts defect feature by first encoder and obtains first defect detection result in combination with first fully connected layer, realize the efficient identification of known defect, while through the comparison of component feature vector and reference feature vector, capture new difference or subtle anomaly not covered by defect feature, make up the limitation of defect detection relying on defect feature, with the mask vector generated by third encoder, strip defect-related components from global feature vector, obtain relatively pure component feature vector, avoid the mixed interference of defect feature and normal feature, make the comparison of component feature vector and reference feature vector more focused on the difference of component itself, improve the detection capability of unknown anomaly, and then improve the generalization and reliability of defect detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of defect detection technology, and specifically to a defect detection method, apparatus, equipment and medium based on visual perception. Background Technology

[0002] In industrial production lines, surface defect detection of electronic components, mechanical parts, and other products is a crucial step in ensuring product quality. Traditional defect detection methods mainly rely on manual visual inspection, which suffers from low efficiency, high subjectivity, high missed detection rates, and incompatibility with high-speed production lines. With the development of machine vision technology, automated defect detection methods based on visual perception are gradually becoming mainstream. Their core is to automatically identify surface defects such as scratches, dents, and stains through image processing and analysis.

[0003] Existing visual perception-based defect detection methods mainly include edge segmentation-based defect detection, template matching-based defect detection, and deep learning-based defect detection. Edge segmentation-based defect detection is difficult to generalize to images with unknown defects, and like template matching-based defect detection, it is greatly affected by environmental factors such as lighting and shadows. Deep learning-based defect detection is difficult to effectively separate defect features from normal features, resulting in a high false positive rate.

[0004] Therefore, improving the generalization and reliability of visual perception-based defect detection has become an urgent problem to be solved. Summary of the Invention

[0005] To achieve the objectives of this invention, the technical solution adopted is as follows: a defect detection method based on visual perception, which includes the following steps:

[0006] S101, Obtain an initial image of the target element, input the initial image into the trained first encoder to obtain a defect feature vector, input the defect feature vector into the trained first fully connected layer to obtain a first defect detection result;

[0007] S102, the initial image is input into the trained second encoder to obtain the global feature vector;

[0008] S103, input the defect feature vector and the global feature vector into the trained third encoder to obtain the mask vector;

[0009] S104, determine the element feature vector based on the mask vector and the global feature vector;

[0010] S105, obtain M reference feature vectors in the storage unit, and determine the feature difference degree between the element feature vector and its closest reference feature vector, where M is a positive integer;

[0011] S106, if the degree of feature difference is less than a preset first difference threshold or greater than a preset second difference threshold, then the first defect detection result shall be used as the target detection result.

[0012] S107, otherwise, determine the difference vector based on the element feature vector and its nearest reference feature vector;

[0013] S108, determine the difference region image based on the difference vector, and use the difference region image and the first defect detection result as the target detection result.

[0014] Compared with the prior art, the beneficial effects of the present invention are:

[0015] 1. By extracting defect features through the first encoder and combining them with the first fully connected layer to obtain the first defect detection result, the efficient identification of known defects is achieved. At the same time, by comparing the component feature vector with the reference feature vector, novel differences or subtle anomalies not covered by defect features are captured, which makes up for the limitations of relying on defect features for defect detection. It takes into account both the identification accuracy of known defects and the detection capability of unknown anomalies, and improves the generalization and reliability of defect detection.

[0016] 2. By using the mask vector generated by the third encoder, defect-related components are extracted from the global feature vector to obtain a relatively pure component feature vector. This avoids the interference of mixed defect features and normal features, and makes the comparison between the component feature vector and the reference feature vector more focused on the differences in the component itself. This improves the detection capability of unknown anomalies, and thus improves the generalization and reliability of defect detection. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a defect detection method based on visual perception according to Embodiment 1 of the present invention.

[0018] Figure 2 This is a schematic diagram of a defect detection device based on visual perception according to Embodiment 2 of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Example 1:

[0021] like Figure 1 As shown, the present invention provides a technical solution: a defect detection method based on visual perception, characterized in that the defect detection method based on visual perception includes the following steps:

[0022] S101, Obtain the initial image of the target element, input the initial image into the trained first encoder to obtain the defect feature vector, input the defect feature vector into the trained first fully connected layer to obtain the first defect detection result.

[0023] The target component can refer to electronic components, mechanical components, or other components that can perform surface defect detection. In this embodiment, the component for surface defect detection is placed on a high-speed conveyor belt and moves with the high-speed conveyor belt. A vision sensor is set above the high-speed conveyor belt. The vision sensor collects images of the high-speed conveyor belt from an orthophoto perspective. By default, a single component can be covered by the field of view of a single image.

[0024] The initial image can refer to the image obtained after image preprocessing of the acquired image of the target element. Image preprocessing can include image processing operations such as image denoising, contrast enhancement, and distortion correction. It should be noted that, in order to isolate the influence of background factors, in this embodiment, image preprocessing also includes element region sub-image extraction, element region sub-image rotation correction, and element region sub-image size normalization. Element region sub-image extraction and element region sub-image rotation correction can be implemented based on the bounding box information extracted by the trained rotating target detection model. The trained rotating target detection model can be the YOLO model.

[0025] Specifically, the first encoder may include a convolutional layer, a pooling layer, a normalization layer, and an activation function layer. The first encoder is used to extract defect feature information of the surface of the target element in the initial image. The defect feature vector is used to characterize the defect feature information of the surface of the target element in the initial image.

[0026] The first fully connected layer can be used to map the defect feature vector extracted by the first encoder into a defect category vector. The defect category vector is processed by the softmax function to obtain the first defect detection result, which can represent the defect type corresponding to the surface of the target component.

[0027] S102, input the initial image into the trained second encoder to obtain the global feature vector.

[0028] The second encoder can be used to extract global feature information of the surface of the target element in the initial image. The global feature vector is used to characterize the global feature information of the surface of the target element in the initial image.

[0029] Specifically, the second encoder may also include convolutional layers, pooling layers, normalization layers, and activation function layers.

[0030] In one specific implementation, the training process for the first encoder and the second encoder includes:

[0031] Obtain several sample images of the i-th defect type, where the sample image of the 1st defect type is a positive sample, and the sample images of the 2nd to the i-th defect types are all negative samples, where i is the defect type identifier, i is an integer in the range [1, I], and I is the number of defect types;

[0032] Use any positive sample or any negative sample as the first training sample, and use any positive sample or any negative sample as the second training sample, wherein the first training sample and the second training sample are different;

[0033] The first training sample is input into the first encoder to obtain the first sample feature vector;

[0034] The second training sample is input into the preset twin encoder to obtain the feature vector of the second sample. The twin encoder is always the same as the first encoder.

[0035] Based on the defect types corresponding to the first and second training samples, the type indicator value is determined.

[0036] The first training loss is calculated based on the first sample feature vector, the second sample feature vector, the type indicator value, and the preset first loss function;

[0037] The first training sample is input into the second encoder to obtain the feature vector of the third sample;

[0038] The feature vector of the third sample is input into the preset auxiliary decoder to obtain the first reconstructed sample;

[0039] The second training sample is input into the second encoder to obtain the feature vector of the fourth sample;

[0040] The feature vector of the fourth sample is input into the auxiliary decoder to obtain the second reconstructed sample;

[0041] The second training loss is calculated based on the first training sample, the first reconstructed sample, the second training sample, the second reconstructed sample, and the preset second loss function;

[0042] The third training loss is calculated based on the first training sample and the defect type identifier corresponding to the first training sample, the feature vector of the first sample, the feature vector of the third sample, the second training sample and the defect type identifier corresponding to the second training sample, the feature vector of the second sample, the feature vector of the fourth sample, and the preset third loss function.

[0043] The target training loss is determined based on the first training loss, the second training loss, and the third training loss.

[0044] Based on the target training loss, the first encoder, the second encoder, and the auxiliary decoder are trained until the target training loss converges, resulting in the trained first encoder, the trained second encoder, and the trained auxiliary encoder.

[0045] The first encoder and the twin encoder form a twin network model architecture. The architecture and model parameters of the twin encoder are always the same as those of the first encoder. The second encoder and the auxiliary decoder form a reconstruction network model architecture.

[0046] Defect types can include no defects, scratches, dents, stains, etc. The implementer can determine the defect type according to the actual component to be tested. For example, when the component is an electronic component, the defect type can also include missing electrodes, solder on resistors, missing resistors, broken corners and edges, and poor coding.

[0047] Specifically, the type indicator value is determined based on the defect types corresponding to the first training sample and the second training sample, respectively. This can mean that when the defect types corresponding to the first training sample and the second training sample are both the first defect type, the type indicator value can be 1; when the defect types corresponding to the first training sample and the second training sample are not the first defect type and the defect types are different, the type indicator value can be 2; when the defect types corresponding to the first training sample and the second training sample are one the first defect type and one a non-first defect type, the type indicator value can be 3; and when the defect types corresponding to the first training sample and the second training sample are not the first defect type and the defect types are the same, the type indicator value can be 4.

[0048] Let the feature vector of the first sample be F1, the feature vector of the second sample be F2, and the type indicator value be a. Then the first loss function can be expressed as:

[0049]

[0050] in, j takes the value of an integer in the range [1,4], σ is the standard deviation coefficient, and σ is set to a minimum value. In this embodiment, σ is set to 0.1. It can be known that when a=j, g j (a)=1, when a≠j, g j(a) is approximately 0, that is, the type indicator value is used to control the portion of the first training loss that actually affects training.

[0051] D = dis(F1, F2), where dis(F1, F2) represents the Euclidean distance between the first sample feature vector and the second sample feature vector, and b1 is the first distance threshold. b1 is set to a minimum value, which is set to 1 in this embodiment. It can be seen that when the defect types corresponding to the first training sample and the second training sample are both the first defect type, that is, when the first training sample and the second training sample are both positive samples, the Euclidean distance between the first sample feature vector and the second sample feature vector is only allowed to have a small difference. Otherwise, the first training loss will be large, thereby weakening the non-defect features.

[0052] b2 is the second distance threshold. b2 is set to a smaller value, which is 100 in this embodiment. When the defect types corresponding to the first training sample and the second training sample are not the first defect type and the defect types are different, that is, when the first training sample and the second training sample are both negative samples and the defect types are different, the Euclidean distance between the feature vector of the first sample and the feature vector of the second sample should be greater than b2. Otherwise, the first training loss will be large, thereby strengthening the defect specificity to a certain extent.

[0053] b3 is the third distance threshold. b3 is set to a large value, which is 500 in this embodiment. When the defect types corresponding to the first training sample and the second training sample are a first defect type and a non-first defect type, that is, when the first training sample and the second training sample are a positive sample and a negative sample, the Euclidean distance between the feature vector of the first sample and the feature vector of the second sample should be greater than b3. Otherwise, the first training loss will be large, thereby strengthening the feature difference under the condition of having or not having defects.

[0054] When the defect types corresponding to the first training sample and the second training sample are not the first defect type and are the same, that is, when the first training sample and the second training sample are both negative samples and have the same defect type, the Euclidean distance between the feature vectors of the first sample and the feature vectors of the second sample should be as small as possible. Otherwise, the first training loss will be large, thereby weakening the difference in defect features between the same defect type. Compared with the first training loss when they are both positive samples, the negative samples of the same defect type do not have a first distance threshold for strong supervision, and the feature vector difference constraint of the negative samples of the same defect type is relatively loose.

[0055] Training based on the first training loss can minimize the difference in feature vectors between positive samples and between negative samples of the same defect type, minimize the difference in feature vectors between negative samples of different defect types, and maximize the difference in feature vectors between positive and negative samples, so that the first encoder pays more attention to defect feature information.

[0056] The second loss function can be expressed as: ,in, As the first training sample, As the first reconstructed sample, As the second training sample, The second reconstructed sample is used as the second loss function to supervise the reconstruction effect, so that the second encoder pays more attention to global feature information.

[0057] The third loss function can be expressed as: Where i1 is the defect type identifier corresponding to the first training sample, and i2 is the defect type identifier corresponding to the second training sample. When i1=1, g1(i1)=1, and when i1≠1, g1(i1) is close to 0. When i2=1, g1(i2)=1, and when i2≠1, g1(i2) is close to 0. The third loss function supervises the feature vectors obtained by the first encoder and the second encoder respectively to be close, so as to enhance the attention of the first encoder to the defect feature information.

[0058] The target training loss can be obtained by adding the first training loss, the second training loss, and the third training loss together.

[0059] Based on the target training loss, gradient descent is used to train the first encoder, the second encoder, and the auxiliary decoder until the target training loss converges, resulting in the trained first encoder, the trained second encoder, and the trained auxiliary encoder.

[0060] S103, input the defect feature vector and the global feature vector into the trained third encoder to obtain the mask vector.

[0061] The mask vector is used to extract defect feature information from the global feature vector.

[0062] Specifically, the third encoder may also include convolutional layers, pooling layers, normalization layers, and activation function layers. The third encoder is used to determine the mask vector based on the defect feature vector and the global feature vector.

[0063] In one specific implementation, the training process of the third encoder includes the following steps:

[0064] The first training sample is input into the trained first encoder to obtain the first reference feature vector;

[0065] The first training sample is input into the trained second encoder to obtain the second reference feature vector;

[0066] The first reference feature vector and the second reference feature vector are fused to obtain the reference fused feature vector.

[0067] The reference fused feature vector is input into the third encoder to obtain the encoded reference vector;

[0068] The reference intermediate feature vector is determined based on the encoding reference vector and the second reference feature vector;

[0069] The fourth training loss is calculated based on the defect type identifier corresponding to the first training sample, the reference intermediate feature vector, the first reference feature vector, the second reference feature vector, and the preset fourth loss function.

[0070] The third encoder is trained based on the fourth training loss until the fourth training loss converges, resulting in a well-trained third encoder.

[0071] In this context, feature fusion of the first reference feature vector and the second reference feature vector can refer to the concatenation operation of the first reference feature vector and the second reference feature vector.

[0072] The second reference feature vector is masked based on the encoded reference vector to obtain the intermediate reference feature vector. For details on the masking process, please refer to the subsequent process of determining the component feature vector based on the mask vector and the global feature vector.

[0073] Specifically, the fourth loss function can be expressed as:

[0074]

[0075] Where U is the reference intermediate feature vector, V1 is the first reference feature vector, V2 is the second reference feature vector, and b4 is the fourth distance threshold, which is set to a large value, 500 in this embodiment. When i1=1, g1(i1)=1; when i1≠1, g1(i1) is close to 0. That is, when the first training sample is a positive sample, the fourth loss function is approximately... At this point, there is no defect feature information. Therefore, both the second reference feature vector and the intermediate reference feature vector represent component feature information. The second reference feature vector and the intermediate reference feature vector should be as close as possible. When the first training sample is a negative sample, the fourth loss function is approximately... At this point, there is defect feature information. Therefore, the reference intermediate feature vector represents the component feature information, and the first reference feature vector represents the defect feature information. Therefore, the feature difference between the first reference feature vector and the reference intermediate feature vector should be large. The fourth training loss is used to enable the third encoder to learn the ability to accurately distinguish between defect features and non-defect features.

[0076] S104, determine the component feature vector based on the mask vector and the global feature vector.

[0077] Among them, the component feature vector can represent the non-defect feature vector.

[0078] In one specific implementation, the mask vector and the global feature vector have the same size;

[0079] Based on the mask vector and the global feature vector, the component feature vector is determined, including:

[0080] For any element, subtract the element value corresponding to that element in the mask vector from the first preset value to obtain the mask coefficient corresponding to that element.

[0081] Multiply the mask coefficient corresponding to the element by the element value corresponding to the element in the global feature vector to obtain the mask value corresponding to the element;

[0082] Add the mask value corresponding to the element to the second preset value to obtain the element feature value corresponding to the element;

[0083] Traverse all elements to obtain the component feature value corresponding to each element, and form the component feature vector from the component feature value corresponding to each element.

[0084] The first preset value can be 1, meaning that any element in the mask vector can take values ​​in the range [0, 1]. The second preset value can be a very small value, such as 0.1, to avoid the situation where the component feature vector is a zero matrix.

[0085] S105, obtain M reference feature vectors in the storage cell, and determine the degree of feature difference between the element feature vector and its closest reference feature vector.

[0086] Where M is a positive integer, the storage unit can be implemented using a cache unit, and the reference feature vector can be extracted from the template image of different components, or it can be a feature vector added to the storage unit in the previous processing flow.

[0087] Specifically, the component feature vector is compared with M reference feature vectors by distance measurement. The distance measurement can be Euclidean distance, cosine similarity, etc. The smallest distance measurement result is determined as the feature difference between the component feature vector and its closest reference feature vector.

[0088] In one specific implementation, obtaining M reference feature vectors from the storage unit includes:

[0089] Obtain the M reference feature vectors in the storage unit and the memory weights corresponding to the M reference feature vectors.

[0090] Among them, the memory weights can characterize the reliability of the reference feature vector.

[0091] S106, if the degree of feature difference is less than the preset first difference threshold or greater than the preset second difference threshold, then the first defect detection result shall be used as the target detection result.

[0092] The first difference threshold is used to filter out component feature vectors that have only a slight difference from the closest reference feature vector. At this point, it can be considered that the component feature vector does not have any abnormalities, and the first defect detection result can be used as the target detection result.

[0093] The second difference threshold is used to filter out component feature vectors that differ too much from the closest reference feature vector. At this point, it can be assumed that the reference feature vector that can be used as the target component template does not exist in the storage unit, and it is impossible to continue the judgment. Alternatively, the first defect detection result can be used as the target detection result.

[0094] In one specific implementation, if the degree of feature difference is less than a preset first difference threshold, or greater than a preset second difference threshold, then the first defect detection result is used as the target detection result, including:

[0095] If the degree of feature difference is less than the preset first difference threshold, the memory weight corresponding to the reference feature vector that is closest to the component feature vector will be increased by one, and the first defect detection result will be used as the target detection result.

[0096] If the degree of feature difference is greater than the preset second difference threshold, the storage capacity N of the storage unit is obtained. When N>M, the component feature vector is added to the storage unit, and the first defect detection result is used as the target detection result. When N=M, the reference feature vector with the smallest memory weight is selected as the feature vector to be replaced. The component feature vector is used to replace the feature vector to be replaced in the storage unit, and the first defect detection result is used as the target detection result.

[0097] If the degree of feature difference is less than the preset first difference threshold, it means that the component feature vector is only slightly different from the closest reference feature vector. The reference feature vector is relatively reliable, and the memory weight of the reference feature vector is increased.

[0098] If the difference in features exceeds the preset second difference threshold, it is considered that the reference feature vector that can be used as the target component template does not exist in the storage unit. Typically, components are placed in batches according to component type on the conveyor belt. In order to improve the reliability of subsequent defect detection, the component feature vector is added to the storage unit as a reference feature vector. When N>M, there is a free area in the storage unit, and it can be stored directly. When N=M, there is no free area in the storage unit, and it needs to be stored in a replacement manner. The reference feature vector with the smallest memory weight is selected as the feature vector to be replaced. When there are multiple reference feature vectors with the smallest memory weight, the reference feature vector stored first is selected as the feature vector to be replaced.

[0099] S107, otherwise, determine the difference vector based on the component's eigenvector and its closest reference eigenvector.

[0100] In this process, the element feature vector is subtracted element by element from its nearest reference feature vector to obtain the difference vector, which can characterize feature anomaly information.

[0101] In one embodiment, when the degree of feature difference is greater than or equal to a preset first difference threshold and less than or equal to a preset second difference threshold, it is considered that there is a certain difference between the component feature vector and its closest reference feature vector, and there may be undetected defects. The memory weight of the reference feature vector closest to the component feature vector is reduced by one. Since it is impossible to determine whether the abnormality is in the reference feature vector or the abnormality in the component feature vector, the component feature vector also needs to be stored in the storage unit to support subsequent defect detection.

[0102] S108, Based on the difference vector, determine the difference region image, and use the difference region image and the first defect detection result as the target detection result.

[0103] The difference region image is only used to represent the approximate abnormal area in the image for manual confirmation by regulatory personnel.

[0104] In one specific implementation, determining the difference region image based on the difference vector includes:

[0105] The difference vector is input into the trained auxiliary decoder to obtain the difference region image.

[0106] Among them, due to the specific feature reconstruction capability of the trained auxiliary decoder, the difference vector can be directly reconstructed into a difference region image using the trained auxiliary decoder. The implementer can also perform binarization processing on the difference region image to highlight the difference region information.

[0107] In this embodiment, defect features are extracted by the first encoder and combined with the first fully connected layer to obtain the first defect detection result, thereby achieving efficient identification of known defects. At the same time, by comparing the component feature vector with the reference feature vector, novel differences or subtle anomalies not covered by defect features are captured, which makes up for the limitations of relying on defect features for defect detection. With the help of the mask vector generated by the third encoder, defect-related components are stripped from the global feature vector to obtain a relatively pure component feature vector, avoiding the mixed interference of defect features and normal features. This makes the comparison between the component feature vector and the reference feature vector more focused on the differences of the component itself, improving the detection capability of unknown anomalies, and thus improving the generalization and reliability of defect detection.

[0108] Example 2:

[0109] like Figure 2 As shown, the present invention provides a technical solution: a defect detection device based on visual perception, characterized in that the defect detection device based on visual perception includes:

[0110] The first detection module 201 is used to acquire an initial image of the target element, input the initial image into the trained first encoder to obtain a defect feature vector, and input the defect feature vector into the trained first fully connected layer to obtain a first defect detection result.

[0111] Feature extraction module 202 is used to input the initial image into the trained second encoder to obtain a global feature vector;

[0112] The mask prediction module 203 is used to input the defect feature vector and the global feature vector into the trained third encoder to obtain the mask vector;

[0113] The mask processing module 204 is used to determine the component feature vector based on the mask vector and the global feature vector;

[0114] The difference determination module 205 is used to obtain M reference feature vectors in the storage unit and determine the degree of feature difference between the element feature vector and its closest reference feature vector, where M is a positive integer;

[0115] The second detection module 206 is used to take the first defect detection result as the target detection result if the feature difference is less than a preset first difference threshold or greater than a preset second difference threshold.

[0116] Vector determination module 207 is used to otherwise determine the difference vector based on the component feature vector and its closest reference feature vector;

[0117] The third detection module 208 is used to determine the difference region image based on the difference vector, and use the difference region image and the first defect detection result as the target detection result.

[0118] It should be noted that the specific limitations of the visual perception-based defect detection device can be found in the limitations of the visual perception-based defect detection method above, and will not be repeated here. The information interaction and execution process between the above modules are based on the same concept as the method embodiments of this invention, and their specific functions and technical effects can be found in the method embodiments section, and will not be repeated here.

[0119] Example 3:

[0120] This invention provides a technical solution: a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the visual perception-based defect detection method described in the above embodiments. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the embodiment of the visual perception-based defect detection device.

[0121] Example 4:

[0122] This invention provides a technical solution: a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the visual perception-based defect detection method described in the above embodiments. Alternatively, when executed by a processor, the computer program implements the functions of each module / unit in the above embodiment of the visual perception-based defect detection device.

[0123] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0124] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0125] The embodiments disclosed herein are preferred embodiments, but are not limited thereto. Those skilled in the art can readily grasp the spirit of the present invention based on the above embodiments and make different extensions and variations, but as long as they do not depart from the spirit of the present invention, they are all within the protection scope of the present invention.

Claims

1. A defect detection method based on visual perception, characterized in that, The visual perception-based defect detection method includes the following steps: S101, Obtain an initial image of the target element, input the initial image into the trained first encoder to obtain a defect feature vector, input the defect feature vector into the trained first fully connected layer to obtain a first defect detection result; S102, the initial image is input into the trained second encoder to obtain the global feature vector; S103, input the defect feature vector and the global feature vector into the trained third encoder to obtain the mask vector; S104, determine the element feature vector based on the mask vector and the global feature vector; S105, obtain M reference feature vectors in the storage unit, and determine the feature difference degree between the element feature vector and its closest reference feature vector, where M is a positive integer; S106, if the degree of feature difference is less than a preset first difference threshold or greater than a preset second difference threshold, then the first defect detection result shall be used as the target detection result. S107, otherwise, determine the difference vector based on the element feature vector and its nearest reference feature vector; S108, determine the difference region image based on the difference vector, and use the difference region image and the first defect detection result as the target detection result.

2. The defect detection method based on visual perception according to claim 1, characterized in that, The training process for the first encoder and the second encoder includes: Obtain several sample images of the i-th defect type, where the sample image of the 1st defect type is a positive sample, and the sample images of the 2nd to the i-th defect types are all negative samples, where i is the defect type identifier, i is an integer in the range [1, I], and I is the number of defect types; Use any positive sample or any negative sample as the first training sample, and use any positive sample or any negative sample as the second training sample, wherein the first training sample and the second training sample are different; The first training sample is input into the first encoder to obtain the first sample feature vector; The second training sample is input into a preset twin encoder to obtain the second sample feature vector. The twin encoder is always the same as the first encoder. Based on the defect types corresponding to the first training sample and the second training sample respectively, determine the type indicator value; The first training loss is calculated based on the first sample feature vector, the second sample feature vector, the type indicator value, and the preset first loss function; The first training sample is input into the second encoder to obtain the feature vector of the third sample; The feature vector of the third sample is input into a preset auxiliary decoder to obtain the first reconstructed sample; The second training sample is input into the second encoder to obtain the fourth sample feature vector; The fourth sample feature vector is input into the auxiliary decoder to obtain the second reconstructed sample; The second training loss is calculated based on the first training sample, the first reconstructed sample, the second training sample, the second reconstructed sample, and a preset second loss function; The third training loss is calculated based on the first training sample and the defect type identifier corresponding to the first training sample, the first sample feature vector, the third sample feature vector, the second training sample and the defect type identifier corresponding to the second training sample, the second sample feature vector, the fourth sample feature vector, and the preset third loss function. The target training loss is determined based on the first training loss, the second training loss, and the third training loss; Based on the target training loss, the first encoder, the second encoder, and the auxiliary decoder are trained until the target training loss converges, resulting in the trained first encoder, the trained second encoder, and the trained auxiliary encoder.

3. The defect detection method based on visual perception according to claim 2, characterized in that, The training process of the third encoder includes the following steps: The first training sample is input into the trained first encoder to obtain the first reference feature vector; The first training sample is input into the trained second encoder to obtain the second reference feature vector; The first reference feature vector and the second reference feature vector are fused to obtain a reference fused feature vector. The reference fused feature vector is input into the third encoder to obtain the encoded reference vector; Based on the encoded reference vector and the second reference feature vector, a reference intermediate feature vector is determined; The fourth training loss is calculated based on the defect type identifier corresponding to the first training sample, the reference intermediate feature vector, the first reference feature vector, the second reference feature vector, and the preset fourth loss function. The third encoder is trained according to the fourth training loss until the fourth training loss converges, thus obtaining the trained third encoder.

4. The defect detection method based on visual perception according to claim 1, characterized in that, The mask vector and the global feature vector have the same size; The step of determining the component feature vector based on the mask vector and the global feature vector includes: For any element, the element value corresponding to that element in the mask vector is subtracted from the first preset value to obtain the mask coefficient corresponding to that element. Multiply the mask coefficient corresponding to the element by the element value corresponding to the element in the global feature vector to obtain the mask value corresponding to the element; Add the mask value corresponding to the element to the second preset value to obtain the element feature value corresponding to the element; Traverse all elements to obtain the component feature value corresponding to each element, and form the component feature vector from the component feature value corresponding to each element.

5. The defect detection method based on visual perception according to claim 1, characterized in that, The acquisition of the M reference feature vectors in the storage unit includes: Obtain the M reference feature vectors in the storage unit and the memory weights corresponding to the M reference feature vectors respectively.

6. The defect detection method based on visual perception according to claim 5, characterized in that, If the degree of feature difference is less than a preset first difference threshold, or greater than a preset second difference threshold, then the first defect detection result is used as the target detection result, including: If the degree of feature difference is less than a preset first difference threshold, then the memory weight corresponding to the reference feature vector that is closest to the component feature vector is increased by one, and the first defect detection result is used as the target detection result. If the degree of feature difference is greater than a preset second difference threshold, the storage capacity N of the storage unit is obtained. When N>M, the component feature vector is added to the storage unit, and the first defect detection result is used as the target detection result. When N=M, the reference feature vector with the smallest memory weight is selected as the feature vector to be replaced. The component feature vector is used to replace the feature vector to be replaced in the storage unit, and the first defect detection result is used as the target detection result.

7. The defect detection method based on visual perception according to claim 2, characterized in that, Determining the difference region image based on the difference vector includes: The difference vector is input into the trained auxiliary decoder to obtain the difference region image.

8. A defect detection device based on visual perception, characterized in that, The visual perception-based defect detection device includes: The first detection module is used to acquire an initial image of the target element, input the initial image into a trained first encoder to obtain a defect feature vector, and input the defect feature vector into a trained first fully connected layer to obtain a first defect detection result. The feature extraction module is used to input the initial image into the trained second encoder to obtain a global feature vector; The mask prediction module is used to input the defect feature vector and the global feature vector into the trained third encoder to obtain the mask vector. A mask processing module is used to determine the component feature vector based on the mask vector and the global feature vector; The difference determination module is used to obtain M reference feature vectors in the storage unit and determine the degree of feature difference between the element feature vector and its closest reference feature vector, where M is a positive integer; The second detection module is used to take the first defect detection result as the target detection result if the degree of feature difference is less than a preset first difference threshold or greater than a preset second difference threshold. A vector determination module is used to otherwise determine a difference vector based on the element feature vector and its closest reference feature vector. The third detection module is used to determine the difference region image based on the difference vector, and use the difference region image and the first defect detection result as the target detection result.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the visual perception-based defect detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the visual perception-based defect detection method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Pedestrian re-recognition model training method and device

    CN116612500A

  • Artificial board surface defect intelligent detection method and system based on machine vision

    CN120355688A