Defect detection model training method, defect detection method, device and equipment
Through the layer-by-layer feature extraction and fusion method, combined with the multi-scale feature extraction and jump connection of the teacher and student models, the problem of shallow information loss in the existing technology is solved, and a more efficient defect detection effect is achieved.
Patent Information
- Application Number
- CN202511006056.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-22
AI Technical Summary
In the existing technology of unsupervised industrial defect detection, shallow feature information is severely lost, resulting in poor detection of high-frequency features such as tiny anomalies and edge cracks. The student network reconstruction ability is limited, making it difficult to accurately locate the defect boundary.
A layer-by-layer feature extraction and fusion method is adopted, combined with the multi-scale feature extraction of the teacher model and the student model. Through the layer-by-layer fusion module and the void space pyramid pooling module, shallow detail information and deep semantic features are retained, and optimization training is performed through skip connection and cosine alignment loss to generate a defect detection model.
It significantly improves the reconstruction capability and accuracy of defect detection, can better identify tiny defects and complex edge structures, and improves the robustness and detection accuracy of the model.
Smart Images

Figure CN120510154B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of industrial visual inspection technology, and in particular to a defect detection model training method, defect detection method, device and equipment. Background Art
[0002] Objective,A class of anomaly detection methods based on reverse distillation is widely used in,unsupervised industrial defect detection tasks.
[0003] This type of method uses a teacher network to generate multi-layer feature maps for industrial images, and then uses a multi-scale feature fusion module (MFF) to downsample shallow features to the deepest feature size and then splice them with the deepest features. A fixed-length embedding vector (One-Class Embedding, OCE) is generated through 1*1 convolution, and then symmetrical reconstruction is performed through a student network.
[0004] However, since this type of method uniformly downsamples all shallow features to the deepest spatial resolution and then fuses them with the deepest features, it seriously destroys information such as shallow texture and edge structure, making the model perform poorly when processing high-frequency features such as tiny anomalies and edge cracks, and easily causing blurred boundaries and positioning offsets in the anomaly detection results.
[0005] In addition, the student network only relies on bottom-up layer-by-layer upsampling to restore features, and its reconstruction ability is limited, especially in the reconstruction of spatial details. Summary of the Invention
[0006] The purpose of the present invention is to provide a defect detection model training method, defect detection method, device and equipment to address the above-mentioned deficiencies in the prior art, so as to effectively retain shallow detail information and deep semantic features and improve the defect detection effect.
[0007] To achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows:
[0008] In a first aspect, an embodiment of the present invention provides a method for training a defect detection model, the method comprising:
[0009] The initial teacher model is used to extract features layer by layer from the sample industrial image to obtain the first feature map of the multi-level sample;
[0010] Using a layer-by-layer fusion module to perform layer-by-layer downsampling and fusion on the first feature map of the multi-level samples to obtain a multi-stage fusion feature map;
[0011] Using an initial student model to perform multi-level upsampling on the multi-stage fusion feature map and the multi-level sample first feature map to obtain a multi-level sample second feature map;
[0012] Calculating a training loss based on the second feature map of the multi-level samples and the first feature map of the multi-level samples;
[0013] Optimizing the initial teacher model and the initial student model according to the training loss to obtain a target teacher model and a target student model;
[0014] A defect detection model is generated according to the target teacher model, the layer-by-layer fusion module and the target student model.
[0015] Optionally, the layer-by-layer fusion module is used to perform layer-by-layer downsampling and fusion on the first feature map of the multi-level samples to obtain a multi-stage fusion feature map, including:
[0016] Downsampling the first feature map of the first level samples by using the layer-by-layer fusion module, and fusing it with the first feature map of the second level samples to obtain a fused feature map;
[0017] The fused feature map is downsampled by using the layer-by-layer fusion module, and is fused with the first feature map of the third layer samples until the fusion is completed with the first feature map of the last layer samples to obtain the multi-stage fused feature map.
[0018] Optionally, the adopting the initial student model to perform multi-level upsampling on the multi-stage fusion feature map and the multi-level sample first feature map to obtain the multi-level sample second feature map includes:
[0019] Upsampling the multi-stage fusion feature map using the first-level upsampling module of the initial student model to obtain a first-level sample second feature map;
[0020] The second feature map of the samples output by each level upsampling module is spliced with the first feature map of the samples at the corresponding level, and the next level upsampling module is used for upsampling to obtain the second feature map of the samples at the next level.
[0021] Optionally, the initial student module further includes: a dilated spatial pyramid pooling module located between adjacent upsampling modules, wherein the second feature map of the samples output by each level of upsampling modules is concatenated with the first feature map of the samples at the corresponding level, and upsampling is performed using the next level upsampling module to obtain the second feature map of the samples at the next level, including:
[0022] The second feature map of the sample output by each level upsampling module is spliced with the first feature map of the sample at the corresponding level to obtain a spliced feature map;
[0023] The spliced feature map is processed using the dilated spatial pyramid pooling module to obtain a context information feature map;
[0024] The context information feature map is upsampled using a next-level upsampling module to obtain a second feature map of the next-level sample.
[0025] Second, an embodiment of the present invention further provides a defect detection method, the method comprising:
[0026] Obtain a target industrial image, perform feature extraction on the target industrial image using a target teacher model and a target student model in a pre-trained defect detection model, and output a plurality of target first feature maps and a plurality of target second feature maps, wherein the defect detection model is trained using the training method of the defect detection model described above;
[0027] generating an abnormality heat map according to the plurality of target first feature maps and the plurality of target second feature maps;
[0028] Defects in the target industrial image are analyzed based on the abnormal thermal map.
[0029] Optionally, generating an abnormal heat map according to the multiple target first feature maps and the multiple target second feature maps includes:
[0030] Calculating pixel cosine distances of the plurality of target first feature maps and the plurality of target second feature maps respectively;
[0031] The abnormal heat map is generated according to the pixel cosine distance.
[0032] Optionally, analyzing defects in the target industrial image according to the abnormal thermal map includes:
[0033] Performing threshold segmentation on the abnormal heat map to obtain a threshold segmentation result;
[0034] Morphological processing is performed on the threshold segmentation result to obtain a defect analysis result of the target industrial image.
[0035] Third, an embodiment of the present invention further provides a training device for a defect detection module, the device comprising:
[0036] The first feature extraction module is used to extract features of the sample industrial image layer by layer using the initial teacher model to obtain a multi-level sample first feature map;
[0037] A feature fusion module is used to downsample and fuse the first feature map of the multi-level samples layer by layer using a layer-by-layer fusion module to obtain a multi-stage fusion feature map;
[0038] A second feature extraction module is used to perform multi-level upsampling on the multi-stage fusion feature map and the multi-level sample first feature map using an initial student model to obtain a multi-level sample second feature map;
[0039] A loss calculation module, configured to calculate a training loss based on the second feature map of the multi-level samples and the first feature map of the multi-level samples;
[0040] A model optimization module, configured to optimize the initial teacher model and the initial student model according to the training loss to obtain a target teacher model and a target student model;
[0041] The model generation module is used to generate a defect detection model based on the target teacher model and the target student model.
[0042] Optionally, the feature fusion module is specifically used to use the layer-by-layer fusion module to downsample the first feature map of the first-level samples, and fuse it with the first feature map of the second-level samples to obtain a fused feature map; use the layer-by-layer fusion module to downsample the fused feature map, and fuse it with the first feature map of the third-level samples, until the fusion is completed with the first feature map of the last-level samples, to obtain the multi-stage fused feature map.
[0043] Optionally, the second feature extraction module is specifically used to use the first-level upsampling module of the initial student model to upsample the multi-stage fusion feature map to obtain the first-level sample second feature map; splice the sample second feature map output by each level upsampling module with the sample first feature map of the corresponding level, and use the next-level upsampling module to upsample to obtain the next-level sample second feature map.
[0044] Optionally, the initial student module also includes: a dilute spatial pyramid pooling module located between adjacent upsampling modules, and the second feature extraction module is further used to splice the sample second feature map output by each level of upsampling module with the sample first feature map of the corresponding level to obtain a spliced feature map; the dilute spatial pyramid pooling module is used to process the spliced feature map to obtain a context information feature map; the next level upsampling module is used to upsample the context information feature map to obtain the next level sample second feature map.
[0045] Fourthly, an embodiment of the present invention further provides a defect detection device, comprising:
[0046] A model output module is used to obtain a target industrial image, use a target teacher model and a target student model in a pre-trained defect detection model to extract features from the target industrial image, and output multiple target first feature maps and multiple target second feature maps. The defect detection model is trained using the training method of the defect detection model described above;
[0047] a heat map generation module, configured to generate an abnormal heat map based on the plurality of target first feature maps and the plurality of target second feature maps;
[0048] The defect analysis module is used to analyze the defects in the target industrial image according to the abnormal thermal map.
[0049] Optionally, the heat map generation module is specifically used to calculate the pixel cosine distances of the multiple target first feature maps and the multiple target second feature maps respectively; and generate the abnormal heat map according to the pixel cosine distances.
[0050] Optionally, the defect analysis module is specifically configured to perform threshold segmentation on the abnormal thermal map to obtain a threshold segmentation result; and perform morphological processing on the threshold segmentation result to obtain a defect analysis result of the target industrial image.
[0051] Fifth, an embodiment of the present invention also provides an electronic device, comprising a processor, a memory, and a communication bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory through the communication bus, and the processor executes the machine-readable instructions to implement any of the methods described above.
[0052] Sixth, an embodiment of the present invention further provides a storage medium, wherein a computer program is stored on the storage medium, and when the computer program is executed by a processor, any of the above methods is executed.
[0053] The beneficial effects of the present invention are:
[0054] The training method, defect detection method, device and equipment of the defect detection model provided by the present invention adopt a teacher model to extract multi-scale features of different depths, wherein shallow features retain rich spatial result information, and deep features contain strong semantic representation capabilities. The multi-scale features are then fused layer by layer based on a layer-by-layer fusion module. In the decoding stage, when upsampling is performed in the student network, the layer-by-layer upsampling results are jump-connected with the features of the teacher network of the same size, and combined with the void space pyramid pooling module, it can effectively retain shallow detail information and deep semantic features, thereby improving the reconstruction capability of the defect area. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0056] Figure 1 Embed anomaly detection network structure for existing technologies;
[0057] Figure 2 Schematic diagram of the training method of the defect detection module provided by the present invention Figure 1 ;
[0058] Figure 3 Schematic diagram of the training method of the defect detection module provided by the present invention Figure 2 ;
[0059] Figure 4 This is an architectural diagram of the defect detection model provided by the present invention;
[0060] Figure 5 Schematic diagram of the training method of the defect detection module provided by the present invention Figure 3 ;
[0061] Figure 6 Schematic diagram of the training method of the defect detection module provided by the present invention Figure 4 ;
[0062] Figure 7 Schematic diagram of the defect detection method provided by the present invention Figure 1 ;
[0063] Figure 8 A schematic structural diagram of a training device for a defect detection module provided by the present invention;
[0064] Figure 9 A schematic structural diagram of a defect detection device provided by the present invention;
[0065] Figure 10 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0066] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0067] The present application is further explained in the following detailed description with reference to the accompanying drawings, wherein:
[0068] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the principles of the application. In the drawings:
[0069] Reference will now be made in detail to some embodiments of the application, one or more examples of which are illustrated in the accompanying drawings. The detailed description, which Figure 1 As shown in the prior art embedded anomaly detection network structure, such as Figure 1 As shown in the prior art embedded anomaly detection network structure, such as
[0070] It can be seen that although the network structure is simple in structure and can balance the semantic and representation embedding expression, due to the loss of shallow structure details and the lack of cross-level skip connection, the student reconstruction ability is limited, and there are obvious short boards in anomaly boundary positioning and fine-grained feature restoration.
[0071] Based on the above problems of the prior art, the present application provides a training method of a defect detection model, which extracts multi-scale features of different depths by using a teacher model, wherein the shallow features retain rich spatial result information, and the deep features contain strong semantic representation ability, then performs layer-by-layer fusion on the multi-scale features based on a layer-by-layer fusion module, and in the decoding stage, performs skip connection on the layer-by-layer up-sampling results and the features of the same size of the teacher network in the student network when up-sampling, and combines a dilated spatial pyramid pooling module, which can effectively retain shallow detail information and deep semantic features, and improve the reconstruction ability of the defect area.
[0072] Further, the student network takes the features extracted by the teacher network as the target, calculates the loss in the training process for supervision, to ensure that the reconstructed features are highly aligned with the teacher network in the direction consistency; and generates an anomaly heat map according to the features extracted by the teacher network and the features extracted by the student network in the inference process.
[0073] It can be seen that the training method of the defect detection model of the present invention fully retains the original structural information of the multi-scale features in the embedded feature construction stage, and realizes the organic combination of hierarchical perception and semantic fusion through step-by-step decoding and jump connection, which significantly improves the expression ability of local defect areas and complex edge structures. Without increasing the computational complexity of the teacher network, it effectively breaks through the performance bottleneck of the existing technology in multi-scale modeling and structural restoration, and has stronger robustness and detection accuracy.
[0074] The specific implementation of the defect detection model training method provided by the present invention is described below in conjunction with specific embodiments.
[0075] Please refer to Figure 2 , which is a schematic diagram of the training method of the defect detection module provided by the present invention Figure 1 ,like Figure 2 As shown, the method may include:
[0076] Step 101: Use the initial teacher model to extract layer-by-layer features of the sample industrial image to obtain a multi-level sample first feature map.
[0077] Specifically, a sample industrial image is obtained and preprocessed, such as normalization and size adjustment. The initial teacher model is used as an encoder to extract features of the preprocessed sample industrial image layer by layer to obtain a multi-level sample first feature map. The multi-level sample first feature maps retain their original spatial resolution to avoid information loss.
[0078] In some embodiments, while retaining the multi-level sample first feature maps, feature statistics are extracted from the multi-level sample first feature maps. The feature statistics exist in the form of additional feature maps with the same size as the feature maps of the corresponding levels. The sample first feature maps of each level are spliced with the additional feature maps of the corresponding level in the channel dimension. The feature statistics can enhance the perception of the noise distribution in the feature map and provide additional discriminant information for subsequent fusion. Among them, the statistical feature quantity can be, for example, the channel standard deviation, the maximum activation value, etc., which is not limited in this embodiment.
[0079] In some embodiments, the initial teacher model may be, for example, a WideResNet-50 backbone network.
[0080] Furthermore, the initial teacher model can also be replaced with other backbone networks with strong multi-scale feature extraction capabilities, such as ResNet-101, EfficientNet series, DenseNet, or the lighter MobileNetV3, to adapt to different computing power environments.
[0081] Step 102: Use a layer-by-layer fusion module to downsample and fuse the first feature map of the multi-level samples layer by layer to obtain a multi-stage fusion feature map.
[0082] Specifically, the layer-by-layer fusion module adopts a Bottleneck structure, inputs the multi-level sample first feature maps into the layer-by-layer fusion module, downsamples the sample first feature maps of each layer to the next layer resolution, and fuses them with the sample first feature maps of the next layer, and further downsamples the fusion results to the next layer resolution, and finally fuses them with the deepest layer sample first feature map, and outputs a multi-stage fusion feature map with the same resolution as the deepest layer sample first feature map.
[0083] In some embodiments, the first feature map of samples at each level is downsampled to the resolution of the next level, and after being concatenated with the first feature map of samples at the next level in the channel dimension, it is compressed and fused through 1*1 convolution.
[0084] Layer-by-layer fusion can not only take into account both shallow details and deep semantics, but also suppress redundant features by normalizing the fusion results at each level, improve the quality of feature expression, and avoid information distortion caused by one-time fusion.
[0085] Furthermore, in addition to the layer-by-layer downsampling Bottleneck structure, the layer-by-layer fusion module can also use the Feature Pyramid Network (FPN) or the adaptive weighted fusion module (such as BiFPN) to achieve multi-scale feature fusion, taking into account both details and semantics.
[0086] Step 103: Use the initial student model to perform multi-level upsampling on the multi-stage fusion feature map and the multi-level sample first feature map to obtain the multi-level sample second feature map.
[0087] Specifically, an initial student model is designed based on the U-Net structure. The initial student model serves as a decoder and includes multiple upsampling modules. The multiple upsampling modules upsample the multi-stage fusion feature maps step by step to restore the spatial resolution. After each upsampling module, the sample first feature map of the same spatial resolution is introduced for skip connection, spliced with the decoded features of each upsampling module and fused through 1*1 convolution compression as the input feature of the next upsampling module to reconstruct multi-scale spatial structure information and realize deep fusion of multi-scale context and spatial details. Multiple upsampling modules output multi-level sample second feature maps.
[0088] In some embodiments, bilinear interpolation is used to perform upsampling on each upsampling module.
[0089] In some embodiments, in addition to using a U-Net structured skip connection decoder as a student model, a Transformer-based decoder (such as a Swin Transformer Decoder), a dual-branch parallel decoder, or a deformable convolutional decoder can also be used to enhance the ability to reconstruct complex shapes and edges.
[0090] Step 104: Calculate the training loss based on the second feature map of the multi-level samples and the first feature map of the multi-level samples.
[0091] Specifically, the multi-level sample first feature map output by the initial teacher model and the multi-level sample second feature map output by the initial student model are aligned at the same spatial resolution, and the training loss is calculated based on the difference between the sample first feature map and the sample second feature map at the same spatial resolution.
[0092] In some embodiments, the training loss is calculated based on the cosine similarity between the sample first feature map and the sample second feature map of the same spatial resolution.
[0093] The cosine similarity-driven directional consistency alignment scheme can effectively suppress non-semantic factors such as illumination changes and disturbing textures, enhance the discrimination of real abnormal deviations, and reduce the false detection and missed detection rates.
[0094] Furthermore, in addition to using cosine similarity loss for directional consistency alignment, L2 distance, structural similarity (SSIM) loss, or contrastive learning paradigms (such as InfoNCE) can also be used to measure the differences in teacher and student features and achieve embedding alignment from different perspectives.
[0095] Step 105: Optimize the initial teacher model and the initial student model according to the training loss to obtain the target teacher model and the target student model.
[0096] Specifically, the model parameters of the initial teacher model and the initial student model are jointly optimized according to the training loss, and the target teacher model and the target student model are obtained when the training rounds reach a preset number of rounds or the training loss converges.
[0097] Step 106: Generate a defect detection model based on the target teacher model, the layer-by-layer fusion module and the target student model.
[0098] Specifically, the target teacher model, the layer-by-layer fusion module and the target student model are combined into a defect detection model, which is used to detect defects in industrial images.
[0099] The training method of the defect detection model provided in the above embodiment adopts multi-scale feature extraction, layer-by-layer feature fusion, jump connection and decoding of the features of the teacher model and the features of the student model, and combines the cosine alignment supervision loss to jointly perform end-to-end joint training of the teacher model and the student model in the same framework, which promotes the collaborative optimization between the modules and significantly improves the convergence speed and stability of the model.
[0100] In one possible implementation, see Figure 3 , which is a schematic diagram of the training method of the defect detection module provided by the present invention Figure 2 ,like Figure 3 As shown, the above step 102 may include:
[0101] Step 121: downsample the first feature map of the first-level samples using a layer-by-layer fusion module, and fuse it with the first feature map of the second-level samples to obtain a fused feature map.
[0102] Step 122: downsample the fused feature map using a layer-by-layer fusion module, and fuse it with the first feature map of the third layer samples until it is fused with the first feature map of the last layer samples to obtain a multi-stage fused feature map.
[0103] For details, please refer to Figure 4 , which is the architecture diagram of the defect detection model provided by the present invention, such as Figure 4 As shown in the figure, the initial teacher model performs layer-by-layer feature extraction on the sample industrial image at multiple spatial resolutions through multiple feature extraction modules to obtain a multi-level sample first feature map.
[0104] This embodiment takes the multi-level sample first feature maps as the shallow feature map feature_a, the middle feature map feature_b, the deep feature map feature_c and the deepest feature map feature_d as an example to illustrate the feature fusion process of the layer-by-layer fusion module.
[0105] Firstly, the layer-by-layer fusion module is used to down-sample the shallow feature map feature_a to the spatial resolution of the middle layer feature map feature_b, and after splicing with the middle layer feature map feature_b, the 1*1 convolution compression fusion is performed to obtain the fusion feature map; then, the layer-by-layer fusion module is used to down-sample the fusion feature map to the spatial resolution of the deep layer feature map feature_c, and after splicing with the deep layer feature map feature_c, the 1*1 convolution compression fusion is performed to obtain a new fusion feature map; then, the layer-by-layer fusion module is used to down-sample the new fusion feature map to the spatial resolution of the deepest layer feature map feature_d, and after splicing with the deepest layer feature map feature_d, the 1*1 convolution compression fusion is performed to obtain the multi-stage fusion feature map. In this way, the layer-by-layer feature fusion can take into account the shallow detail information and deep semantic information, and avoid the distortion of information caused by one-time fusion.
[0106] The training method of the defect detection model provided by the above embodiment retains the shallow texture structure and deep semantic information through the layer-by-layer fusion module, facilitates the completion of the retention in the process of embedding and constructing the multi-stage fusion feature map, and ensures that the shallow details and deep semantics are fully aligned in the student model through the up-sampling decoding with the jump connection, so that the spatial hierarchical structure can be effectively restored, the defect boundary is clearer, and the continuity is better.
[0107] In a possible implementation, referring to Figure 5 , the schematic diagram of the training method of the defect detection module provided by the present application is shown Figure 3 , as shown in the figure Figure 5 , the above step 103 can include:
[0108] Step 131, using the first level up-sampling module of the initial student model to up-sample the multi-stage fusion feature map to obtain the sample second feature map of the first level.
[0109] Step 132, splicing the sample second feature map output by each level up-sampling module with the sample first feature map of the corresponding level, and using the next level up-sampling module to up-sample to obtain the sample second feature map of the next level.
[0110] Specifically, the multi-level up-sampling module in the initial student model is used to continuously up-sample the multi-stage fusion feature map, and the sample second feature map output by each level up-sampling module is jump-connected with the sample first feature map of the same spatial resolution as the input of the next level up-sampling module, and the next level up-sampling module is used to continue up-sampling to obtain the sample second feature map of the next level, and the layer-by-layer up-sampling is performed until the sample second feature map of the same size as the sample first feature map of the first level is obtained, and the decoding of the initial student model is completed.
[0111] For example, the first-level upsampling module is used to upsample the multi-stage fusion feature module to obtain the first-level sample second feature map, and the first-level sample second feature map is spliced with the deep feature map feature_c as the first-level decoding feature; the second-level upsampling module is used to upsample the first-level decoding feature to obtain the second-level sample second feature map, and the second-level sample second feature map is spliced with the middle feature map feature_b as the second-level decoding feature; the third-level upsampling module is used to upsample the second-level decoding feature to obtain the third-level sample second feature map, and the third-level sample second feature image and the shallow feature map feature_a are at the same spatial resolution.
[0112] The training method of the defect detection module provided in the above embodiment introduces the corresponding teacher features after upsampling and recovery at each level of the student model and decodes them through jump connections, deeply fusing the details of the original scale with the fusion features, significantly enhancing the spatial structure ring energy and boundary alignment accuracy, which is the key to improving the detection of small targets and complex boundaries, effectively restoring the spatial hierarchical structure, making the defect boundaries clearer and more continuous.
[0113] In one possible implementation, Figure 4 As shown, the initial student module also includes: a dilated spatial pyramid pooling module located between adjacent upsampling modules, please refer to Figure 6 , which is a schematic diagram of the training method of the defect detection module provided by the present invention Figure 4 ,like Figure 6 As shown, the above step 132 may include:
[0114] Step 201: Splice the second feature map of the samples output by each level of upsampling module with the first feature map of the samples at the corresponding level to obtain a spliced feature map.
[0115] Step 202: Use a dilated spatial pyramid pooling module to process the spliced feature map to obtain a contextual information feature map.
[0116] Step 203: Use the next-level upsampling module to upsample the context information feature map to obtain a next-level sample second feature map.
[0117] Specifically, the multi-stage upsampling module in the initial student module is used to continuously upsample the multi-stage fusion feature map, and the second feature map of the sample output by each level upsampling module is jump-connected with the first feature map of the sample with the same spatial resolution as the splicing feature map of each level.
[0118] An Atrous Spatial Pyramid Pooling (ASPP) module is arranged between adjacent up-sampling modules. The ASPP module includes a plurality of dilated convolution analysis with different dilated rates and a global average pooling branch. The ASPP module is used to process each level of the spliced feature map, and context information of different receptive fields is obtained in parallel. The outputs of the branches are spliced, and then subjected to 1*1 convolution and batch normalization processing, so as to further strengthen the perception ability of the complex background and the micro defect and enhance the embedded multi-scale semantic representation.
[0119] In some embodiments, the ASPP module can be replaced by a Pyramid Spatial Pooling (PSP) module, a dilated attention module or a separable dilated convolution module. A lightweight multi-scale attention (such as PPM+SE) can also be used to supplement the global and local context under different receptive fields.
[0120] The training method of the defect detection model provided in the above embodiments can extract a plurality of receptive field information in parallel through the multi-scale context enhancement module of the ASPP module, so that the model can simultaneously focus on local microscopic cracks and global semantic distribution, thereby significantly improving the recognition ability of micro defects and noise artifacts and enhancing the edge continuity and clarity. In combination with the layer-by-layer fusion module and the ASPP, the original resolution features of each layer are reserved and integrated into the multi-receptive field context. In combination with the skip connection mechanism and the cosine alignment, the student model can capture not only the low-level texture such as micro cracks but also high-level semantics through the step-by-step reconstruction, so as to realize high-universal detection of different scale abnormalities. Further, the student model is provided with the lightweight layer-by-layer fusion module and the ASPP module, so that the overall model parameter quantity and calculation quantity are controllable, the end-to-end training and inference are fast and stable, and the real-time detection demand of high-resolution industrial images is met.
[0121] Based on the training method of the defect detection module, the specific implementation manner of the defect detection method using the defect detection model is described in combination with the embodiments.
[0122] Please refer to Figure 7 The schematic diagram of the defect detection method provided by the present application is shown in Figure 1 As Figure 7 indicated, the method can include the following steps.
[0123] In step 301, a target industrial image is obtained, and a teacher model and a student model in a pre-trained defect detection model are used to extract features of the target industrial image respectively, so as to output a plurality of target first feature maps and a plurality of target second feature maps.
[0124] In step 302, an abnormal heat map is generated according to the plurality of target first feature maps and the plurality of target second feature maps.
[0125] Step 303: Analyze defects in the target industrial image based on the abnormal thermal map.
[0126] Specifically, the defect detection model is trained using the training method of the above-mentioned defect detection model, and the teacher model in the defect detection model is used to extract features of the target industrial image to obtain multiple target first feature maps. The layer-by-layer fusion module in the defect detection model is used to fuse the multiple target first feature maps layer by layer to obtain a target multi-stage fusion image. The student model in the defect detection model is used to upsample the target multi-stage fusion image layer by layer, and the upsampling result of each layer is jump-connected with the target first feature map of the same spatial resolution, and spatial pooling is performed through the ASPP module to output multiple target second feature maps.
[0127] In some embodiments, step 302 may include:
[0128] Calculate the pixel cosine distances of multiple target first feature maps and multiple target second feature maps respectively; and generate an abnormal heat map based on the pixel cosine distances.
[0129] like Figure 4 As shown in the figure, the cosine distance is calculated pixel by pixel for the first feature map and the second feature map of the target with the same size. According to the sum of multiple cosine distances, color mapping is used to generate an abnormal heat map. The higher the sum of the cosine distances, the higher the abnormality of the corresponding pixel.
[0130] Furthermore, anomaly heatmap generation can be combined with adaptive thresholds, conditional random fields (CRFs), or graph-cut-based optimization strategies to further improve boundary accuracy and spatial consistency. High-error areas can also be secondary purified through clustering-based methods (such as DBSCAN).
[0131] In a possible implementation, step 303 may include:
[0132] Perform threshold segmentation on the abnormal thermal map to obtain the threshold segmentation result; perform morphological processing on the threshold segmentation result to obtain the defect analysis result of the target industrial image.
[0133] Specifically, the abnormal heat map is subjected to dynamic or fixed threshold processing, and combined with morphological post-processing, a binary defect detection result is output. Based on the binary defect detection result, the defects of the target industrial image can be analyzed to obtain a defect analysis result.
[0134] The defect detection method provided in the above embodiment performs defect detection based on a defect detection model composed of a teacher model, a student model, a layer-by-layer fusion module, and a feature jump connection. It can effectively identify defects in industrial images and intuitively locate defects by highlighting and reconstructing deviation areas through abnormal heat maps.
[0135] Based on the above embodiment, the present invention also provides a training device for a defect detection module. Figure 8 , which is a structural diagram of the training device for the defect detection module provided by the present invention, such as Figure 8 As shown, the device may include:
[0136] The first feature extraction module 401 is used to extract features layer by layer from the sample industrial image using the initial teacher model to obtain a multi-level sample first feature map;
[0137] A feature fusion module 402 is configured to perform layer-by-layer downsampling and fusion on the first feature map of the multi-level samples using a layer-by-layer fusion module to obtain a multi-stage fused feature map;
[0138] A second feature extraction module 403 is configured to perform multi-level upsampling on the multi-stage fusion feature map and the multi-level sample first feature map using an initial student model to obtain a multi-level sample second feature map;
[0139] A loss calculation module 404 is configured to calculate a training loss based on the second feature map of the multi-level samples and the first feature map of the multi-level samples;
[0140] A model optimization module 405 is configured to optimize the initial teacher model and the initial student model according to the training loss to obtain a target teacher model and a target student model;
[0141] The model generation module 406 is used to generate a defect detection model based on the target teacher model and the target student model.
[0142] Optionally, the feature fusion module 402 is specifically used to use the layer-by-layer fusion module to downsample the first feature map of the first-level samples, and fuse it with the first feature map of the second-level samples to obtain a fused feature map; use the layer-by-layer fusion module to downsample the fused feature map, and fuse it with the first feature map of the third-level samples, until the fusion is completed with the first feature map of the last-level samples, to obtain the multi-stage fused feature map.
[0143] Optionally, the second feature extraction module 403 is specifically used to use the first-level upsampling module of the initial student model to upsample the multi-stage fusion feature map to obtain the first-level sample second feature map; splice the sample second feature map output by each level upsampling module with the sample first feature map of the corresponding level, and use the next-level upsampling module to upsample to obtain the next-level sample second feature map.
[0144] Optionally, the initial student module also includes: a dilated space pyramid pooling module located between adjacent upsampling modules, and the second feature extraction module 403 is further used to splice the sample second feature map output by each level of upsampling module with the sample first feature map of the corresponding level to obtain a spliced feature map; the dilated space pyramid pooling module is used to process the spliced feature map to obtain a context information feature map; the next level upsampling module is used to upsample the context information feature map to obtain the next level sample second feature map.
[0145] The embodiment of the present invention also provides a defect detection device. Figure 9 , which is a structural diagram of the defect detection device provided by the present invention, such as Figure 9 As shown, the device may include:
[0146] A model output module 501 is configured to acquire a target industrial image, perform feature extraction on the target industrial image using a target teacher model and a target student model in a pre-trained defect detection model, and output a plurality of target first feature maps and a plurality of target second feature maps. The defect detection model is trained using the above-described defect detection model training method.
[0147] A heat map generating module 502 is configured to generate an abnormal heat map based on the plurality of target first feature maps and the plurality of target second feature maps;
[0148] The defect analysis module 503 is configured to analyze defects in the target industrial image based on the abnormal thermal map.
[0149] Optionally, the heat map generation module 502 is specifically configured to respectively calculate pixel cosine distances of the plurality of target first feature maps and the plurality of target second feature maps; and generate the abnormal heat map according to the pixel cosine distances.
[0150] Optionally, the defect analysis module 503 is specifically configured to perform threshold segmentation on the abnormal thermal map to obtain a threshold segmentation result; and perform morphological processing on the threshold segmentation result to obtain a defect analysis result of the target industrial image.
[0151] In a possible implementation, the embodiment of the present invention further provides an electronic device, please refer to Figure 10 , is a schematic diagram of an electronic device provided by an embodiment of the present invention, such as Figure 10As shown, the electronic device can include a processor 601, a memory 602 and a communication bus 603, wherein the memory 602 stores machine readable instructions executable by the processor 601, and the processor 601 communicates with the memory 602 through the communication bus 603 when the electronic device is running, and the processor 601 executes the machine readable instructions to implement the method of the above embodiments.
[0152] In a possible implementation, the embodiment of the present application further provides a storage medium, and the storage medium stores a computer program, and the computer program is run by a processor to execute the method of the above embodiment.
[0153] The above merely is a specific implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, and all should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for training a defect detection model, characterized in that: The method comprises: The initial teacher model is used to extract features layer by layer from the sample industrial image to obtain the first feature map of the multi-level sample; Using a layer-by-layer fusion module to perform layer-by-layer downsampling and fusion on the first feature map of the multi-level samples to obtain a multi-stage fusion feature map; Using an initial student model to perform multi-level upsampling on the multi-stage fusion feature map and the multi-level sample first feature map to obtain a multi-level sample second feature map; Calculating a training loss based on the second feature map of the multi-level samples and the first feature map of the multi-level samples; Optimizing the initial teacher model and the initial student model according to the training loss to obtain a target teacher model and a target student model; generating a defect detection model according to the target teacher model, the layer-by-layer fusion module, and the target student model; The layer-by-layer fusion module is used to perform layer-by-layer downsampling and fusion on the first feature map of the multi-level samples to obtain a multi-stage fusion feature map, including: Downsampling the first feature map of the first-level samples by using the layer-by-layer fusion module, and fusing it with the first feature map of the second-level samples to obtain a fused feature map; The fused feature map is downsampled by the layer-by-layer fusion module, and is fused with the first feature map of the third level samples until the fusion is completed with the first feature map of the last level samples to obtain the multi-stage fused feature map; The method of using the initial student model to perform multi-level upsampling on the multi-stage fusion feature map and the multi-level sample first feature map to obtain a multi-level sample second feature map includes: Upsampling the multi-stage fusion feature map using the first-level upsampling module of the initial student model to obtain a first-level sample second feature map; The second feature map of the sample output by each level upsampling module is spliced with the first feature map of the sample at the corresponding level, and upsampling is performed using the next level upsampling module to obtain the second feature map of the sample at the next level; The initial student model further includes: a dilated spatial pyramid pooling module located between adjacent upsampling modules, wherein the second feature map of the samples output by each upsampling module is concatenated with the first feature map of the samples at the corresponding level, and upsampling is performed using the next upsampling module to obtain the second feature map of the samples at the next level, including: The second feature map of the sample output by each level upsampling module is spliced with the first feature map of the sample at the corresponding level to obtain a spliced feature map; The spliced feature map is processed using the dilated spatial pyramid pooling module to obtain a context information feature map; The context information feature map is upsampled using a next-level upsampling module to obtain a second feature map of the next-level sample.
2. A defect detection method, characterized in that: The method comprises: Acquire a target industrial image, perform feature extraction on the target industrial image using a target teacher model and a target student model in a pre-trained defect detection model, and output a plurality of target first feature maps and a plurality of target second feature maps, wherein the defect detection model is trained using the method according to claim 1; generating an abnormality heat map according to the plurality of target first feature maps and the plurality of target second feature maps; Defects in the target industrial image are analyzed based on the abnormal thermal map.
3. The method according to claim 2, characterized in that Generating an abnormal heat map according to the plurality of target first feature maps and the plurality of target second feature maps includes: Calculating pixel cosine distances of the plurality of target first feature maps and the plurality of target second feature maps respectively; The abnormal heat map is generated according to the pixel cosine distance.
4. The method according to claim 3, characterized in that Analyzing defects in the target industrial image according to the abnormal thermal map includes: Performing threshold segmentation on the abnormal heat map to obtain a threshold segmentation result; Morphological processing is performed on the threshold segmentation result to obtain a defect analysis result of the target industrial image.
5. A training device for a defect detection module, characterized in that: The device comprises: The first feature extraction module is used to extract features of the sample industrial image layer by layer using the initial teacher model to obtain a multi-level sample first feature map; A feature fusion module is used to downsample and fuse the first feature map of the multi-level samples layer by layer using a layer-by-layer fusion module to obtain a multi-stage fusion feature map; A second feature extraction module is used to perform multi-level upsampling on the multi-stage fusion feature map and the multi-level sample first feature map using an initial student model to obtain a multi-level sample second feature map; A loss calculation module, configured to calculate a training loss based on the second feature map of the multi-level samples and the first feature map of the multi-level samples; A model optimization module, configured to optimize the initial teacher model and the initial student model according to the training loss to obtain a target teacher model and a target student model; A model generation module, configured to generate a defect detection model based on the target teacher model and the target student model; The feature fusion module is specifically used to use the layer-by-layer fusion module to downsample the first feature map of the first level samples and fuse it with the first feature map of the second level samples to obtain a fused feature map; use the layer-by-layer fusion module to downsample the fused feature map and fuse it with the first feature map of the third level samples until it is completely fused with the first feature map of the last level samples to obtain the multi-stage fused feature map; The second feature extraction module is specifically configured to upsample the multi-stage fusion feature map using the first-level upsampling module of the initial student model to obtain a first-level sample second feature map; concatenate the sample second feature map output by each level upsampling module with the sample first feature map of the corresponding level, and upsample using the next level upsampling module to obtain a next-level sample second feature map; The initial student model also includes: a dilated spatial pyramid pooling module located between adjacent upsampling modules; the second feature extraction module is further used to splice the second feature map of the samples output by each level of upsampling module with the first feature map of the samples at the corresponding level to obtain a spliced feature map; the dilated spatial pyramid pooling module is used to process the spliced feature map to obtain a contextual information feature map; and the next level upsampling module is used to upsample the contextual information feature map to obtain the second feature map of the samples at the next level.
6. A defect detection device, characterized in that: The device comprises: a model output module, configured to acquire a target industrial image, perform feature extraction on the target industrial image using a target teacher model and a target student model in a pre-trained defect detection model, and output a plurality of target first feature maps and a plurality of target second feature maps, wherein the defect detection model is trained using the method of claim 1; a heat map generation module, configured to generate an abnormal heat map based on the plurality of target first feature maps and the plurality of target second feature maps; The defect analysis module is used to analyze the defects in the target industrial image according to the abnormal thermal map.
7. An electronic device, characterized in that: The electronic device comprises a processor, a memory and a communication bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the communication bus, and the processor executes the machine-readable instructions to implement the method according to claim 1, or the method according to any one of claims 2 to 4.
Citation Information
Patent Citations
Smart classroom recording and broadcasting privacy protection method and system based on deep learning
CN118158477A
Defect classification and segmentation method and system based on unsupervised and weak supervised combination
CN120198703A