Dual-mode industrial defect detection method and device based on SqueezeNet network and medium
By introducing a dual-mode Fire unit and an adaptive hybrid attention mechanism into the SqueezeNet network, the feature collaborative optimization problem of traditional networks under multi-light source conditions is solved, and high-precision defect detection of lightweight networks in complex industrial scenarios is realized.
Patent Information
- Application Number
- CN202510838826.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
Traditional lightweight networks such as SqueezeNet and MobileNet only support single-light image input, resulting in the inability to coordinate the optimization of cross-modal features and the detection accuracy of a single model is not high, especially in complex industrial scenarios, it is difficult to effectively utilize defective features under different lighting solutions.
The dual-modal industrial defect detection method based on the SqueezeNet network is adopted. By acquiring the image sets of multiple workpieces under two lighting conditions, the dual-modal Fire unit is used to dynamically fusion of features, including independent compression layer, cross attention fusion layer and multi-scale expansion layer, combining adaptive hybrid attention and orthogonal loss function, the efficient fusion of multi-modal features is achieved.
While maintaining the advantages of lightweight networks, the detection accuracy is significantly improved, and it is suitable for online defect detection in industrial scenarios such as metal processing parts, achieving efficient multimodal feature fusion and high-precision detection.
Smart Images

Figure CN120355709A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a dual-modal industrial defect detection method, device, and medium based on the SqueezeNet network. Background Art
[0002] The accuracy of industrial defect detection highly depends on the adaptability between the physical characteristics of defects and the lighting scheme. Different lighting schemes can enhance the imaging features of specific types of defects through optical path design (such as incident angle, light intensity distribution).
[0003] However, traditional lightweight networks (such as SqueezeNet, MobileNet) only support single-light source image input. To cover diverse defects in complex industrial scenarios (such as scratches on metal parts, solder joint voids, coating peeling, etc.), it is necessary to independently construct training datasets, annotation systems, and dedicated detection models for each lighting scheme (bright field, dark field, polarized light, etc.), and there is no parameter sharing or feature interaction between models.
[0004] Due to the independent training of detection models in the prior art, cross-modal features cannot be jointly optimized. For example, the tiny scratches on the edge of a metal part present blurred features due to reflection interference in the bright field, but the clear contour features in the dark field are not jointly utilized, resulting in low accuracy of single-model detection. Summary of the Invention
[0005] Based on this, it is necessary to provide a dual-modal industrial defect detection method, device, and medium based on the SqueezeNet network for the above technical problems.
[0006] In a first aspect, an embodiment of the present application provides a dual-modal industrial defect detection method based on the SqueezeNet network, and the method includes:
[0007] Obtain a first image set and a second image set respectively collected for multiple workpieces under two lighting conditions; wherein, the first image set includes multiple first images, and the second image set includes multiple second images;
[0008] Input the first image set and the second image set into the SqueezeNet network for training to obtain a defect detection model. The SqueezeNet network includes a first convolutional layer, a dual-modal Fire unit, a second convolutional layer, and a classifier connected in sequence. The first convolutional layer is used to perform convolutional operations on each of the first images and each of the second images to obtain each first feature map and each second feature map. The dual-modal Fire unit is used to generate each first weight map and each second weight map through adaptive hybrid attention based on each of the first feature maps and each of the second feature maps, and obtain each third feature map based on each of the first feature maps and each of the first weight maps and each fourth feature map based on each of the second feature maps and each of the second weight maps. The second convolutional layer is used to perform convolutional operations on each of the third feature maps and each of the fourth feature maps and then input them into the classifier to obtain the defect types of each workpiece.
[0009] Input the third image and the fourth image collected for the workpiece to be detected under the two lighting conditions into the defect detection model to obtain the defect type of the workpiece to be detected.
[0010] In one embodiment, the dual-modal Fire unit includes a plurality of dual-modal Fire modules connected in sequence. The output of the current dual-modal Fire module is used as the input of the next dual-modal Fire module. The dual-modal Fire module includes an independent compression layer, a cross-attention fusion layer, and a multi-scale expansion layer.
[0011] The independent compression layer is used to perform low-redundancy feature extraction on the input first feature map and second feature map respectively to obtain a fifth feature map and a sixth feature map.
[0012] The cross-attention fusion layer is used to splice the fifth feature map and the sixth feature map to obtain a spliced feature map, and generate a first weight map and a second weight map through adaptive hybrid attention based on the spliced feature map; and obtain a seventh feature map based on the fifth feature map and the first weight map, and obtain an eighth feature map based on the sixth feature map and the second weight map.
[0013] The multi-scale expansion layer is used to perform feature extraction on the seventh feature map based on a convolutional kernel to obtain the third feature map, and perform feature extraction on the eighth feature map based on a convolutional kernel to obtain the fourth feature map.
[0014] In one embodiment, the cross-attention fusion layer includes a channel attention module, a spatial attention module, and a gated fusion network.
[0015] The channel attention module is used to obtain channel attention features based on the spliced feature map through a channel attention mechanism;
[0016] The spatial attention module is used to obtain spatial attention features based on the spliced feature map through a spatial attention mechanism;
[0017] The gated fusion network is used to obtain a first fusion weight map and a second fusion weight map based on the spliced feature map, and combine the channel attention features and the spatial attention features to obtain the first weight map and the second weight map.
[0018] In one embodiment, in the channel attention module, the channel attention feature M C (X) is calculated by the following formula:
[0019]
[0020] In the spatial attention module, the spatial attention feature M S (X) is calculated by the following formula:
[0021]
[0022] In the gated fusion network, the first weight map A1 and the second weight map A2 are calculated by the following formula:
[0023]
[0024] Where: are learnable parameters, is the ReLU activation function, is the Sigmoid function, X is the spliced feature map, GAP is the global average pooling operation, Conv is the convolution operation, W C ,W S are the first fusion weight map and the second fusion weight map respectively.
[0025] In one embodiment, in the independent compression layer, the fifth feature map S1 and the sixth feature map S2 are calculated by the following formula:
[0026]
[0027] In the cross-attention fusion layer, the seventh feature map and the eighth feature map are calculated by the following formula:
[0028]
[0029] In the multi-scale expansion layer, the third feature map Expand(F1) and the fourth feature map Expand(F2) are calculated by the following formula:
[0030]
[0031] where F m and F d are the first feature map and the second feature map respectively.
[0032] In one embodiment, the SqueezeNet network is trained by a total loss function, and the total loss function is composed of a multi-level orthogonality loss function and a cross-entropy loss function weighted. The calculation formula of the total loss function is as follows:
[0033]
[0034] where the is the cross-entropy loss function, and the is the multi-level orthogonality loss function.
[0035] In one embodiment, the calculation formula of the multi-level orthogonality loss function is as follows:
[0036] The dual-modal Fire unit is divided into three levels, and the feature pairs output by each level are obtained respectively;
[0037] Based on the feature pairs output by each level, the orthogonal loss corresponding to each level is calculated, and the orthogonal loss of each level The calculation formula is as follows:
[0038]
[0039] where F1 and F2 are the feature pairs output by each level, and B is the number of batches in model training;
[0040] Based on the weights of each level and the corresponding orthogonal losses, weighted summation is performed to obtain the multi-level orthogonality loss function.
[0041] In a second aspect, an embodiment of the present application further provides a dual-modal industrial defect detection device based on a SqueezeNet network. The device includes:
[0042] An acquisition module, configured to acquire a first image set and a second image set respectively collected under two lighting conditions for a plurality of workpieces; wherein, the first image set includes a plurality of first images, and the second image set includes a plurality of second images;
[0043] A training module, configured to input the first image set and the second image set into a SqueezeNet network for training to obtain a defect detection model. The SqueezeNet network includes a first convolutional layer, a dual-modal Fire unit, a second convolutional layer, and a classifier connected in sequence. The first convolutional layer is configured to perform convolutional operations on each of the first images and each of the second images respectively to obtain each first feature map and each second feature map. The dual-modal Fire unit is configured to generate each first weight map and each second weight map through adaptive hybrid attention based on each of the first feature maps and each of the second feature maps, and obtain each third feature map based on each of the first feature maps and each of the first weight maps and obtain each fourth feature map based on each of the second feature maps and each of the second weight maps. The second convolutional layer is configured to perform convolutional operations on each of the third feature maps and each of the fourth feature maps respectively and then input them into the classifier to obtain the defect types of each of the workpieces.
[0044] A detection module, configured to input the third image and the fourth image respectively collected for the workpiece to be detected under the two lighting conditions into the defect detection model to obtain the defect type of the workpiece to be detected.
[0045] In a third aspect, an embodiment of the present application further provides a computer device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method as described in the first aspect above.
[0046] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium. A computer program is stored in the storage medium, where the computer program, when executed by a processor, implements the method as described in the first aspect above.
[0047] The above-mentioned dual-modal industrial defect detection method, device and medium based on the SqueezeNet network propose a dual-modal dynamic fusion network structure based on the lightweight SqueezeNet network architecture for defect images collected under different lighting schemes (such as bright field, dark field, multi-angle polarized light), which can efficiently fuse multi-modal features and improve the detection accuracy, and is applicable to online defect detection in industrial scenarios such as metal processing parts.
[0048] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0050] Figure 1 It is a hardware structure block diagram of a terminal device for a dual-modal industrial defect detection method based on the SqueezeNet network in an embodiment;
[0051] Figure 2 It is a schematic flowchart of a dual-modal industrial defect detection method based on the SqueezeNet network in an embodiment;
[0052] Figure 3 It is a comparison diagram of images taken under different lighting schemes in an embodiment;
[0053] Figure 4 It is a comparison diagram of images taken under different lighting schemes in another embodiment;
[0054] Figure 5 It is a structure diagram of the SqueezeNet network in an embodiment;
[0055] Figure 6 It is a network structure diagram of the dual-modal Fire module in an embodiment;
[0056] Figure 7 It is a network architecture diagram of the cross-attention fusion layer in an embodiment;
[0057] Figure 8 It is a structure block diagram of a dual-modal industrial defect detection device based on the SqueezeNet network in an embodiment;
[0058] Figure 9 It is a schematic diagram of the structure of a computer device in an embodiment. Specific implementation manners
[0059] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be described and explained below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0060] The method embodiments provided in this embodiment can be executed on a terminal, a computer or a similar computing device. For example, when running on a terminal, Figure 1 It is a hardware structure block diagram of the terminal of the dual-modal industrial defect detection method based on the SqueezeNet network in this embodiment. As Figure 1 shown, the terminal may include one or more ( Figure 1Only one) processor 102 and a memory 104 for storing data are shown. Among them, the processor 102 may include, but is not limited to, processing devices such as a microprocessor MCU or a field-programmable gate array FPGA. The above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 The structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than Figure 1 shown in, or have a different configuration from Figure 1 shown.
[0061] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the dual-modal industrial defect detection method based on the SqueezeNet network in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0062] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by the communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0063] The embodiment of the present application provides a dual-modal industrial defect detection method based on the SqueezeNet network. Taking the method applied to the Figure 1 terminal in as an example for illustration, as Figure 2 shown, the method includes the following steps:
[0064] Step 201, obtain a first image set and a second image set respectively collected under two lighting conditions for a plurality of workpieces; wherein, the first image set includes a plurality of first images, and the second image set includes a plurality of second images.
[0065] Images collected under different lighting conditions contain complementary physical information, such as the global texture of the bright field and the local topography of the dark field. As Figure 3 and Figure 4 are comparison images of different lighting schemes. As shown in Figure 3 , the left side is dark field lighting (low-angle directional light source), and the right side is bright field lighting (high-angle uniform light source). It can be seen from Figure 3 that the bruise in the red area is more obvious under dark field lighting; as shown in Figure 4 , the left side is bright field lighting (high-angle uniform light source), and the right side is dark field lighting (low-angle directional light source). It can be seen from Figure 4 that the bump in the red area is more obvious under bright field lighting.
[0066] Exemplarily, the two lighting conditions of the present application are bright field lighting and dark field lighting respectively. The first image set is a plurality of images collected under bright field lighting conditions, and the second image set is a plurality of images collected under dark field lighting conditions.
[0067] Step 202: Input the first image set and the second image set into the SqueezeNet network for training to obtain a defect detection model. Figure 5 is the SqueezeNet network structure diagram of the present application. As shown in Figure 5 , the SqueezeNet network includes a first convolutional layer, a dual-modal Fire unit 200, a second convolutional layer, and a classifier connected in sequence. The first convolutional layer is used to perform convolutional operations on each of the first images and each of the second images to obtain each first feature map and each second feature map; the dual-modal Fire unit is used to generate each first weight map and each second weight map through adaptive hybrid attention based on each of the first feature maps and each of the second feature maps, and obtain each third feature map based on each of the first feature maps and each of the first weight maps and obtain each fourth feature map based on each of the second feature maps and each of the second weight maps; the second convolutional layer is used to perform convolutional operations on each of the third feature maps and each of the fourth feature maps and then input them into the classifier to obtain the defect types of each workpiece.
[0068] Among them, the network architecture of the SqueezeNet of the present application improves the existing SqueezeNet network architecture by constructing a dual-modal Fire unit. The dual-modal Fire unit is the core unit for dynamically interacting and fusing dual-modal features collected under two lighting conditions, which solves the problem of insufficient utilization of cross-light source complementary features in the traditional SqueezeNet network single-modal architecture.
[0069] Step 203: Input the third image and the fourth image collected for the workpiece to be detected under the two polishing conditions into the defect detection model to obtain the defect type of the workpiece to be detected.
[0070] In the above steps S201 to S203, for defect images collected under different polishing schemes (such as bright field, dark field, multi-angle polarized light), a dual-modal dynamic fusion network structure is proposed based on the lightweight SqueezeNet network architecture, which can efficiently fuse multi-modal features and improve detection accuracy, and is applicable to on-line defect detection in industrial scenarios such as metal processed parts.
[0071] Specifically, the dual-modal Fire unit includes a plurality of sequentially connected dual-modal Fire modules, and the output of the current dual-modal Fire module is used as the input of the next dual-modal Fire module. Figure 6 It is the network structure diagram of the dual-modal Fire module (DualModalFire) in this application. The dual-modal Fire module includes an independent compression layer, a cross-attention fusion layer, and a multi-scale expansion layer.
[0072] The independent compression layer (Dual Squeeze) is used to perform low-redundancy feature extraction on the input first feature map F m and the second feature map F d respectively to obtain a fifth feature map S1 and a sixth feature map S2.
[0073] Mathematical expression:
[0074]
[0075] Among them, the 1×1 convolution kernel compresses the number of channels to 1 / 2 of the original. F m and F d are the first feature map and the second feature map respectively, and S1 and S2 are the fifth feature map and the sixth feature map.
[0076] The independent compression layer of this application performs low-redundancy feature extraction on the inputs (the first feature map ( )) and the second feature map ( )) under the two polishing conditions respectively, avoiding feature confusion between modalities, and has the following advantages: independently processing modal differences (such as the high dynamic range of the bright field and the low signal-to-noise ratio of the dark field), retaining the physical characteristics of the light source; lightweight compression, avoiding interference from redundant features.
[0077] The Cross-Attention Fusion layer is used to splice the fifth feature map S1 and the sixth feature map S2 to obtain a spliced feature map, and generate a first weight map A1 and a second weight map A2 based on the spliced feature map through adaptive hybrid attention; and obtain a seventh feature map based on the fifth feature map S1 and the first weight map A1 , and obtain an eighth feature map based on the sixth feature map S2 and the second weight map A2 .
[0078] Mathematical expression:
[0079] (1) Feature splicing: First, splice the two features obtained by the independent compression layer and to obtain a spliced feature map .
[0080]
[0081] (2) Attention weight generation: Generate weight maps for the spliced features through adaptive hybrid attention to obtain a first weight map A1 and a second weight map A2.
[0082]
[0083] (3) Residual symmetric fusion: Fuse the obtained weight maps with the two features obtained by the independent compression layer through the following formula and to obtain a seventh feature map and an eighth feature map .
[0084]
[0085]
[0086] The cross-attention fusion layer of this application dynamically learns the correlation of features in different modalities (bright field and dark field) through the adaptive hybrid attention mechanism, realizes defect type-driven weight allocation, and the weight is adaptively adjusted according to the defect type, and the residual connection retains the original feature distribution, improving the training convergence speed.
[0087] The multi-scale expansion layer (Expand) is used to extract features from the seventh feature map based on a convolutional kernel to obtain the third feature map Expand(F1), and extract features from the eighth feature map Feature extraction is performed to obtain the fourth feature map Expand(F2).
[0088] Mathematical expression:
[0089]
[0090]
[0091] The multi-scale expansion layer of this application fuses global semantics and local details through multi-scale convolutional kernels, enhancing the ability to express defect features. Respectively use convolutional kernels of size and and to perform feature extraction on
[0092] respectively, and extract features through different receptive fields. Among them, the 1×1 convolution extracts global semantic features, the 3×3 convolution captures local texture details, and the output channel number is restored to 4 times that before compression, maintaining compatibility with the original SqueezeNet network.
[0093] In the industrial quality inspection scenario, the traditional attention mechanism faces two core challenges: First, single-dimensional attention (such as only channel or only space) is difficult to handle the feature differences of multi-form defects. For example, scratches on the metal surface require feature enhancement in the channel dimension, while the position offset of electronic components requires precise positioning in the space dimension; Second, the multi-attention fusion method with fixed weights lacks the adaptive ability to complex backgrounds and is prone to false activation in scenarios with uneven illumination or texture interference. Therefore, the cross-attention fusion layer of this application adopts a double-layer dynamic coupling attention architecture Figure 7 shows the network architecture of the cross-attention fusion layer. Its core innovation lies in the space-adaptive weight allocation mechanism. Different from the traditional method (such as CBAM) that adds fixed weights, this application generates pixel-level fusion weights through a gating network, enabling each spatial position to autonomously select the optimal attention combination mode according to the local feature complexity.
[0094] Such as Figure 7As shown, the cross-attention fusion layer includes a channel attention module, a spatial attention module, and a gated fusion network.
[0095] The channel attention module is used to obtain channel attention features based on the spliced feature map through a channel attention mechanism.
[0096] Mathematical expression:
[0097]
[0098] Where: are learnable parameters, is the ReLU activation function, is the Sigmoid function, X is the spliced feature map , GAP is the global average pooling operation.
[0099] The channel attention module of the present application enables the weight calculation of each channel to perceive the statistical characteristics of the entire image area through the global average pooling (GAP) operation; through the dimensionality reduction - dimensionality increase structure constructed by two fully connected layers, it effectively captures the non-linear dependence relationship between channels.
[0100] The spatial attention module is used to obtain spatial attention features based on the spliced feature map through a spatial attention mechanism.
[0101] Mathematical expression:
[0102]
[0103] Where: Conv is the convolution operation.
[0104] The spatial attention module of the present application uses a 3×3 convolution kernel instead of a larger receptive field, aiming to retain the fine structure of the defect edge; the design of single-channel output enables the module to focus on the significance of spatial positions rather than the channel dimension combination; it has a very small number of parameters, only containing C×1×3×3 = 9C parameters, and only 4,608 parameters are required when C = 512, which is suitable for deployment on edge computing devices.
[0105] The gated fusion network is used to obtain the first fusion weight map and the second fusion weight map based on the spliced feature map, and combine the channel attention features and the spatial attention features to obtain the first weight map and the second weight map.
[0106] Mathematical expression:
[0107]
[0108] Where, are the first fusion weight map and the second fusion weight map respectively.
[0109] The gated fusion network of this application has spatial adaptability: the gating network generates independent fusion weights for each pixel point , achieving fine-grained attention combination control; it has noise robustness: through the probability normalization characteristic of the function, the gating network can still maintain the stability of the weight distribution under noise interference; it has scene perception ability: the convolutional layer of the gating network can implicitly learn the characteristic patterns of different material surfaces.
[0110] In industrial defect detection, similar defects show high intra-class differences due to factors such as deformation and lighting (such as the length and direction changes of metal scratches), while there are local similarities between different types of defects (such as the local texture similarity between cracks and scratches), resulting in a confusing area in the feature space. The traditional Cross-Entropy Loss only focuses on class separability and lacks explicit constraints on the geometric structure of the feature space.
[0111] This application introduces a loss function based on matrix orthogonality to constrain the feature space from both algebraic and geometric perspectives. Algebraic perspective: Force the covariance matrices of different class feature vectors in the embedding space to approach orthogonality, maximizing the distinguishability between classes. Geometric perspective: Through hypersphere mapping, make the same-class feature vectors be tightly distributed on the sphere and maximize the angle between different-class feature vectors.
[0112] In one embodiment, the total loss function is composed of a multi-level orthogonality loss function and a cross-entropy loss function weighted, and the SqueezeNet network is trained through the total loss function.
[0113] Among them, the calculation process of the multi-level orthogonality loss function is as follows:
[0114] Divide the bimodal Fire units in the SqueezeNet network into three levels, as Figure 5 shown. The 0th layer 210 includes bimodal Fire module 2, bimodal Fire module 3, and bimodal Fire module 4; the 1st layer 220 includes bimodal Fire module 5, bimodal Fire module 6, bimodal Fire module 7, and bimodal Fire module 8; the 2nd layer 230 includes bimodal Fire module 9. Then, respectively obtain the feature pairs (F1 and F2) output by each level, and simultaneously constrain the orthogonality of shallow (edge / texture) and deep (semantic) features through the multi-stage features of the SqueezeNet network.
[0115] The orthogonality calculation process is as follows:
[0116] Matrix expansion: Flatten the 4D feature tensor [B, C, H, W] into a 2D matrix [B, CHW], where B is the number of batches for model training, C is the number of channels, H is the height of the feature map, and W is the width of the feature map.
[0117] Correlation matrix calculation: Calculate the feature correlations between samples through matrix multiplication:
[0118]
[0119] Orthogonal loss calculation: Use the Frobenius norm to measure the strength of the non - diagonal elements of the correlation matrix:
[0120]
[0121] Calculate the orthogonal loss for each level through the above method After that, perform a weighted sum of the orthogonal losses of each level with the corresponding weights to obtain the multi - level orthogonality loss .
[0122] Among them, the weights of each level adopt a hierarchical decay mechanism, which can emphasize the decoupling of shallow features (such as edges / textures) and weaken the orthogonal constraints on deep semantic features.
[0123] Adopt an exponential decay mechanism, and the loss weights of different levels decrease according to , for example:
[0124] Weight of the 0th layer: 0.8^0 = 1.0;
[0125] Weight of the 1st layer: 0.8^1 = 0.8;
[0126] Weight of the 2nd layer: 0.8^2 = 0.64;
[0127] The collaborative training of the above multi - level orthogonality loss and cross - entropy loss, the total loss function is linearly weighted by the cross - entropy loss (Cross - Entropy Loss) and the multi - level orthogonal loss (Orthogonal Loss), and the calculation formula of the total loss function is as follows:
[0128]
[0129] Among them: is the standard cross - entropy loss, as the main classification loss, is the multi - level orthogonality loss, is the collaborative weight coefficient (for example, set to 0.1).
[0130] In an example embodiment, images of two lighting schemes are synchronously collected to construct an industrial defect dataset of 2,000 images, which is trained using the SqueezeNet network of the present application. The cross-entropy loss and the orthogonal loss are jointly optimized, and the training period is 100 iteration times. The experimental results are compared in Table 1.
[0131] Table 1
[0132]
[0133] Based on the co-design of the improved dual-modal Fire module of SqueezeNet, its dynamic fusion, adaptive attention (adaptive hybrid attention mechanism with channel-space decoupling), and orthogonal loss, the present application achieves a breakthrough in the industrial defect detection accuracy of 98% while maintaining the lightweight advantage of SqueezeNet (with only 1.40M parameters), providing a high-precision and low-cost intelligent detection solution for edge computing scenarios.
[0134] It should be noted that the steps shown in the above process or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0135] In one embodiment, as Figure 8 shown, a dual-modal industrial defect detection device based on the SqueezeNet network is provided. The device includes:
[0136] An acquisition module 10, configured to acquire a first image set and a second image set respectively collected for multiple workpieces under two lighting conditions; wherein, the first image set includes multiple first images, and the second image set includes multiple second images;
[0137] The training module 20 is configured to input the first image set and the second image set into the SqueezeNet network for training to obtain a defect detection model. The SqueezeNet network includes a first convolutional layer, a dual-modal Fire module, a second convolutional layer, and a classifier connected in sequence. The first convolutional layer is configured to perform convolutional operations on each of the first images and each of the second images to obtain respective first feature maps and second feature maps. The dual-modal Fire module is configured to generate respective first weight maps and second weight maps through adaptive hybrid attention based on the first feature maps and the second feature maps, and obtain respective third feature maps based on the first feature maps and the first weight maps and obtain respective fourth feature maps based on the second feature maps and the second weight maps. The second convolutional layer is configured to perform convolutional operations on each of the third feature maps and each of the fourth feature maps and then input them into the classifier to obtain the defect types of the workpieces.
[0138] The detection module 30 is configured to input the third image and the fourth image respectively collected for the workpiece to be detected under the two lighting conditions into the defect detection model to obtain the defect type of the workpiece to be detected.
[0139] In one embodiment, the dual-modal Fire unit includes a plurality of dual-modal Fire modules connected in sequence. The output of the current dual-modal Fire module serves as the input of the next dual-modal Fire module. The dual-modal Fire module includes an independent compression layer, a cross-attention fusion layer, and a multi-scale expansion layer. The independent compression layer is configured to perform low-redundancy feature extraction on the input first feature map and second feature map respectively to obtain a fifth feature map and a sixth feature map. The cross-attention fusion layer is configured to splice the fifth feature map and the sixth feature map to obtain a spliced feature map, and generate a first weight map and a second weight map through adaptive hybrid attention based on the spliced feature map. And obtain a seventh feature map based on the fifth feature map and the first weight map, and obtain an eighth feature map based on the sixth feature map and the second weight map. The multi-scale expansion layer is configured to perform feature extraction on the seventh feature map based on a convolutional kernel to obtain the third feature map, and perform feature extraction on the eighth feature map based on a convolutional kernel to obtain the fourth feature map.
[0140] In one embodiment, the cross-attention fusion layer includes a channel attention module, a spatial attention module, and a gated fusion network. The channel attention module is configured to obtain channel attention features based on the spliced feature map through a channel attention mechanism. The spatial attention module is configured to obtain spatial attention features based on the spliced feature map through a spatial attention mechanism. The gated fusion network is configured to obtain a first fusion weight map and a second fusion weight map based on the spliced feature map, and combine the channel attention features and the spatial attention features to obtain the first weight map and the second weight map.
[0141] In one embodiment, in the channel attention module, the channel attention feature M C (X) is calculated by the following formula:
[0142]
[0143] In the spatial attention module, the spatial attention feature M S (X) is calculated by the following formula:
[0144]
[0145] In the gated fusion network, the first weight map A1 and the second weight map A2 are calculated by the following formula:
[0146]
[0147] Where: are learnable parameters, is the ReLU activation function, is the Sigmoid function, X is the spliced feature map, GAP is the global average pooling operation, W C , W S are the first fusion weight map and the second fusion weight map, respectively.
[0148] In one embodiment, in the independent compression layer, the fifth feature map S1 and the sixth feature map S2 are calculated by the following formula:
[0149]
[0150] In the cross-attention fusion layer, the seventh feature map and the eighth feature map are calculated by the following formula:
[0151]
[0152] In the multi-scale expansion layer, the third feature map Expand(F1) and the fourth feature map Expand(F2) are calculated by the following formula:
[0153]
[0154] where F m and F d are the first feature map and the second feature map respectively.
[0155] In one embodiment, the training module 20 further includes training the SqueezeNet network through a total loss function, and the total loss function is composed of a multi-level orthogonality loss function and a cross-entropy loss function weighted, and the calculation formula of the total loss function is as follows:
[0156]
[0157] where the is the cross-entropy loss function, and the is the multi-level orthogonality loss function.
[0158] In one embodiment, the calculation formula of the multi-level orthogonality loss function is as follows:
[0159] Divide the dual-modal Fire unit into three levels, and respectively obtain the feature pairs output by each level;
[0160] Based on the feature pairs output by each level, calculate the orthogonal loss corresponding to each level, and the orthogonal loss of each level The calculation formula is as follows:
[0161]
[0162] where F1 and F2 are the feature pairs output by each level, and B is the number of batches of model training;
[0163] Based on the weights of each level and the corresponding orthogonal losses, perform weighted summation to obtain the multi-level orthogonality loss function.
[0164] It should be noted that the above-mentioned modules can be functional modules or program modules, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned modules can be located in the same processor; or the above-mentioned modules can also be located in different processors in any combined form.
[0165] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be asFigure 9 As shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a dual-modal industrial defect detection method based on the SqueezeNet network.
[0166] Those skilled in the art can understand that Figure 9 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0167] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in any of the above embodiments of the dual-modal industrial defect detection method based on the SqueezeNet network.
[0168] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0169] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.
[0170] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A dual-modal industrial defect detection method based on the SqueezeNet network, characterized in that The method includes: Obtaining a first image set and a second image set respectively collected for multiple workpieces under two polishing conditions; wherein, the first image set includes multiple first images, and the second image set includes multiple second images; Inputting the first image set and the second image set into a SqueezeNet network for training to obtain a defect detection model. The SqueezeNet network includes a first convolutional layer, a dual-modal Fire unit, a second convolutional layer, and a classifier connected in sequence. The first convolutional layer is used to perform convolutional operations on each of the first images and each of the second images to obtain each first feature map and each second feature map; the dual-modal Fire unit is used to generate each first weight map and each second weight map through adaptive hybrid attention based on each of the first feature maps and each of the second feature maps, and obtain each third feature map based on each of the first feature maps and each of the first weight maps and obtain each fourth feature map based on each of the second feature maps and each of the second weight maps; the second convolutional layer is used to perform convolutional operations on each of the third feature maps and each of the fourth feature maps and then input them into the classifier to obtain the defect types of each of the workpieces; Inputting a third image and a fourth image respectively collected for the workpiece to be detected under the two polishing conditions into the defect detection model to obtain the defect type of the workpiece to be detected.
2. The method according to claim 1, wherein The dual-modal Fire unit includes multiple dual-modal Fire modules connected in sequence. The output of the current dual-modal Fire module is used as the input of the next dual-modal Fire module. The dual-modal Fire module includes an independent compression layer, a cross-attention fusion layer, and a multi-scale expansion layer; The independent compression layer is used to perform low-redundancy feature extraction on the input first feature map and second feature map respectively to obtain a fifth feature map and a sixth feature map; The cross-attention fusion layer is used to splice the fifth feature map and the sixth feature map to obtain a spliced feature map, and generate a first weight map and a second weight map through adaptive hybrid attention based on the spliced feature map; and obtain a seventh feature map based on the fifth feature map and the first weight map, and obtain an eighth feature map based on the sixth feature map and the second weight map; The multi-scale expansion layer is used to perform feature extraction on the seventh feature map based on a convolutional kernel to obtain the third feature map, and perform feature extraction on the eighth feature map based on a convolutional kernel to obtain the fourth feature map.
3. The method according to claim 2, wherein The cross-attention fusion layer includes a channel attention module, a spatial attention module, and a gated fusion network, The channel attention module is used to obtain channel attention features based on the spliced feature map through a channel attention mechanism; The spatial attention module is used to obtain spatial attention features based on the spliced feature map through a spatial attention mechanism; The gated fusion network is used to obtain a first fusion weight map and a second fusion weight map based on the spliced feature map, and combine the channel attention feature and the spatial attention feature to obtain the first weight map and the second weight map.
4. The method according to claim 3, wherein In the channel attention module, the channel attention feature M C (X) is calculated by the following formula: ; In the spatial attention module, the spatial attention feature M S (X) is calculated by the following formula: ; In the gated fusion network, the first weight map A1 and the second weight map A2 are calculated by the following formula: ; Wherein: is a learnable parameter, is the ReLU activation function, is the Sigmoid function, X is the concatenated feature map, GAP is the global average pooling operation, Conv is the convolution operation, W C , W S are the first fusion weight map and the second fusion weight map respectively.
5. The method according to claim 3, wherein In the independent compression layer, the fifth feature map S1 and the sixth feature map S2 are calculated by the following formula: ; In the cross-attention fusion layer, the seventh feature map and the eighth feature map are calculated by the following formula: ; In the multi-scale expansion layer, the third feature map Expand(F1) and the fourth feature map Expand(F2) are calculated by the following formula: ; Among them, F m and F d are the first feature map and the second feature map respectively.
6. The method according to claim 2, wherein The SqueezeNet network is trained through a total loss function, and the total loss function is composed of a multi-level orthogonality loss function and a cross-entropy loss function weighted. The calculation formula of the total loss function is as follows: ; Among them, is the cross-entropy loss function, is the multi-level orthogonality loss function.
7. The method according to claim 6, wherein The calculation formula of the multi-level orthogonality loss function is as follows: The dual-modal Fire unit is divided into three levels, and the feature pairs output by each level are obtained respectively; Calculate the orthogonal loss corresponding to each level based on the feature pairs output by each of the said levels, the orthogonal loss of each level The calculation formula is as follows: ; wherein, F1 and F2 are the feature pairs output by each level, and B is the number of batches for model training; Based on the weights of each level and the corresponding orthogonal losses, weighted summation is performed to obtain the multi-level orthogonality loss function.
8. A dual-modal industrial defect detection device based on the SqueezeNet network, characterized in that, The device includes: An acquisition module, configured to acquire a first image set and a second image set respectively collected for a plurality of workpieces under two lighting conditions; wherein, the first image set includes a plurality of first images, and the second image set includes a plurality of second images; A training module, configured to input the first image set and the second image set into a SqueezeNet network for training to obtain a defect detection model. The SqueezeNet network includes a first convolutional layer, a dual-modal Fire unit, a second convolutional layer, and a classifier connected in sequence. The first convolutional layer is configured to perform convolutional operations on each of the first images and each of the second images to obtain each first feature map and each second feature map; the dual-modal Fire unit is configured to generate each first weight map and each second weight map through adaptive hybrid attention based on each of the first feature maps and each of the second feature maps, and obtain each third feature map based on each of the first feature maps and each of the first weight maps and obtain each fourth feature map based on each of the second feature maps and each of the second weight maps; the second convolutional layer is configured to perform convolutional operations on each of the third feature maps and each of the fourth feature maps and then input them into the classifier to obtain the defect types of each of the workpieces; A detection module, configured to input a third image and a fourth image respectively collected for a workpiece to be detected under the two lighting conditions into the defect detection model to obtain the defect type of the workpiece to be detected.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Workpiece defect detection method and device fusing multi-attention mechanism
CN113822885A
Intelligent fan blade defect detection method based on bimodal fusion
CN114429457A
Multispectral target detection method and system, computer equipment and storage medium
CN114998667A
Edge detection method and device, electronic equipment and storage medium
CN115511863A
Method and system for detecting surface defects of pultrusion plate for wind power blade
CN119399092A
Cited By
Intelligent AF hardened film surface defect detection method combined with machine vision
CN121207986A