Metal surface defect classification model construction method based on structured state space model and application
By using a structured state space model combined with deep convolution and bidirectional Mamba module in metal surface defect classification, the problems of low classification accuracy and large computing overhead in the existing technology are solved, efficient and accurate defect classification is achieved, and production costs are reduced.
Patent Information
- Application Number
- CN202510196723.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The prior art has the problem of low classification accuracy in metal surface defect classification, especially when dealing with complex defects and variable data sets, the calculation overhead is large and the local detail capture capability is insufficient.
A neural network based on structured state space model is adopted, combined with deep convolution and bidirectional Mamba modules, a multi-scale feature extraction and fusion module is built, and the classification performance and generalization capabilities of the model are improved through transfer learning and pyramid structure network architecture.
It significantly improves the accuracy and efficiency of metal surface defect classification, reduces production costs, and shows good generalization capabilities on variable data sets.
Smart Images

Figure CN120164016A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial defect detection, and more specifically, relates to a method for constructing a metal surface defect classification model based on a structured state space model and its application. Background Art
[0002] Industrial defect detection is a key link in modern manufacturing production. Especially in the process of metal processing production, it is of great significance for the control of metal quality. In modern production, deep learning is widely used in industrial defect detection. By constructing a deep neural network to extract the deep features of images, automatic defect detection is completed. The deep neural network has a strong fitting ability and is very suitable for processing irregular and variable defect samples. Without professional knowledge, efficient, accurate and robust defect detection can be achieved by collecting sample data. Automatic defect detection is mainly divided into two types of tasks: defect classification and defect localization. Among them, the defect classification task aims to judge whether there are defects in the image and determine the defect category; the defect localization task is to analyze the position of the defects in the image, mark one or more defects with a bounding box and give a confidence score. Complete defect detection includes two parts: defect localization and defect classification. Efficiently and accurately completing the defect classification task is the basis for realizing real-time defect detection.
[0003] The existing neural networks applied to defect classification mainly include convolutional neural networks (CNNs), generative adversarial networks (GANs), Transformer networks, etc. CNNs have strong local feature extraction capabilities, but there are limitations in their global understanding ability due to inductive bias. Vision architectures based on Transformer include Vision Transformer (ViT), Swin-Transformer, etc.; ViT divides the image into small patches and processes these patches through the self-attention mechanism to capture and integrate global context information; Swin-Transformer combines local and global features through a sliding window mechanism. Networks in the Transformer series can achieve higher classification performance with fewer parameters and can handle extremely complex defect situations. However, the quadratic complexity of the self-attention mechanism introduces a large amount of computational overhead, and dividing the image into fixed-size patches results in insufficient local detail capture ability of the model. Moreover, the defects in the metal defect dataset are relatively large, and there may be significant differences in the appearance of the same type of defect, such as inclined scratches and horizontal scratches, and there are differences in lighting, background, and materials during shooting, and the gray levels of the same type of defect will also vary; at the same time, there is a certain degree of similarity between different types of defects. Directly using a neural network to classify defects also has the problem of low classification accuracy. Summary of the Invention
[0004] In view of the above defects or improvement requirements of the prior art, the present invention provides a method and application for constructing a metal surface defect classification model based on a structured state space model, aiming to improve the accuracy and efficiency of metal surface defect classification and reduce production costs.
[0005] To achieve the above object, the present invention provides a method for constructing a metal surface defect classification model based on a structured state space model, including:
[0006] Pre-train a neural network based on a structured state space model using an image dataset, and fine-tune the pre-trained neural network using a metal surface defect dataset to obtain a metal surface defect classification model; wherein, the neural network based on the structured state space model includes:
[0007] A Patch layer for downsampling the input image sample and increasing the number of channels.
[0008] A multi-scale feature extraction module including depth convolutional layers and bidirectional Mamba layers of different scales; the depth convolutional layers are used to parallelly capture local spatial feature maps of different scales of the output image of the Patch layer in the spatial dimension; the bidirectional Mamba layer includes two Mamba structures after removing the fully connected layers and a fully connected layer, one Mamba structure is used to obtain a forward output feature based on the one-dimensional data obtained by downsampling and flattening the local spatial feature map, and the other Mamba structure is used to obtain a reverse output feature based on the data obtained by flipping the one-dimensional data; the fully connected layer is used to linearize the feature after fusing the forward output feature and the reverse output feature to obtain the feature map output by the bidirectional Mamba layer.
[0009] A multi-scale feature fusion module for multi-scale fusion of the feature maps output by the depth convolutional layers and the feature maps output by the bidirectional Mamba layer to obtain a fused feature map.
[0010] A classification head for generating a classification result based on the fused feature map.
[0011] Further, the depth convolutional layer is an MBConv layer, and the multi-scale feature extraction module includes four Stages connected in a pyramid structure in sequence: Stage0 - Stage3, and the scales of the feature maps output by Stage0 - Stage3 gradually decrease, and the corresponding number of feature channels doubles;
[0012] Among them, Stage0 includes 2 MBConv layers connected in sequence, Stage1 includes 6 MBConv layers connected in sequence, Stage2 includes 14 bidirectional Mamba layers connected in sequence, and Stage3 includes 2 bidirectional Mamba layers connected in sequence;
[0013] The multi-scale feature fusion module is used to fuse the feature maps output by Stage0 - Stage3 to obtain the fused feature map.
[0014] Further, the forward output feature y forward is:
[0015] y forward = x proj ⊙ SiLU(z)
[0016] x proj = SSM(SiLU(Conv1d(x)))
[0017] The backward output feature y backward is:
[0018]
[0019] where SSM(·) represents the output of the main path of the Mamba structure after removing the fully connected layer, SiLU(·) is the SiLU activation function, Conv1d(·) represents the one-dimensional convolution operation, and ⊙ represents the dot product operation; x and z respectively represent the inputs of the main path and the residual path in the Mamba structure obtained by linearly mapping the one-dimensional data; x b and z b are respectively the inputs of the main path and the residual path in the Mamba structure obtained by linearly mapping the data after flipping the one-dimensional data.
[0020] Further, the multi-scale feature fusion module is a multi-scale feature fusion module M2SF with learnable scaling factors, specifically including: an SE layer, a normalization layer, an average pooling layer, a scaling factor learned through training, and a merging layer;
[0021] The SE layer is used to perform channel adaptive adjustment on the feature maps output by each Stage. After the adjusted feature maps are sequentially processed by the corresponding normalization layer and average pooling layer, feature maps of the same size are obtained; after the feature maps are multiplied by the corresponding scaling factors, they are merged in the channel dimension by the merging layer to obtain the fused feature map.
[0022] Further, the metal surface defect dataset is a dataset obtained by merging multiple metal surface defect datasets of different sizes.
[0023] Further, based on the metal surface defect dataset, the pre-trained neural network is fine-tuned using the method of transfer learning to obtain a metal surface defect classification model.
[0024] The present invention also provides a method for classifying metal surface defects, including:
[0025] Inputting the metal surface defect image to be classified into the metal surface defect classification model to obtain a defect classification result; wherein, the metal surface defect classification model is constructed by the metal surface defect classification model construction method described in any one of the above.
[0026] The present invention also provides an electronic device, including a computer-readable storage medium and a processor;
[0027] The computer-readable storage medium is used for storing executable instructions;
[0028] The processor is used for reading the executable instructions stored in the computer-readable storage medium to execute the metal surface defect classification model construction method described in any one of the above, or execute the metal surface defect classification method described above.
[0029] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the metal surface defect classification model construction method described in any one of the above, or implements the metal surface defect classification method described above.
[0030] The present invention also provides a computer program product, including a computer program, and when the computer program runs on a computer, it causes the computer to execute the metal surface defect classification model construction method described in any one of the above, or execute the metal surface defect classification method described above.
[0031] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0032] (1) The present invention combines CNN with a structured state space model Mamba that has linear computational complexity and long sequence modeling capabilities to construct a neural network model structure that combines deep convolution and bidirectional Mamba modules. Specifically, when considering the use of Mamba for understanding metal surface defect images, the causal pattern inherent in its unidirectional modeling has an adverse effect on the understanding of global features. To mitigate this adverse factor, the bidirectional Mamba designed in the present invention processes the input from two directions. Compared with the causal pattern of the original Mamba, the bidirectional Mamba in the present invention is closer to the fully visible pattern of Transformer, avoiding the limitations of causal constraints to a certain extent, being more suitable for the global understanding of images, and avoiding the large computational overhead and insufficient local detail capture ability brought by using a Transformer-based network structure. Therefore, the present invention uses deep convolution to parallelly capture the local spatial features of image data in the spatial dimension, and then integrates the global information through the long sequence modeling advantage of bidirectional Mamba, enabling the model structure to simultaneously focus on local and global features, greatly improving the accuracy of the metal surface defect classification model while taking into account computational efficiency, and reducing production costs.
[0033] (2) Preferably, the present invention connects Mamba modules and deep convolution modules of different scales together in a pyramid structure to build a new network architecture with four Stages, which simultaneously has the optimal local two-dimensional spatial feature capture ability and the optimal global long sequence modeling ability, and avoids the quadratic computational complexity of the attention mechanism, achieving efficient and accurate classification performance on the metal surface defect dataset.
[0034] (3) Preferably, the multi-scale features extracted by the feature pyramid are fused through a multi-scale feature fusion module M2SF with learnable scaling factors, mixing features of different levels, adaptively adjusting the weights of each level for datasets of different scales, and enhancing the generalization ability of the classification model.
[0035] (4) Preferably, in order to cope with the changing situations in the industrial production process, the images collected may not guarantee a unified size. The metal surface defect dataset used in the present invention is a dataset obtained by merging metal surface defect datasets of various different sizes, further enhancing the generalization ability of the model.
[0036] (5) Due to issues such as cost and standardization, the scale of industrial defect datasets is generally small. Using a pre-trained model generated on a general dataset for transfer learning reduces the training time and computational resource cost of the metal surface defect classification model, speeds up the model convergence speed, and is more suitable for industrial scenario applications. Description of the Drawings
[0037] Figure 1 This is a flowchart of the method for constructing a metal surface defect classification model based on a structured state space model in an embodiment of the present invention.
[0038] Figure 2 This is a schematic diagram of the Mamba structure and the bidirectional Mamba structure (Bi-Mamba) provided in an embodiment of the present invention.
[0039] Figure 3 This is a neural network structure diagram of the hybrid of deep convolution and bidirectional Mamba provided in an embodiment of the present invention.
[0040] Figure 4 This is a structure diagram of the multi-scale feature fusion module M2SF provided in an embodiment of the present invention.
[0041] Figure 5 This is the defect classification confusion matrix of the Resnet18 model.
[0042] Figure 6 This is the defect classification confusion matrix of the ViT-B model.
[0043] Figure 7 This is the defect classification confusion matrix of the model without the M2SF module in an embodiment of the present invention.
[0044] Figure 8 This is the defect classification confusion matrix of the model including the M2SF module in an embodiment of the present invention. Detailed implementation manners
[0045] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0046] Embodiment 1
[0047] As Figure 1 - Figure 2 shown, the method for constructing a metal surface defect classification model based on a structured state space model in an embodiment of the present invention includes:
[0048] Pre-train a neural network based on a structured state space model using an image dataset, and fine-tune the pre-trained neural network using a metal surface defect dataset to obtain a metal surface defect classification model.
[0049] In the embodiments of the present invention, the image dataset for pre-training the neural network is the ImageNet1K dataset. In other embodiments, other image datasets can also be selected to pre-train the neural network based on the structured state space model.
[0050] Preferably, the present invention takes into account that some existing methods perform well on defect datasets of the same size, but when faced with defect data of multiple sizes in the dataset, the accuracy decreases and the generalization ability is not strong. In order to cope with the changing situations in the industrial production process, the images collected may not guarantee a unified size. The metal surface defect dataset adopted in the present invention is a dataset obtained by merging metal surface defect datasets of multiple different sizes, so as to improve the generalization ability of the model.
[0051] In the embodiments of the present invention, the obtained two metal surface defect datasets with different sizes and annotations are merged and divided into a training set and a validation set according to a ratio. The dataset is from the Surface Defect Database of Northeastern University, which is a relatively authoritative public dataset in the field of metal surface defect detection. One hot-rolled steel plate defect dataset is NEU-CLS, with a unified pixel size of 200×200, collecting six typical surface defects of hot-rolled steel strips, namely, rolling scale (Rs), patches (Pa), cracks (Cr), pitted surface (Ps), inclusions (In) and scratches (Sc). There are 300 samples for each category, totaling 1800 image samples. Another dataset is NEU-CLS-64, with a unified pixel size of 64×64, which is an extension of the NEU-CLS dataset, expanding the defect categories to 9 categories (Cr, Gg, In, Pa, Ps, Rp, Rs, Sc, Sp), where Gg represents grinding marks, Rp represents pressed-in scale, and Sp represents scale spalling.
[0052] The neural network in the embodiments of the present invention is a model structure combined with a deep convolutional MBConv module and a bidirectional Mamba module, specifically including a Patch layer, a feature extraction module, a multi-scale feature fusion module, and a classification head.
[0053] The Patch layer is used to downsample the input original image and increase the number of channels.
[0054] The multi-scale feature extraction module includes depth convolution layers and bidirectional Mamba layers of different scales. The depth convolution layer in the embodiment of the present invention is an MBConv layer, that is, a mobile inverted bottleneck convolution (MBConv). The depth convolution layers of different scales are used to parallelly capture local spatial features of different scales of the image processed by the Patch layer through depth convolution. After obtaining sufficient local two-dimensional spatial information, the obtained local spatial features are downsampled and flattened into one-dimensional data, and then input into the bidirectional Mamba layer (bidirectional Mamba encoder). The bidirectional Mamba layer includes: two Mamba structures after removing the fully connected layer and a fully connected layer. One Mamba structure is used to input the flattened one-dimensional data to obtain a forward output feature. The other Mamba structure is used to input the data obtained by flipping the flattened one-dimensional data to obtain a reverse output feature. The fully connected layer is used to linearize the feature after fusing the forward output feature and the reverse output feature to obtain the feature map output by the bidirectional Mamba layer.
[0055] The multi-scale feature fusion module is used to fuse the feature map output by the depth convolution layer and the feature map output by the bidirectional Mamba layer to obtain a fused feature map.
[0056] The classification head is used to generate a classification result based on the fused feature map. In the embodiment of the present invention, the classification head is a linear output layer that outputs 9-way prediction values.
[0057] In the embodiment of the present invention, Mamba is a structured state space model that introduces a state space model into the sequence modeling task of deep learning. In the task of image understanding, the Transformer series models usually divide the image data into multiple Patches of fixed size, use the Patches as Tokens and input them into the model, and establish global dependencies through the self-attention mechanism. When Mamba is used for image understanding, this setting can also be followed, but the inherent causal pattern of its unidirectional modeling will impose additional causal constraints on the mixing of Tokens, which has an adverse effect on global understanding. To alleviate this adverse factor, the present invention designs a bidirectional Mamba to process the input from two directions. Specifically, assume that the recurrence process of the existing unidirectional Mamba is y′ = SSM(x), and assume that the input sequence is input (in the embodiment of the present invention, it is the one-dimensional data flattened after passing through the MBConv layer above). After being input into the Mamba module, it is mapped through a linear layer to obtain the input x of the main path and the input z of the residual path. The calculation process of the main path is as follows:
[0058] x proj = SSM(SiLU(Conv1d(x))) (1)
[0059] Among them, SiLU(·) is the SiLU activation function, and Conv1d(·) represents a one-dimensional convolution operation.
[0060] After the hybrid residual path, the forward output feature y is obtained forward :
[0061] y forward = x proj ⊙SiLU(z) (2)
[0062] Among them, ⊙ represents the dot product operation.
[0063] Then, the input is flipped in the sequence dimension to obtain input_b, and x b and z b are obtained through linear layer mapping. Similarly, the reverse output feature y is obtained backward :
[0064]
[0065] After fusing the features in two directions and passing through the linear output layer, the output y of the bidirectional Mamba module can be obtained:
[0066]
[0067] Compared with the causal mode of the original Mamba, the bidirectional Mamba is closer to the fully visible mode of the Transformer, avoiding the limitations of causal constraints to a certain extent and being more suitable for the global understanding of images.
[0068] In the embodiment of the present invention, the Patch layer is composed of 2 3×3 convolutional layers. The stride of the first convolutional layer is 2, and the original 3-channel image is output into a 48-channel feature map with halved size. The second convolutional layer does not change the number of channels and size of the feature map, further enhancing the expression ability of the network.
[0069] Preferably, as Figure 3As shown, the feature extraction module in the embodiment of the present invention includes four Stages connected in series in sequence. Stage0 includes 2 MBConv layers connected in series, Stage1 includes 6 MBConv layers connected in series, Stage2 includes 14 bidirectional Mamba layers, and Stage3 includes 2 bidirectional Mamba layers. The MBConv layer is used to capture two-dimensional local spatial features, including a 1×1 convolutional layer, a 3×3 depth convolutional layer that expands the number of channels by 4 times, an SE module, and a 1×1 convolutional layer that restores the number of channels to the original value. Among them, the SE module is used to perform global average pooling on the spatial dimension of the input feature map, generate the weights of each channel through two fully connected layers and a non-linear activation function, and then act on each channel of the original feature map to achieve adaptive adjustment of channel information. The first MBConv layer of each Stage additionally undertakes the function of downsampling, which is achieved by changing the stride of the first convolutional layer from 1 to 2.
[0070] Preferably, in order to maintain good performance when processing data of different sizes, a bottom-up pyramid structure is used to build the network. The four Stages are connected together in a pyramid structure, and the scales of the feature maps processed by Stage0 - Stage3 gradually decrease, and the corresponding number of feature channels doubles. Arranging the four Stages in a pyramid structure enables features to flow among depth convolutions (MBConv layers) and bidirectional Mamba modules of different widths, obtaining features of different resolutions and semantic levels, which can enable the model to process small targets and large targets simultaneously, enhancing the generalization ability and classification performance of the model.
[0071] As a further design of the present invention, the multi-scale feature fusion module in the embodiment of the present invention is a multi-scale feature fusion module M2SF with learnable scaling factors, as Figure 4 shown, specifically including: an SE layer (Squeeze-and-Excitation Layer), a normalization layer, an average pooling layer, a scaling factor learned through training, and a merging layer.
[0072] The SE layer is used to perform channel adaptive adjustment on the feature maps output by each Stage; in the embodiment of the present invention, there are four SE layers, which are respectively used to perform channel adaptive adjustment on the feature maps output by the four Stages. The adjusted feature maps of different scales are sequentially processed by the normalization process of the normalization layer and the average pooling layer. The average pooling layer uses kernels of different sizes to average pool the feature maps of different scales to the same size, obtaining feature maps of the same size; each Stage is associated with a learnable scaling factor, multiplying the scaling factor by the corresponding feature map, and the merging layer is used to merge the multiplied feature maps in the channel dimension to obtain the finally fused feature map. The globally averaged pooled fused feature map is input into the classification head to generate a classification result.
[0073] In the embodiment of the present invention, the multi-scale feature fusion module M2SF with a learnable scaling factor uses the channel self-adaptation of the SE layer to enhance the feature expression ability of each stage, scales the enhanced features to adapt to data of different scales. If the target is small, the scaling factors corresponding to the lower-level stages (such as Stage0 and Stage1) are increased to obtain more attention. If the target is large, the scaling factors corresponding to the higher-level stages (such as Stage2 and Stage3) can be increased.
[0074] The network with a pyramid structure constructed in the embodiment of the present invention is shown in Table 1.
[0075] Table 1 Network with a pyramid structure constructed in the embodiment of the present invention
[0076]
[0077]
[0078] In Table 1, E and D respectively represent the scale and the number of channels of the feature map.
[0079] Preferably, the present invention uses the method of transfer learning to fine-tune the pre-trained model on ImageNet1K on the defect classification dataset. Since the scale of the metal surface defect classification dataset is small, in order to enhance the performance of the model, the model is pre-trained on the image classification dataset ImageNet1K in advance, and the learned knowledge is transferred to the metal surface defect classification task. Since some general feature representations have been learned on the large-scale dataset, general low-level and middle-level features can be extracted. Therefore, when fine-tuning on the metal surface defect dataset, the model can converge quickly, reducing a large amount of training time and computing resources, and can also effectively avoid problems such as overfitting and difficult convergence caused by the too small scale of the task dataset.
[0080] The ImageNet1K dataset contains 1.28M training images and 50K validation images from 1000 categories. The training settings mainly follow DeiT. Specifically, random cropping, random horizontal flipping, label smoothing regularization, mixing, and random erasing are applied as data augmentation. When training the input image with a pixel size of 224×224, the present invention sets the momentum to 0.9, the total batch size to 512, and the weight decay to 0.05 to optimize the model. The learning rate uses cosine decay, 3×10 -4The initial learning rate and EMA are trained for 300 epochs (due to computational resource limitations, the total batch size is reduced from 1024 in DeiT to 512, so the initial learning rate is also reduced accordingly). The training is carried out on 4 RTX 4090 GPUs, and the best model weights after training are saved to a.pth file.
[0081] For fine-tuning training, it is necessary to first load the saved pre-trained model, and then change the linear output layer of the network, changing the number of output channels from 1000 classes in ImageNet1K to 9 classes of metal surface defects. The metal surface defect dataset is input for fine-tuning training. Since the metal surface defect data is quite different from the ImageNet1K data, no neural layers are frozen. The mixed metal surface defect dataset first undergoes data augmentation operations (random cropping, flipping, rotation) and normalization, and is resized to a unified size of 224×224. Since it is a grayscale image, the grayscale values of the RGB three color channels are the same. During training, the AdamW optimizer is used, and the learning rate is set to a fixed value of 1×10 -4 , the batch size is set to 8, the number of epochs is set to 25, and the fine-tuning training is carried out on a single RTX 4090 GPU. The training data input to the model is a matrix with the shape of 8×224×224×3, and the output is an 8×9 vector. Each vector corresponds to 1 type of defect, and the vector with the largest output value is the final classification result.
[0082] In the embodiments of the present invention, the deep neural network has 24 neural layers with decreasing widths to achieve the above-mentioned 9 defect classifications. Among them, the Patch layer consists of 2 convolutional layers of 3×3. The stride of the first convolutional layer is 2, and the original 3-channel image is output into a feature map with 48 channels and halved size. The second convolutional layer does not change the number of channels and size of the feature map, further enhancing the network's expressive ability. At this time, the tensor shape changes to 8×112×112×48. Stage0 has 2 MBConv layers. The first MBConv layer contains a max-pooling layer for downsampling. After passing through the network layers of Stage0, the tensor shape changes to 8×56×56×48; Stage1 has 6 MBConv layers. The first MBConv layer also contains a downsampling layer. After passing through the network layers of Stage1, the tensor shape changes to 8×28×28×96; Stage2 has 14 bidirectional Mamba layers. An additional 3×3 convolutional downsampling layer is added before the first bidirectional Mamba layer. After tensor downsampling, it becomes 8×14×14×192. To adapt to the one-dimensional data input format of the Mamba layer, the tensor is flattened from two-dimensional to one-dimensional 8×196×192 and layer normalization is performed. After passing through 14 bidirectional Mamba layers, the shape remains unchanged; Stage3 has 2 bidirectional Mamba layers, restoring the one-dimensional data shape to two-dimensional format. After downsampling, the tensor shape is 8×7×7×384, and then flattened to 8×49×384. After passing through 2 bidirectional Mamba layers, the shape remains unchanged. The outputs of Stage2 and Stage3 are converted into two-dimensional forms of feature maps, and are simultaneously processed by the SE layer and normalized by BatchNorm with the feature maps output by Stage0 and Stage1, and then average pooling is performed to obtain 4 feature maps with the same size but different numbers of channels. Each feature map is multiplied by the corresponding scaling factor, and they are fused in the channel dimension. The resulting tensor shape is 8×7×7×720, flattened to 8×49×720, and after layer normalization, average pooling is performed in the second dimension to obtain a vector of shape 8×720. After passing through the output linear layer, a vector of 8×9 is output.
[0083] The performance of the classification model is evaluated using the validation set. The data in the validation set is only adjusted in size and normalized using the mean and standard deviation of the training set to ensure stable data distribution during validation and consistent with the actual usage scenario. To visually display the classification effect, the present invention uses indicators such as accuracy, precision, recall, and F1-score as evaluation criteria, which are expressed by the formula:
[0084]
[0085]
[0086] Among them, TP, FP, FN, and TN represent the numbers of true positives, false positives, false negatives, and true negatives, respectively.
[0087] The indicators of the metal surface defect classification model constructed in the embodiments of the present invention are shown in Table 2, where support represents the number of defect samples of this type participating in the test.
[0088] Table 2 Indicators of the 9-class metal surface defect classification model
[0089]
[0090] It can be observed that the models provided in the embodiments of the present invention have good effects in classifying 9 types of defects, and the overall accuracy rate reaches 0.983228. Comparing the performance of the model proposed by the present invention with that of Resnet18 and ViT, the model of the present invention also achieves the highest classification accuracy rate. Since the Rp type of defect is the most difficult type to distinguish for each model, the 3 classification indicators of Rp are also compared separately, as shown in Table 3.
[0091] Table 3 Comparison of the classification performance of various models
[0092]
[0093]
[0094] Analysis shows that the model of the present invention has good effects in all 3 indicators. When the M2SF module is removed, the recall rate indicator is the best. Although the accuracy rate of the model is not as good as that of Resnet18 and ViT, the computational efficiency of the model is relatively high; after adding the M2SF module, the model of the present invention achieves the best effect, proving the improvement of the M2SF proposed by the present invention on the model generalization ability and classification ability.
[0095] In addition, Figure 5 - Figure 8 The confusion matrices of the classification of four models, namely Resnet18, ViT-B, the model of the present invention without the M2SF module, and the model of the present invention with the M2SF module, are respectively drawn. Each column of the confusion matrix represents the predicted category, the total number of each column represents the number of data predicted as this category, each row represents the true attribution category of the data, the total number of data in each row represents the number of data instances of this category, and the value in each column represents the number of true data predicted as this category. The darker the color on the main diagonal, the better the classification effect. It can be seen that the model of the present invention can obtain a good classification effect.
[0096] The present invention focuses on the defect classification task and proposes a novel network architecture to improve the classification performance and provide reliable support for the complete defect detection process. Aiming at the limitations of CNN and Transformer, the present invention combines CNN with the structured state space model Mamba that has linear computational complexity and long sequence modeling ability to construct a model structure combining deep convolution and bidirectional Mamba modules. The deep convolution is used to capture the local spatial features of the image data in parallel in the spatial dimension, and then the upper limit and capacity of the model are improved through the long sequence modeling ability of Mamba. The design of the bidirectional Mamba is to reduce the adverse impact on the global understanding caused by the unidirectional causal dependence of the RNN-like structure of Mamba. This model structure pays attention to both local and global features, greatly improving the accuracy of the classification model while taking into account the computational efficiency. During the construction of the network, the network is divided into four parts, each part having different feature scales and decreasing gradually. M2SF is used to fuse these features along the channel dimension, making full use of the feature information of each scale and assigning adaptive weights to this information, which can alleviate the impact of insufficient defect samples to a certain extent and enable the network to be robustly trained on multiple datasets with different sizes at the same time, enhancing the generalization ability of the model.
[0097] Embodiment 2
[0098] An embodiment of the present invention provides a method for classifying metal surface defects, including:
[0099] Input the metal surface defect image to be classified into the metal surface defect classification model to obtain the defect classification result. Among them, the metal surface defect classification model is constructed by the method for constructing a metal surface defect classification model based on a structured state space model in Embodiment 1.
[0100] For related technical solutions, refer to the description in Embodiment 1 and will not be elaborated here.
[0101] Embodiment 3
[0102] An embodiment of the present invention provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method for constructing a metal surface defect classification model based on a structured state space model in Embodiment 1 or the method for classifying metal surface defects in Embodiment 2.
[0103] The electronic device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The so-called processor can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application-Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory.
[0104] The related technical solutions are the same as above and will not be elaborated here.
[0105] Embodiment 3
[0106] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for constructing a metal surface defect classification model based on a structured state space model in Embodiment 1 or the metal surface defect classification method in Embodiment 2 are realized.
[0107] Specifically, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0108] The related technical solutions are the same as above and will not be elaborated here.
[0109] Embodiment 4
[0110] An embodiment of the present application provides a computer program product, including a computer program. When the computer program runs on a computer, the computer is caused to execute the steps of the method for constructing a metal surface defect classification model based on a structured state space model in Embodiment 1 or the metal surface defect classification method in Embodiment 2.
[0111] The related technical solutions are the same as above and will not be elaborated here.
[0112] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for constructing a metal surface defect classification model based on a structured state space model, characterized in that: include: The image data set is used to pre-train a neural network based on a structured state space model, and the metal surface defect data set is used to fine-tune the pre-trained neural network to obtain a metal surface defect classification model; wherein the neural network based on the structured state space model includes: Patch layer, used to downsample the input image samples and increase the number of channels; A multi-scale feature extraction module, comprising deep convolution layers and bidirectional Mamba layers of different scales; the deep convolution layers are used to capture local spatial feature maps of different scales of the Patch layer output image in parallel in the spatial dimension; the bidirectional Mamba layer comprises two Mamba structures after removing the fully connected layer and a fully connected layer, one Mamba structure is used to obtain a forward output feature based on one-dimensional data after downsampling and flattening the local spatial feature map, and the other Mamba structure is used to obtain a reverse output feature based on data after flipping the one-dimensional data; the fully connected layer is used to linearize the features after the forward output features and the reverse output features are fused to obtain the feature map output by the bidirectional Mamba layer; A multi-scale feature fusion module, used for performing multi-scale fusion on the feature map output by the deep convolution layer and the feature map output by the bidirectional Mamba layer to obtain a fused feature map; A classification head is used to generate a classification result based on the fused feature map.
2. The method for constructing a metal surface defect classification model according to claim 1, characterized in that: The deep convolution layer is an MBConv layer, and the multi-scale feature extraction module includes four stages connected in sequence in a pyramid structure: Stage0-Stage3, the scales of the feature maps output by Stage0-Stage3 gradually decrease, and the corresponding number of feature channels doubles; Among them, Stage0 includes 2 MBConv layers connected in sequence, Stage1 includes 6 MBConv layers connected in sequence, Stage2 includes 14 bidirectional Mamba layers connected in sequence, and Stage3 includes 2 bidirectional Mamba layers connected in sequence; The multi-scale feature fusion module is used to fuse the feature maps output by Stage0-Stage3 to obtain a fused feature map.
3. The method for constructing a metal surface defect classification model according to claim 1 or 2, characterized in that: The positive output feature y forward for: y forward =x proj ⊙SiLU(z) x proj =SSM(SiLU(Conv1d(x))) The reverse output feature y backward for: Wherein, SSM(·) represents the output of the main path of the Mamba structure after removing the fully connected layer, SiLU(·) is the SiLU activation function, Conv1d(·) represents a one-dimensional convolution operation, ⊙ represents a dot product operation; x and z represent the input of the main path and the input of the residual path in the Mamba structure obtained by mapping the one-dimensional data through the linear layer, respectively; x b 、z b The data after flipping the one-dimensional data are respectively mapped through a linear layer to obtain the input of the main path and the input of the residual path in the Mamba structure.
4. The method for constructing a metal surface defect classification model according to claim 3, characterized in that: The multi-scale feature fusion module is a multi-scale feature fusion module M2SF with a learnable scaling factor, specifically comprising: an SE layer, a normalization layer, an average pooling layer, a scaling factor learned through training, and a merging layer; The SE layer is used to perform channel adaptive adjustment on the feature map output by each Stage. The adjusted feature map is processed by the corresponding normalization layer and the average pooling layer in sequence to obtain a feature map of the same size. After the feature map is multiplied by the corresponding scaling factor, it is merged in the channel dimension by the merging layer to obtain the fused feature map.
5. The method for constructing a metal surface defect classification model according to claim 1 or 4, characterized in that: The metal surface defect dataset is a dataset obtained by merging metal surface defect datasets of multiple different sizes.
6. The method for constructing a metal surface defect classification model according to claim 5, characterized in that: Based on the metal surface defect dataset, the pre-trained neural network was fine-tuned using the transfer learning method to obtain a metal surface defect classification model.
7. A method for classifying metal surface defects, characterized in that: include: The metal surface defect image to be classified is input into a metal surface defect classification model to obtain a defect classification result; wherein the metal surface defect classification model is constructed by the metal surface defect classification model construction method described in any one of claims 1 to 6.
8. An electronic device, characterized in that: comprising a computer readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium to execute the metal surface defect classification model construction method described in any one of claims 1 to 6, or to execute the metal surface defect classification method described in claim 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for constructing a metal surface defect classification model as described in any one of claims 1 to 6 is implemented, or the method for classifying metal surface defects as described in claim 7 is implemented.
10. A computer program product, characterized in that It includes a computer program, which, when running on a computer, enables the computer to execute the metal surface defect classification model construction method described in any one of claims 1 to 6, or execute the metal surface defect classification method described in claim 7.
Citation Information
Patent Citations
Polycrystalline photovoltaic cell defect identification method based on attention mechanism and multi-scale feature fusion
CN117876339A
Visual representation method and device based on bidirectional state space model
CN117876845A
SSVEP (Steady-State Visual Evoked Potential) signal classification method based on Mama model
CN118760946A
Adaptive voiceprint recognition system based on Mama model
CN119274560A
Cited By
Medical image classification method and system of structure perception state space model
CN121505366A