Steel flaw detection method based on Mamba network
Through the hierarchical feature aggregation and dual attention mechanism of the Mamba network, the problems of low efficiency and insufficient accuracy in steel defect detection are solved, efficient and accurate defect detection are achieved, and the stability and production continuity of the system are improved.
Patent Information
- Application Number
- CN202510378285.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art has low efficiency, missed inspection, difficulty in dealing with complex surface textures and various defect types, high computational complexity, strict hardware requirements, and inability to meet the needs of fast, accurate and efficient inspection.
The steel defect detection method based on Mamba network is adopted, including image data input, encoder module, jump connection module, decoder module and system monitoring and management module. Through hierarchical feature aggregation, parallel Mamba module, dual attention mechanism and real-time system monitoring, coordinated extraction of deep and shallow features and accurate capture of defect features are achieved.
It significantly improves the accuracy of defect detection, reduces missed and missed detection, improves the stability and reliability of the system, ensures production continuity, and reduces costs.
Smart Images

Figure CN120298358A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and image processing, and in particular to a method for applying an advanced Mamba network to steel defect detection. Background Art
[0002] In steel production, defect detection is crucial for ensuring product quality and relies on computer vision and image processing technologies. However, with the expansion of production scale and the improvement of quality requirements, the problems of traditional detection methods have become prominent.
[0003] Early manual detection was inefficient and greatly affected by subjective factors. Traditional image processing algorithms were difficult to accurately extract defect features in the face of complex steel surface textures, lighting changes, and diverse defect types, and missed detections and false detections were common. Convolutional neural networks (CNNs) had certain achievements in steel defect detection, but problems such as gradient disappearance and overfitting occurred when the network was deepened, resulting in poor generalization ability. At the same time, CNNs were weak in processing long-range dependence relationships and had poor detection effects on large-area or discrete defects. The Transformer model based on the attention mechanism could better capture long-range dependence relationships, but had a high computational complexity and demanding hardware performance requirements, and was limited in application in actual steel production scenarios. In addition, new network architectures had deficiencies in balancing shallow and deep feature mining and multi-scale feature processing. The shapes and sizes of steel defects varied greatly, and the existing technologies could not meet the requirements of steel production for fast, accurate, and efficient defect detection.
[0004] Therefore, developing new steel defect detection methods is of great significance, which helps to improve product quality, reduce costs, and enhance the competitiveness of enterprises. Summary of the Invention
[0005] The purpose of the present invention is to provide a steel defect detection method based on the Mamba network to solve the problems raised in the above background art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A steel defect detection method based on the Mamba network, characterized by comprising:
[0007] An image data input module for reading image data from multiple data sources and performing preprocessing, where the data sources include local storage media, network cameras, remote servers based on network protocols, and databases; the preprocessing includes image format conversion, size adjustment, pixel value normalization, and single-channel image channel expansion;
[0008] An encoder module, whose hierarchical architecture contains 6 layers. The first 3 layers use a hierarchical feature aggregation module (HFA) to mine shallow features, and the last 3 layers use a parallel Mamba module (PMLayer) to extract deep features;
[0009] The skip connection module is used to perform average pooling and max pooling operations on the multi-scale features of each layer of the encoder respectively. After concatenating the operation results along the channel dimension, the concatenated results are input into the shared_conv2d convolutional layer to generate a spatial attention map. Then, channel attention processing is performed through the ECA module. Finally, the processed features are added to the original features to output the fused features.
[0010] The decoder module is used to restore and fuse the fused features output by the skip connection module, and gradually restore the resolution of the image through upsampling operations, and output image features that meet the actual task requirements.
[0011] The result output module is used to convert the image features output by the decoder into the results of the actual task, and perform output and storage.
[0012] The system monitoring and management module is used to monitor the running status of the entire system in real time, adjust parameters, and handle faults.
[0013] Preferably, when the image data input module reads images from the local storage medium, it reads directly according to the file path; when obtaining images from a network camera, it captures in real time by calling the camera driver and related interfaces; when reading images from a network server, it sends requests and receives data following the network protocol.
[0014] Preferably, in the hierarchical feature aggregation module (HFA), after the image features are input, they are sequentially processed through nn.BatchNorm2d normalization, HierarchicalFeatureAggregation module processing, DropPath module to randomly discard some features, and again through nn.BatchNorm2d normalization. Finally, after being processed by the ChannelAggregationFFN module, features rich in shallow information are output. Among them, in the HierarchicalFeatureAggregation module, proj_1 reshapes the feature dimension through 1x1 convolution, the gating branch generates gating weights using the gate convolution, the aggregation branch uses MultiOrderDWConv to capture context information, and the outputs of the two branches are multiplied after being processed by the activation function and then added to the cloned value of the input features to achieve residual connection; the ChannelAggregationFFN module first increases the dimension, then goes through depth convolution, activation function, dropout operation, feature decomposition and correction, and finally reduces the dimension and makes a residual connection with the original input features.
[0015] Preferably, in the parallel Mamba module (PMLayer), the input features are first flattened and transposed, then normalized by nn.LayerNorm, and then evenly split into 4 branches and processed through the Mamba module respectively. The processing results of each branch are added to the original branch features multiplied by the learnable skip connection scaling factor skip_scale. After the 4 branches are processed, they are concatenated along the channel dimension, normalized by nn.LayerNorm again, projected to the specified output dimension through the proj fully connected layer, and finally transposed and reshaped into the spatial structure of the original image to output the deep features.
[0016] Preferably, in the skip connection module, average pooling is used to obtain the global average information of the features, and max pooling is used to highlight the important local information in the features; the ECA module first performs global average pooling on the features, then models the local channel relationship through 1D convolution, and finally uses the Sigmoid function to output the channel attention weights and multiply them element-wise with the input features to complete channel attention weighting.
[0017] Preferably, the first 3 layers of the decoder module use conventional convolution operations. After processing the input features using PMLayer, they are added to the fused features of the corresponding layers in the skip connection, and then upsampled through operations such as nn.Conv2d and F.interpolate. After multiple upsampling and convolution operations, the image feature representation that meets the target size is output.
[0018] Preferably, in the image segmentation task, the result output module maps the image features to a segmentation mask; in the object detection task, the image features are converted into the bounding box coordinates and class information of the objects; the output results can be stored in a local storage device, sent to a remote server through a network protocol, or directly displayed on a display device.
[0019] Preferably, the system monitoring and management module monitors the resource usage of each module of the system and the performance metrics of model training and inference in real time; dynamically adjusts the key parameters of the system according to the monitoring data and actual task requirements; has the ability to detect and repair faults, and takes corresponding repair measures and records the fault information when a fault is detected.
[0020] A steel defect detection method based on the Mamba network, characterized by comprising the following steps:
[0021] S1. The image data input module reads image data from a specified data source, performs preprocessing, and passes the processed image data to the encoder module;
[0022] S2. The encoder module processes the image data layer by layer. The first 3 layers of the HFA module extract shallow features, and the last 3 layers of the PMLayer module extract deep features, and pass the extracted features to the skip connection module;
[0023] S3, the skip connection module fuses the multi-scale features of each layer of the encoder, and after spatial attention and channel attention processing, passes the fused features to the decoder module;
[0024] S4, the decoder module uses PMLayer to optimize and adjust the input features, adds them to the fusion features of the corresponding layer in the jump connection, restores the image resolution through upsampling and convolution operations, and outputs the processed image features to the result output module;
[0025] S5, the result output module converts the image features into actual task results and stores, transmits or displays them according to the requirements;
[0026] S6. The system monitoring and management module monitors the system operation status in real time, adjusts parameters and handles faults based on the monitoring data to ensure stable operation of the system.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] 1. The encoder module of the present invention adopts a unique layered design. The hierarchical feature aggregation module (HFA) of the first three layers can deeply mine shallow features and carefully capture the subtle texture and preliminary feature information of the steel surface; the parallel Mamba module (PMLayer) of the last three layers accurately extracts deep semantics, effectively captures long-distance dependencies, and fully grasps the global features of steel images. Compared with traditional methods, this method of collaborative extraction of deep and shallow features can more comprehensively and accurately obtain features related to steel defects, provide a solid data foundation for subsequent detection, and greatly improve the accuracy of defect detection.
[0029] 2. The jump connection module innovatively introduces a dual attention mechanism. Through average pooling and maximum pooling operations, it obtains the global average information and important local information of the features respectively, and fuses the two to generate a spatial attention map. At the same time, the ECA module is used to perform channel attention processing to automatically adjust the degree of attention to different channel features. This feature fusion method optimized by the dual attention mechanism can enable the model to more keenly capture the detailed information of the defects in the steel defect detection task, significantly improve the ability to identify defects, and effectively avoid missed detection and false detection.
[0030] 3. The system monitoring and management module monitors the resource usage of each module in the system in real time, including CPU, GPU usage rates, memory occupancy, etc. At the same time, it closely monitors the performance metrics of model training and inference, such as loss values, accuracy rates, response times, etc. Based on the monitoring data, this module can dynamically adjust the key parameters of the system, such as optimizing the learning rate, adjusting convolution kernel parameters, etc., to ensure that the model always maintains the best performance state. In addition, when software or hardware failures occur in the system, it can respond quickly and take repair measures, while recording the failure information in detail, providing a strong basis for subsequent system optimization, greatly ensuring the continuous and stable operation of the system, reducing production interruptions and data losses caused by failures, and improving the reliability and practicality of the entire detection system. Brief Description of the Drawings
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the structures shown in these drawings.
[0032] Figure 1 It is a flowchart of a steel defect detection method based on the Mamba network of the present invention;
[0033] Figure 2 It is an overall framework diagram of a steel defect detection method based on the Mamba network of the present invention;
[0034] Figure 3 It is a framework diagram of the HFA module in a steel defect detection method based on the Mamba network of the present invention;
[0035] Figure 4 It is a framework diagram of the PM layer in a steel defect detection method based on the Mamba network of the present invention;
[0036] The implementation, functional features, and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. Detailed Embodiments
[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0038] The present invention presents a technical solution, a steel defect detection method based on the Mamba network, mainly including the following modules: an image data input module, an encoder module, a skip connection module, a decoder module, a result output module, and a system monitoring and management module.
[0039] As a preferred solution of the present invention, in the image data input module, this module undertakes the important responsibility of acquiring and preprocessing image data. In terms of data acquisition, it has the ability to read images from a variety of data sources, including local storage media (such as hard disks, solid-state drives), network cameras, remote servers based on network protocols (such as HTTP, FTP), and databases (such as MySQL, MongoDB, etc., provided that the database stores image data or image data path information). For different data sources, the module adopts corresponding reading strategies. For example, when reading from the local file system, it directly reads according to the file path; when obtaining images from a network camera, it captures images in real time by calling the camera driver and related interfaces; when reading from a network server, it sends requests and receives data following the network protocol.
[0040] In the image preprocessing stage, the module first performs format conversion and size adjustment on the read images. It supports the conversion of common image formats such as JPEG, PNG, and BMP, and uniformly converts the images into a format suitable for model processing. At the same time, according to the model input requirements, the images are scaled to a specified size. For example, images of any size are scaled to an H×W resolution, and algorithms such as bilinear interpolation and bicubic interpolation are used to ensure the quality of the scaled images. In addition, the pixel values of the image data are normalized, mapping them to the [-1,1] or [0,1] interval, making the data distribution more in line with the model training requirements and accelerating the model convergence speed. For single-channel images, the module will perform channel expansion operations to expand them into three-channel images to ensure consistency with the model input dimension.
[0041] As a preferred solution of the present invention, in the encoder module, the encoder hierarchical architecture includes 6 layers, and the number of channels in each layer is 8, 16, 24, 32, 48, and 64 respectively. The first 3 layers use the hierarchical feature aggregation module (HFA) to mine shallow features, and the last 3 layers use the parallel Mamba module (PMLayer) to extract deep features.
[0042] In the Hierarchical Feature Aggregation (HFA) module, after the image features are input into the HFA module, they first undergo nn.BatchNorm2d normalization. This operation normalizes based on the mean and variance of each channel's data, making the data distribution more stable and effectively accelerating model convergence. Then, the processed features enter the Hierarchical Feature Aggregation module, where proj_1 reshapes the feature dimensions through 1x1 convolutional projection to prepare for subsequent processing. The gating branch uses gate convolution to generate gating weights, which determine the screening and circulation of information; the aggregation branch adopts MultiOrderDWConv, and through depth convolutions with different dilation rates (1, 2, 3), captures context information from different scales to enhance the feature representation ability. The outputs of the two branches are multiplied after being processed by activation functions such as SiLU, and then added to the cloned value of the input features to achieve a residual connection, retaining the key information in the original features. Subsequently, the DropPath module randomly discards some features according to the set probability (such as drop_path_rate), effectively preventing model overfitting and enhancing the model's generalization ability. When drop_path_rate is 0, the module directly outputs the input features. After random depth dropout, the features are normalized again by nn.BatchNorm2d to further stabilize the data distribution. Then, they enter the Channel Aggregation FFN module, which first uses fc1 for 1x1 convolutional upsampling to broaden the feature representation space; dwconv depth convolution is used to capture local context information; the GELU activation function introduces non-linearity to enhance the model's representation ability; the dropout operation further prevents overfitting; feature decomposition and correction are performed through decompose convolution and decompose_act activation function to mine global information; finally, fc2 performs 1x1 convolutional downsampling to restore to the original number of channels and makes a residual connection with the original input features again, outputting features rich in shallow information.
[0043] In the parallel Mamba module (PMLayer), the features input to the PMLayer module are first flattened and transposed to convert them into a format suitable for nn.LayerNorm processing. nn.LayerNorm normalizes the features, adjusts the mean and variance of the data, and ensures the stability of model training. The normalized features are evenly split into 4 branches, and each branch is independently processed through the Mamba module. The Mamba module can effectively capture the dependencies in long sequence data, and its ability to capture remote spatial information is particularly prominent when processing image features. The processing result of each branch is added to the original branch features multiplied by the learnable skip connection scaling factor skip_scale to complete the residual connection, enhancing the stability and representational ability of the features. After the 4 branches are processed, their features are concatenated along the channel dimension to restore the original number of channels. The concatenated features are normalized again by nn.LayerNorm to ensure the stability of the data, and then projected to the specified output dimension through the proj fully connected layer. Finally, the projected features are transposed and reshaped into the spatial structure of the original image, and features containing rich deep information are output.
[0044] As a preferred solution of the present invention, in the skip connection module (SC_Att_Bridge), the module first performs average pooling and max pooling operations on the multi-scale features of each layer of the encoder respectively. Average pooling calculates the average value of the feature map in the spatial dimension to obtain the global average information of the features, reflecting the overall feature trend of the image; max pooling selects the maximum value of the feature map in the spatial dimension to highlight the important local information in the features. The results of average pooling and max pooling are concatenated along the channel dimension and then input into the shared_conv2d convolutional layer composed of a 7x7 convolutional layer and a Sigmoid activation function. The convolutional layer performs a convolutional operation on the concatenated features to fuse the information of average pooling and max pooling, generating a spatial attention map. The spatial attention map is multiplied element-wise with the original features to achieve spatial attention adjustment, making the model pay more attention to the important regions in the image.
[0045] The features adjusted by spatial attention enter the ECA module for channel attention processing. The ECA module first performs global average pooling on the features to compress the features in the spatial dimension and obtain the global statistical information of each channel. Then, it models the local channel relationships through 1D convolution to explore the local correlations between channels. The kernel size of the 1D convolution can be set according to the actual situation (e.g., the default value is 3), and the weight relationships between channels are learned through the convolution operation. Finally, the Sigmoid function is used to output the channel attention weights, which reflect the importance of each channel in the current image features. The channel attention weights are multiplied element-wise with the input features to complete the channel attention weighting, enabling the model to automatically adjust the attention degree to different channel features, highlight the features of important channels, and suppress the interference of irrelevant channels.
[0046] After completing the spatial and channel attention processing, the features weighted by channel attention are added to the original features to output the fused features. This fusion method fully utilizes the original feature information and the feature information processed by the attention mechanism, provides richer and more accurate context information for the decoder, and helps to improve the overall performance of the model.
[0047] As a preferred solution of the present invention, in the decoder module, the fused features are restored and fused, and the resolution of the image is gradually restored through upsampling operations to output the image features that meet the actual task requirements. The first 3 layers of the decoder use conventional convolution operations to achieve the restoration and fusion of features. First, the input features are processed using the PMLayer, and the processing process is the same as that of the PMLayer in the encoder, including normalization, branch processing, branch merging, secondary normalization and projection, shape restoration, etc. Through these steps, the deep features are further optimized and adjusted to make them more suitable for subsequent feature restoration and fusion operations.
[0048] The processed features are added to the fused features of the corresponding layer in the skip connection to achieve the fusion and restoration of features. This fusion method can combine the feature information of different levels of the encoder, utilize the rich context information transmitted by the skip connection, and gradually restore the details and resolution of the image. For example, when processing an image segmentation task, the fused features can more accurately reflect the boundary and region information of different objects.
[0049] After feature fusion and restoration, the decoder performs upsampling through operations such as nn.Conv2d and F.interpolate. For example, bilinear interpolation upsampling using F.interpolate doubles the size of the feature map. During the upscaling process, convolutional operations are combined to further fuse and optimize the features. Convolutional operations can refine and adjust the features while increasing the size of the feature map, enabling the features to more accurately reflect the detailed information of the image. Through multiple such upsampling and convolutional operations, an image feature representation that matches the target size is finally output, providing high-quality feature data support for subsequent image-related tasks.
[0050] As a preferred embodiment of the present invention, in the result output module, the result output module is responsible for converting the image features output by the decoder into the results of actual tasks, and performing output and storage. In an image segmentation task, the module maps the image features to a segmentation mask, and through threshold segmentation, clustering algorithms, or the output activation function of a deep learning model (such as Softmax), the features are converted into pixel-level annotations of different object classes. In an object detection task, the module converts the image features into the bounding box coordinates and class information of the objects, uses regression algorithms to predict the position and size of the bounding boxes, and determines the class of the objects through classification algorithms.
[0051] The output results can be processed in various ways according to actual needs. For local application scenarios, the results can be stored in local storage devices (such as hard disks, solid-state drives), and common file formats (such as JSON, XML for storing object detection results, PNG, JPEG for storing segmentation mask images) are used for saving. For network application scenarios, the results can be sent to a remote server through network protocols (such as HTTP, WebSocket) for use by other systems or users. At the same time, the results can also be directly displayed on local or remote display devices for convenient viewing and analysis by users.
[0052] As a preferred embodiment of the present invention, in the system monitoring and management module, the system monitoring and management module is responsible for real-time monitoring, parameter adjustment, and fault handling of the operating status of the entire system to ensure the stable and efficient operation of the system. In terms of monitoring, the module monitors the resource usage of each module of the system in real time, including CPU usage, GPU usage, memory occupancy, etc., and obtains this information through system call interfaces provided by the operating system or third-party monitoring tools (such as NVIDIA-SMI for monitoring GPU status). At the same time, performance metrics such as loss values and accuracies during training, and response times and throughputs during inference of the model are monitored.
[0053] In terms of parameter adjustment, the module dynamically adjusts the key parameters of the system according to the monitoring data and actual task requirements. For example, during the model training process, if it is found that the model convergence speed is too slow, the learning rate can be appropriately adjusted; if overfitting occurs, the regularization parameter can be adjusted or the data augmentation strategy can be increased. For parameters such as the convolution kernel size and stride in the encoder and decoder, they can also be optimized and adjusted according to the characteristics of the image data and task requirements.
[0054] In terms of fault handling, the module has the ability to detect and repair faults. By real-time monitoring of the system operation status, potential faults can be detected in a timely manner, such as hardware faults (such as GPU overheating, memory faults), software faults (such as model training crashes, inference errors). Once a fault is detected, the module will take corresponding repair measures, such as restarting the faulty module, switching to standby hardware devices, reloading the model, etc. At the same time, the system will also record the fault information, including the time, type, and cause of the fault occurrence, etc., for subsequent fault analysis and system optimization.
[0055] As a preferred solution of the present invention, the working process of the steel defect detection method based on the Mamba network is as follows:
[0056] S1. The image data input module reads the image data from the specified data source and performs preprocessing, and transfers the processed image data to the encoder module.
[0057] S2. The encoder module processes the image data layer by layer. The first 3 layers of HFA modules extract shallow features, and the last 3 layers of PMLayer modules extract deep features, and transfer the extracted features to the skip connection module.
[0058] S3. The skip connection module fuses the multi-scale features of each layer of the encoder. After spatial attention and channel attention processing, it transfers the fused features to the decoder module.
[0059] S4. The decoder module uses PMLayer to optimize and adjust the input features, adds them to the fused features of the corresponding layers in the skip connection, restores the image resolution through upsampling and convolution operations, and outputs the processed image features to the result output module.
[0060] S5. The result output module converts the image features into actual task results and stores, transmits, or displays them according to the requirements.
[0061] S6. The system monitoring and management module monitors the system operation status in real time, performs parameter adjustment and fault handling according to the monitoring data, and ensures the stable operation of the system.
[0062] The above embodiments are only used to illustrate the technical concept and characteristics of the present invention. The purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly, and it should not be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A steel defect detection method based on the Mamba network, characterized in that It includes: An image data input module, which is used to read image data from multiple data sources and perform preprocessing. The data sources include local storage media, webcameras, remote servers based on network protocols, and databases; the preprocessing includes image format conversion, size adjustment, pixel value normalization, and single-channel image channel expansion; An encoder module, whose hierarchical architecture contains 6 layers. The first 3 layers use a hierarchical feature aggregation module (HFA) to mine shallow features, and the last 3 layers use a parallel Mamba module (PMLayer) to extract deep features; A skip connection module, which is used to perform average pooling and max pooling operations on the multi-scale features of each layer of the encoder respectively, splice the operation results along the channel dimension and input them into a shared_conv2d convolutional layer to generate a spatial attention map, then perform channel attention processing through an ECA module, and finally add the processed features to the original features to output the fused features; A decoder module, which is used to restore and fuse the fused features output by the skip connection module, and gradually restore the resolution of the image through upsampling operations, and output image features that meet the actual task requirements; A result output module, which is used to convert the image features output by the decoder into the results of actual tasks, and perform output and storage; A system monitoring and management module, which is used to monitor the running state of the entire system in real time, adjust parameters, and handle faults.
2. The steel flaw detection method based on the Mamba network according to claim 1, wherein When the image data input module reads images from the local storage medium, it reads directly according to the file path; when it obtains images from a webcamera, it captures them in real time by calling the camera driver and related interfaces; When reading images from a network server, it sends requests and receives data according to the network protocol.
3. The method for detecting steel defects based on the Mamba network according to claim 1, characterized in that, In the hierarchical feature aggregation module (HFA), after the image features are input, they are sequentially processed by nn.BatchNorm2d normalization, HierarchicalFeatureAggregation module, DropPath module to randomly discard some features, and then nn.BatchNorm2d normalization again, and finally output features rich in shallow information after being processed by the ChannelAggregationFFN module; among them, in the HierarchicalFeatureAggregation module, proj_1 reshapes the feature dimension through 1x1 convolution, the gating branch generates gating weights using the gate convolution, the aggregation branch uses MultiOrderDWConv to capture context information, and the outputs of the two branches are multiplied after being processed by the activation function and then added to the cloned value of the input features to achieve a residual connection; the ChannelAggregationFFN module first increases the dimension, then goes through depth convolution, activation function, dropout operation, feature decomposition and correction, and finally reduces the dimension and makes a residual connection with the original input features.
4. The steel flaw detection method based on the Mamba network according to claim 1, wherein, In the parallel Mamba module (PMLayer), the input features are first flattened and transposed, then normalized by nn.LayerNorm, and then evenly split into 4 branches to be processed by the Mamba module respectively. The processed results of each branch are added to the original branch features multiplied by the learnable skip connection scaling factor skip_scale. After the 4 branches are processed, they are concatenated along the channel dimension, normalized again by nn.LayerNorm, projected to the specified output dimension through the proj fully connected layer, and finally transposed and reshaped into the spatial structure of the original image to output the deep features.
5. The steel flaw detection method based on the Mamba network according to claim 1, wherein In the skip connection module, average pooling is used to obtain the global average information of the features, and max pooling is used to highlight the important local information in the features; the ECA module first performs global average pooling on the features, then models the local channel relationship through 1D convolution, and finally uses the Sigmoid function to output the channel attention weights and multiply them element-wise with the input features to complete the channel attention weighting.
6. The steel defect detection method based on the Mamba network according to claim 1, characterized in that The first 3 layers of the decoder module use conventional convolution operations. After processing the input features using PMLayer, they are added to the fused features of the corresponding layers in the skip connection, and then upsampled through operations such as nn.Conv2d and F.interpolate. After multiple upsampling and convolution operations, the image feature representation that meets the target size is output.
7. The steel flaw detection method based on the Mamba network according to claim 1, characterized in that In the image segmentation task, the result output module maps the image features to a segmentation mask; in the object detection task, the image features are converted into the bounding box coordinates and class information of the objects; the output results can be stored in the local storage device, sent to a remote server through a network protocol, or directly displayed on a display device.
8. The steel defect detection method based on the Mamba network according to claim 1, characterized in that, The system monitoring and management module monitors the resource usage of each module of the system and the performance metrics of model training and inference in real time; dynamically adjusts the key parameters of the system according to the monitoring data and actual task requirements; has the ability to detect and repair faults, and takes corresponding repair measures and records the fault information when a fault is detected.
9. A steel flaw detection method based on the Mamba network, characterized in that, It includes the following steps: S1. The image data input module reads the image data from the specified data source, preprocesses it, and passes the processed image data to the encoder module; S2. The encoder module processes the image data layer by layer. The first 3 layers of the HFA module extract the shallow features, and the last 3 layers of the PMLayer module extract the deep features. The extracted features are passed to the skip connection module; S3. The skip connection module fuses the multi-scale features of each layer of the encoder. After spatial attention and channel attention processing, the fused features are passed to the decoder module; S4. The decoder module uses PMLayer to optimize and adjust the input features, adds them to the fused features of the corresponding layers in the skip connection, restores the image resolution through upsampling and convolution operations, and outputs the processed image features to the result output module; S5. The result output module converts the image features into the actual task results and stores, transmits, or displays them according to the requirements; S6. The system monitoring and management module monitors the running status of the system in real time, adjusts the parameters and processes the faults according to the monitoring data to ensure the stable operation of the system.