Steel flaw detection method, device, medium and equipment

Through generative adversarial network denoising processing and multi-layer neural network feature extraction, combined with multi-scale feature fusion and regional proposal network detection, the problem of low detection accuracy in steel defect detection is solved, and more efficient defect detection is achieved.

CN120088246AActive Publication Date: 2025-06-03XI'AN POLYTECHNIC UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510557056.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-06-03
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The prior art has low detection accuracy in steel defect detection, making it difficult to effectively distinguish complex noise from natural textures, resulting in loss of image edge information and excessive smooth details, affecting detection accuracy.

Method used

Generative adversarial networks are used for denoising, combined with pre-trained Swin Transformer and convolutional neural network models for feature extraction, and fusion features are integrated through adaptive multi-scale feature fusion modules. Finally, the regional proposal network and the regional network of interest are used for detection of defect types and locations.

Benefits of technology

It improves the denoising performance, retains image details, solves the problem of high-frequency feature loss during image denoising, and optimizes the accuracy of feature extraction through multi-scale feature fusion, improving the accuracy of steel defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088246A_ABST
    Figure CN120088246A_ABST
Patent Text Reader

Abstract

The invention discloses a steel defect detection method and device, a medium and equipment, and relates to the technical field of image recognition. The method comprises the following steps: acquiring an image training data set of steel and an image of to-be-detected steel; taking an image in the image training data set as input, taking the denoised steel image as output, and training the generative adversarial network to obtain a denoising model; inputting the image of the steel to be detected into a denoising model to obtain a denoised image; inputting the denoised steel image into a pre-trained swin transformer model and a pre-trained convolutional neural network model for feature extraction to obtain a first feature and a second feature of the denoised image, and performing feature fusion on the first feature and the second feature to obtain a feature map about the to-be-detected steel; and processing the feature map by using the regional proposal network and the region-of-interest network to obtain the defect type and the defect position of the steel to be detected. According to the scheme, the accuracy of defect detection on the steel can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly to a method, device, medium and equipment for detecting steel defects. Background Art

[0002] As a core material for construction and infrastructure construction, steel is not only the cornerstone of modern industry but also an important factor driving social development and economic growth. The surface of steel is usually disturbed by various factors. For example, the surface of steel often has complex textures or reflections, especially the metal surface has specular reflections, making it difficult to distinguish defects from the background. Moreover, during the process of image acquisition, it may be affected by factors such as light and sensors, resulting in a lot of irrelevant noise in the image, which affects the defect recognition ability of the model.

[0003] Traditional steel defect detection mainly relies on manual inspection one by one, but this detection method is a labor-intensive process. With the rapid development of deep learning and the improvement of device computing power, deep convolutional neural networks have achieved remarkable results in many fields and are currently the mainstream direction for detecting surface defects of industrial products. Existing methods for detecting steel surface defects based on deep convolutional neural networks need to first find the prior information of the image and then use an optimization algorithm to iteratively solve the model. The time and computational cost spent on its complex optimization process are very large. At the same time, it is also easy to cause image blurring and detail loss. In addition, existing denoising methods mainly use synthetic noise images for research, and the types of noise are relatively single. However, for the images taken during the steel detection process, the types of noise may be a superposition of multiple types, with a relatively complex distribution and unknown intensity.

[0004] With the application of deep learning in image denoising, the denoising effect has been improved. However, when dealing with real noise images, due to problems such as noise interference, diverse types of steel surface defects, different shapes, large scale differences, and small contrast between defects and the background, it is difficult to distinguish complex noise from natural textures, resulting in problems such as loss of image edge information and over-smoothing of details, leading to low accuracy in detecting steel defects. Summary of the Invention

[0005] Based on this, it is necessary to provide a method, device, medium and equipment for detecting steel defects in view of the technical problem of low accuracy in detecting steel defects in the prior art.

[0006] The present invention adopts the following technical solutions: In the first aspect, the present invention provides a method for detecting steel defects, the method comprising: Obtain an image training data set related to steel, and obtain an image of the steel to be detected; Taking the images in the image training dataset as inputs and the denoised steel images as outputs, training a generative adversarial network to obtain a denoising network model; Inputting the image of the steel to be detected into the denoising network model to obtain the denoised image of the steel to be detected; inputting the denoised image of the steel to be detected into a pre-trained swin transformer model and a pre-trained convolutional neural network model respectively for feature extraction to obtain the first feature and the second feature of the denoised image of the steel to be detected, and performing multi-scale feature fusion on the first feature and the second feature to obtain a feature map of the steel to be detected; Using a region proposal network and a region of interest network to process the feature map to obtain the defect type and defect location of the steel to be detected.

[0007] Further, performing multi-scale feature fusion on the first feature and the second feature is achieved through an adaptive multi-scale feature fusion module, and the expression of the adaptive multi-scale feature fusion is: ; where P i is the fused feature, C i is the feature map of the i-th layer, and C i+1 is the feature map of the (i + 1)-th layer.

[0008] Further, the adaptive multi-scale feature fusion module includes a cross-attention branch and a selectable convolution kernel branch; where the cross-attention branch is used to determine the association between adjacent convolutional layers and fuse the feature information of different convolutional layers through a cross-layer attention mechanism; The selectable convolution kernel branch is used to adaptively select different convolution kernels in each layer of the feature map and adjust the receptive field of the convolution target in each layer of the feature map.

[0009] Further, the using a region proposal network and a region of interest network to process the feature map to obtain the defect type and defect location of the steel to be detected specifically includes: Inputting the feature map into the region proposal network to obtain the processing result of the feature map; Inputting the processing result of the feature map into the region of interest network for further processing to obtain the defect type and defect location of the steel to be detected.

[0010] Further, taking the images in the image training dataset as inputs and the denoised steel images as outputs, training a generative adversarial network to obtain a denoising network model specifically includes: Input the Nth image in the image training dataset into the generator of the generative adversarial network to obtain a first denoised image of the Nth image, where N is a positive integer; Input the Nth image and the first denoised image of the Nth image into the discriminator of the generative adversarial network to obtain the similarity probability between the Nth image and the first denoised image of the Nth image; Input the similarity probability into the loss function of the generative adversarial network to obtain a loss value, and input the loss value into the generator for adaptive moment estimation iterative training to obtain the denoising network model.

[0011] Further, the loss function is obtained by weighted summation of an adversarial loss, a perceptual loss, and a loss function that combines the image texture feature matching of the discriminator, and its calculation expression is: ; where Loss is the output of the loss function, is the adversarial loss, is the perceptual loss, is the loss function that combines the image texture feature matching of the discriminator, is 's weight, is 's weight, is 's weight The 's calculation expression is: ; where, represents the i-th convolutional layer of the discriminator, G represents the generator, represents the data distribution of the original noise samples, represents the data distribution of samples without noise, z represents the noise, T represents the number of layers of the discriminator, represents the number of elements in each layer, k represents the discrimination scale, represents the .

[0012] Further, the pre-trained convolutional neural network model is trained based on EfficientNet.

[0013] In a second aspect, the present invention provides a steel flaw detection device, including: An acquisition module, configured to acquire an image training dataset related to steel, and acquire an image of the steel to be detected; A training module, which is used to take the images in the image training dataset as inputs, take the denoised steel images as outputs, train a generative adversarial network, and obtain a denoising network model; A feature extraction module, which is used to input the image of the steel to be detected into the denoising network model to obtain the denoised image of the steel to be detected; input the denoised image of the steel to be detected into a pre-trained swin transformer model and a pre-trained convolutional neural network model respectively for feature extraction, obtain the first feature and the second feature of the denoised image of the steel to be detected, and perform multi-scale feature fusion on the first feature and the second feature to obtain a feature map of the steel to be detected; A detection module, which is used to process the feature map by using a region proposal network and a region of interest network to obtain the defect type and defect location of the steel to be detected.

[0014] The present invention provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned steel defect detection method is implemented.

[0015] The present invention provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned steel defect detection method is implemented.

[0016] The at least one technical solution adopted by the present invention can achieve the following beneficial effects: The present invention takes an image training dataset related to steel as an input and a denoised steel image as an output to train a generative adversarial network, which can improve the denoising performance while retaining the details of the image, and solve the problem of losing high-frequency features during the denoising process of the image. And by using a pre-trained swin transformer model and a pre-trained convolutional neural network model to extract the first feature and the second feature of the denoised image of the steel to be detected respectively, and performing multi-scale feature fusion on the first feature and the second feature, the local features and global features of the image of the steel to be detected are fused, and the problems of feature loss and feature redundancy during the feature extraction process of the image are solved, thereby optimizing the accuracy of feature extraction for the image of the steel to be detected. Finally, by using a region proposal network and a region of interest network to process the feature map, the defect type and defect location of the steel to be detected can be accurately obtained, thereby improving the accuracy of detecting steel defects. Description of the Drawings

[0017] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0018] Figure 1 Schematic flow diagram of a steel flaw detection method provided by the present invention; Figure 2 Schematic technical route diagram of a steel flaw detection method provided by the present invention; Figure 3 Schematic diagram of the denoising network model architecture provided by the present invention; Figure 4 Schematic diagram of the structure of the pre-trained convolutional neural network model and the pre-trained swin transformer model provided by the present invention; Figure 5 Schematic diagram of the structure of the adaptive multi-scale fusion network provided by the present invention; Figure 6 Schematic diagram of the structure of the selectable convolution kernel module provided by the present invention; Figure 7 Schematic diagram of the structure of the MBconv network provided by the present invention; Figure 8 Schematic diagram of a steel flaw detection device provided by the present invention; Figure 9 Schematic diagram of a computer device for implementing a steel flaw detection method provided by the present invention. Detailed implementation manners

[0019] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0020] At present, the server mentioned in the present invention can be a server set up on a business platform or a device such as a desktop computer or a laptop computer that can execute the solution of the present invention. For the convenience of description, only the server will be used as the execution subject for description below. The technical solutions provided by each embodiment of the present invention will be described in detail below with reference to the drawings.

[0021] Refer to Figure 1 , a steel flaw detection method in the present invention specifically includes the following steps: S10: Obtain an image training dataset related to steel, and obtain an image of the steel to be detected.

[0022] In this embodiment, the image training dataset related to steel is an original image dataset containing noise, and the image of the steel to be detected is obtained by photographing the steel that needs to be flaw-detected.

[0023] S20: Use the images in the image training dataset as inputs and the denoised steel images as outputs to train a generative adversarial network to obtain a denoising network model. The generator is used to denoise the images in the image training dataset to obtain images without noise. The discriminator is used to calculate the similarity probability between the images in the image training dataset and the images without noise.

[0024] In this embodiment, referring to Figure 2 , use each image in the image training dataset as an input and the denoised steel image as an output to train a generative adversarial network (GAN) to obtain a denoising network model.

[0025] In this embodiment, use the images in the image training dataset as inputs and the denoised steel images as outputs to train a generative adversarial network to obtain a denoising network model, which specifically includes: Input the Nth image in the image training dataset into the generator of the generative adversarial network to obtain a first denoised image of the Nth image, where N is a positive integer.

[0026] Input the Nth image and the first denoised image of the Nth image into the discriminator of the generative adversarial network to obtain the similarity probability between the Nth image and the first denoised image of the Nth image.

[0027] Input the similarity probability into the loss function of the generative adversarial network to obtain a loss value, and input the loss value into the generator for adaptive moment estimation (Adam) iterative training to obtain a denoising network model.

[0028] Specifically, the structure of the generative adversarial network is as Figure 3As shown, it includes a generation network, a discriminative network, and a loss function. Among them, the generation network contains a generator G, which is used to remove the noise in the noisy images (each image in the image training dataset) to obtain a denoised image (the first denoised image). The discriminative network contains a discriminator D, which takes the first denoised image generated by the generator G and the noisy image as inputs, outputs a value in the range of [0, 1], which is used to represent the similarity probability between the denoised image and the noisy image, and uses the similarity probability as the input of the loss function. The loss function determines the loss value based on the similarity probability between the noisy image and the first denoised image, and inputs the loss value into the generator G for adaptive moment estimation iterative training to obtain a denoising network model. Using the denoising network model, the noise in the noisy image can be removed to obtain the final denoised image. Among them, the first denoised image refers to the image output by the generator G, and the denoised image of the steel to be detected refers to the final denoised image output by the denoising network model.

[0029] Specifically, the generator G is composed of a parallel neural network with residual dense blocks, which can extract more feature information without increasing the time cost. The discriminative network adopts a fully convolutional network structure to achieve pixel-level classification of images, so as to improve the performance of the discriminator D.

[0030] S30: Input the image of the steel to be detected into the denoising network model to obtain the denoised image of the steel to be detected; input the denoised image of the steel to be detected into the pre-trained swin transformer model and the pre-trained convolutional neural network model respectively for feature extraction to obtain the first feature and the second feature of the steel to be detected, and perform multi-scale feature fusion on the first feature and the second feature to obtain the feature map of the steel to be detected.

[0031] In this embodiment, referring to Figure 2 , the denoised image of the steel to be detected is input into the pre-trained swin transformer model and the pre-trained convolutional neural network model respectively for feature extraction. Among them, the first feature refers to the global feature obtained by extracting the features of the denoised image of the steel to be detected through the pre-trained swin transformer model, and the second feature refers to the local feature obtained by extracting the features of the denoised image of the steel to be detected through the pre-trained convolutional neural network model.

[0032] Specifically, referring to Figure 4, Swin Transformer and CNN are used as the backbone networks of two branches, dynamically selecting, weighting, and fusing features from the two branches of CNN and Swin Transformer. By using the local features extracted by CNN and the global features extracted by Swin Transformer, the problems of information loss or feature redundancy that may exist in traditional methods are avoided. The importance of different features is adaptively adjusted according to the task requirements, thereby optimizing the detection effect. Among them, the swintransformer model includes a patch partition module, a linear embedding module Linear Embeding, a first shifted window transformer module (Swin Transformer Block), a first path fusion module (PatchMerging), a second shifted window transformer module (Swin Transformer Block), a second path fusion module (Patch Merging), a third shifted window transformer module (Swin Transformer Block), a third path fusion module (Patch Merging), and a fourth shifted window transformer module (Swin Transformer Block).

[0033] The convolutional neural network model includes five convolutional layers. Among them, the second, third, fourth, and fifth convolutional layers are each composed of three residual blocks.

[0034] S40: Use the region proposal network and the region of interest network to process the feature map to obtain the defect type and defect location of the steel to be detected.

[0035] Based on Figure 1A steel flaw detection method is shown. In the present invention, an image training data set related to steel is used as input, and the denoised steel image is used as output to train a generative adversarial network, which can improve the denoising performance while retaining the details of the image, and solve the problem of losing high-frequency features during the denoising process of the image. Moreover, the first feature and the second feature of the image of the steel to be detected after denoising are respectively extracted by a pre-trained swin transformer model and a pre-trained convolutional neural network model, and multi-scale feature fusion is performed on the first feature and the second feature, realizing the fusion of the local features and the global features of the image of the steel to be detected, and solving the problems of feature loss and feature redundancy during the feature extraction process of the image, thereby optimizing the accuracy of feature extraction for the image of the steel to be detected. Finally, the region proposal network and the region of interest network are used to process the feature map, and the flaw type and flaw position of the steel to be detected can be accurately obtained, thereby improving the accuracy of detecting steel flaws.

[0036] When applying a steel flaw detection method provided by the present invention, it is not necessary to execute according to Figure 1 the order of the steps shown. The specific execution order of each step can be determined according to needs, and the present invention does not limit this.

[0037] In addition, in one or more embodiments of the present invention, the region proposal network and the region of interest network are used to process the feature map to obtain the flaw type and flaw position of the steel to be detected, specifically including: Input the feature map into the region proposal network to obtain the processing result of the feature map.

[0038] Input the processing result of the feature map into the region of interest network for further processing to obtain the flaw type and flaw position of the steel to be detected.

[0039] In this embodiment, referring to Figure 2 , the feature map is first input into the region proposal network (Region Proposal Network, RPN) for preliminary processing, and then the feature map after preliminary processing is input into the region of interest network (Region of Interest, ROI) for processing to obtain the flaw type and flaw position of the steel to be detected.

[0040] The solution of this embodiment can combine the advantages of RPN and ROI (ROI can focus on specific regions in the image, avoiding indiscriminate feature extraction for the entire image, thereby improving the efficiency and pertinence of feature extraction and helping to more accurately identify and locate the target object; when generating candidate regions, RPN uses a sliding window or convolution method to perform rapid calculations on the image feature map and can generate a large number of candidate regions in a short time), improving the accuracy and efficiency of obtaining the defect type and defect location of the steel to be detected.

[0041] Among them, the Region Proposal Network includes an anchor box layer, a classification regression layer, and a non-maximum suppression layer; the Region of Interest Network includes a Region of Interest pooling layer, a fully connected layer, and a classification regression layer.

[0042] In addition, in one or more embodiments of the present invention, the loss function is obtained by weighted summation of an adversarial loss, a perceptual loss, and a loss function for matching the image texture features of the discriminator, and its calculation expression is: ; where Loss is the output of the loss function, is the adversarial loss, is the perceptual loss, is the loss function for matching the image texture features of the discriminator, is the weight of is the weight of is the weight of

[0043] In this embodiment, , and The specific values of are obtained through experiments. First, observe the individual training of each loss function, and set the benchmark weights according to the order of magnitude of the loss functions when each loss function is approaching convergence. For example, assuming that the first loss function is around 0.02, and the second and third loss functions are around 0.2, then set the benchmark weights to 1:10.

[0044] This embodiment constructs a composite loss function from aspects such as the pixel, structure, and semantic features of the image, enabling the denoising network model to have network stability, denoising effectiveness, and the ability to maintain the stability of the image content.

[0045] Specifically, The calculation expression of is: ; where represents the i-th convolutional layer of the discriminator, G represents the generator, Represents the data distribution of the original noise samples, Represents the data distribution of the samples without noise, z represents the noise, and T represents the number of layers of the discriminator, Represents the number of elements in each layer, and k represents the discrimination scale, Represents at the k-th discrimination scale , and G represents.

[0046] In this embodiment, the Euclidean distance is used to measure the feature expression of the noisy image and the denoised image at different scales, and the comparison results are summarized in the total loss.

[0047] In addition, in one or more embodiments of the present invention, the multi-scale feature fusion of the first feature and the second feature is implemented through an adaptive multi-scale feature fusion module. The expression of the adaptive multi-scale feature fusion is: ; where P i is the fused feature, C i is the feature map of the i-th layer, and C i+1 is the feature map of the (i + 1)-th layer.

[0048] In this embodiment, an adaptive multi-scale feature fusion module (Adaptive Multi-Scale Feature Fusion, AMFF) is constructed based on the Feature Pyramid Network (FPN). AMFF not only inherits the core idea of FPN, but also introduces an adaptive receptive field adjustment mechanism on this basis, as well as strengthens the interaction between the current layer and the adjacent layer, so as to more effectively capture the diverse features of the target.

[0049] Specifically, referring to Figure 5 , the adaptive multi-scale feature fusion module includes a cross-attention branch (CEAM) and a large-scale kernel (LSK) branch.

[0050] Among them, the cross-attention branch is used to determine the association between adjacent convolutional layers and fuse the feature information of different convolutional layers through a cross-layer attention mechanism.

[0051] In this embodiment, the CEAM branch aims to find the internal connection between adjacent convolutional layers and dynamically fuse the feature information of different levels through a cross-layer attention mechanism. The specific expression is: ; where is the output of the CEAM branch, and sig is the sigmoid activation function, It is pixel-by-pixel multiplication. Under the guidance of the features of the i-th layer, this branch provides more valuable context interaction information for fusion. Branch integration: After the above effective operations, the outputs of LSK and CEAM are added to the original i-th layer feature map in a residual manner as

[0052] ; where is pixel-by-pixel summation. In summary, the features of the i-th layer are adaptively adjusted through intra-layer and cross-layer attention mechanisms. Compared with the traditional FPN, the AMFF module has higher flexibility and adaptability. It can adaptively adjust the fusion method and receptive field size according to the specific characteristics of the target, so as to more accurately capture the feature information of the target. At the same time, the AMFF module also improves the model's detection ability for multi-scale targets by strengthening the interaction between layers.

[0053] The LSK branch is used to adaptively select different convolutional kernels in the feature maps of each layer and adjust the receptive fields of the convolutional targets in the feature maps of each layer.

[0054] Refer to Figure 6 , the LSK branch allows the adaptive multi-scale feature fusion module to adaptively select different large convolutional kernels on the i-th layer feature map and adjust the receptive field of each convolutional target in the space of the feature maps of each layer as needed. This structure is mainly composed of large-kernel convolution and spatial selection mechanism.

[0055] To achieve effective modeling and adaptive selection of multiple long-range contexts, the LSK branch constructs a series of depth convolution sequences, which have larger growing kernels and increasing dilation rates.

[0056] Specifically, the kernel size k, dilation rate d, and receptive field RF dilation of the i-th depth convolution in this series are defined as follows: ; ; The increase in kernel size and dilation rate ensures that the receptive field expands fast enough, and by setting an upper limit on the dilation rate, it is ensured that the dilated convolution does not introduce gaps between the feature maps. When extracting rich context information from different ranges of the input, the LSK branch uses a depth convolution sequence with different receptive fields instead of a single large convolutional kernel. This method can generate multiple features with different large receptive fields, facilitating subsequent kernel selection. In addition, through sequential decomposition, that is, using a stack of multiple small convolutional kernels, compared with a single large convolutional kernel, while maintaining the same theoretical receptive field, the number of parameters is significantly reduced, improving the efficiency and flexibility of the model.

[0057] In addition, in one or more embodiments of the present invention, the pre-trained convolutional neural network model is constructed and trained based on EfficientNet (Efficient Network).

[0058] In this embodiment, the EfficientNet is a lightweight model designed by Google for edge computing devices. As an efficient CNN architecture, EfficientNet has a good balance between performance and computing resources, and is particularly suitable for the lightweight CNN branch in steel defect detection. Through compound scaling and lightweight design, EfficientNet can reduce the computational overhead and memory occupancy while ensuring high accuracy, and is very suitable for real-time inference in industrial environments. Combining EfficientNet with Swin Transformer can further improve the feature extraction and global modeling capabilities, thus better coping with complex scenarios in the steel defect detection task.

[0059] Specifically, referring to Figure 6 , the EfficientNet feature extraction network consists of MBConv blocks, and MBConv is composed of SENet (Squeeze and Excitation Network) and a depth convolutional network. In the structure of the MBConv block, first use a 1×1 convolution to increase the dimension of the image, then perform depth convolution and SENet in sequence, and finally use 1×1 to reduce the dimension of the image output. EfficientNet achieves a more balanced network architecture by balancing the three dimensions of network depth, network width, and image resolution.

[0060] The above is a steel defect detection method provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding steel defect detection device, as Figure 8 shown, including: An acquisition module, configured to acquire an image training data set related to steel and acquire an image of the steel to be detected.

[0061] A training module, configured to use the images in the image training data set as inputs and the denoised steel images as outputs to train a generative adversarial network to obtain a denoising network model.

[0062] A feature extraction module, configured to input an image of the steel to be detected into a denoising network model to obtain a denoised image of the steel to be detected; input the denoised image of the steel to be detected into a pre-trained swin transformer model and a pre-trained convolutional neural network model respectively for feature extraction to obtain a first feature and a second feature of the denoised image of the steel to be detected, and perform multi-scale feature fusion on the first feature and the second feature to obtain a feature map of the steel to be detected.

[0063] A detection module, configured to process the feature map by using a region proposal network and a region of interest network to obtain the defect type and defect location of the steel to be detected.

[0064] For the specific limitations of a steel defect detection device, reference may be made to the limitations of a steel defect detection method in the foregoing text, which will not be elaborated herein. Each module in the steel defect detection device may be implemented in whole or in part by software, hardware, and their combination. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the foregoing modules.

[0065] The present invention also provides a computer-readable storage medium storing a computer program, which can be used to execute the Figure 1 provided steel defect detection method.

[0066] The present invention also provides Figure 9 a schematic structural diagram of the computer device shown in, as Figure 9 shown, at the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, other hardware required for other services may also be included. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the Figure 1 provided steel defect detection method.

[0067] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the described embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0068] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded by the present invention.

Claims

1. A steel defect detection method, characterized in that: include: Obtain an image training data set related to steel, and obtain an image of the steel to be detected; Using the images in the image training data set as input and the denoised steel images as output, training the generative adversarial network to obtain a denoising network model; Inputting the image of the steel to be detected into the denoising network model to obtain a denoised image of the steel to be detected; inputting the denoised image of the steel to be detected into a pre-trained swin transformer model and a pre-trained convolutional neural network model respectively for feature extraction to obtain a first feature and a second feature of the denoised image of the steel to be detected, and performing multi-scale feature fusion on the first feature and the second feature to obtain a feature map of the steel to be detected; The feature map is processed using a region proposal network and a region of interest network to obtain the defect type and defect location of the steel to be detected.

2. A steel defect detection method according to claim 1, characterized in that: The multi-scale feature fusion of the first feature and the second feature is implemented by an adaptive multi-scale feature fusion module. The expression of the adaptive multi-scale feature fusion is: ; Among them, P i is the fused feature, C i is the feature map of the i-th layer, C i+1 is the feature map of the i+1th layer.

3. A steel defect detection method according to claim 2, characterized in that: The adaptive multi-scale feature fusion module includes a cross attention branch and a selectable convolution kernel branch; The cross-attention branch is used to determine the association between adjacent convolutional layers and fuse the feature information of different convolutional layers through the cross-layer attention mechanism; The selectable convolution kernel branch is used to adaptively select different convolution kernels in each layer of feature maps and adjust the receptive field of the convolution target in each layer of feature maps.

4. A steel defect detection method according to claim 1, characterized in that: The method of processing the feature map using a region proposal network and a region of interest network to obtain the defect type and defect location of the steel to be detected specifically includes: Inputting the feature map into the region proposal network to obtain a processing result of the feature map; The processing result of the feature map is input into the region of interest network for further processing to obtain the defect type and defect position of the steel to be detected.

5. A steel defect detection method according to claim 1, characterized in that: The images in the image training data set are used as input, and the denoised steel images are used as output, and the generative adversarial network is trained to obtain a denoising network model, which specifically includes: Inputting the Nth image in the image training data set into the generator of the generative adversarial network to obtain a first denoised image of the Nth image, where N is a positive integer; Inputting the Nth image and the first denoised image of the Nth image into the discriminator of the generative adversarial network to obtain a similarity probability between the Nth image and the first denoised image of the Nth image; The similarity probability is input into the loss function of the generative adversarial network to obtain a loss value, and the loss value is input into the generator for adaptive moment estimation iterative training to obtain the denoising network model.

6. A steel defect detection method according to claim 1, characterized in that: The loss function is obtained by weighted summation of adversarial loss, perceptual loss and loss function of image texture feature matching combined with the discriminator, and its calculation expression is: ; Among them, Loss is the output of the loss function, To combat losses, is the perceived loss, is the loss function for image texture feature matching combined with the discriminator, yes The weight of yes The weight of yes Weight Said The calculation expression is: ; in, represents the i-th convolutional layer of the discriminator, G represents the generator, represents the data distribution of the original noise samples, represents the data distribution of samples without noise, z represents noise, T represents the number of layers of the discriminator, represents the number of elements in each layer, k represents the scale of discrimination, represents the kth discriminant scale .

7. A steel defect detection method according to claim 1, characterized in that: The pre-trained convolutional neural network model is trained based on EfficientNet.

8. A steel defect detection device, characterized in that: include: An acquisition module, used to acquire an image training data set related to steel materials and acquire images of steel materials to be detected; A training module, used to take the images in the image training data set as input and the denoised steel images as output, train the generative adversarial network, and obtain a denoising network model; A feature extraction module is used to input the image of the steel to be detected into the denoising network model to obtain a denoised image of the steel to be detected; input the denoised image of the steel to be detected into a pre-trained swintransformer model and a pre-trained convolutional neural network model respectively for feature extraction to obtain a first feature and a second feature of the denoised image of the steel to be detected, and perform multi-scale feature fusion on the first feature and the second feature to obtain a feature map of the steel to be detected; The detection module is used to process the feature map using a region proposal network and a region of interest network to obtain the defect type and defect position of the steel to be detected.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the steel defect detection method according to any one of claims 1 to 7 is implemented.

10. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, a steel defect detection method as claimed in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Injection molding part defect detection method and device based on normal sample auxiliary feature extraction and medium

    CN115082386A

  • Industrial product surface flaw detection method

    CN117522821A

  • Steel surface defect detection method and system based on image denoising

    CN118840346A

  • Steel surface defect detection method and device

    CN119399125A

  • PPY-YOLO-based steel surface defect detection method and system

    CN119672031A