Steel defect detection method adopting domain adversarial adaptive multi-scale feature fusion

By adopting the method of domain-adaptive adaptive multi-scale feature fusion, the problem of insufficient feature representation capability and difficulty in cross-domain feature migration in the graph neural network steel defect detection model is solved, and more efficient and accurate steel defect detection is achieved.

CN119941664APending Publication Date: 2025-05-06PUTIAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510005471.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing steel defect detection model based on graph neural networks has problems such as insufficient retention and computational efficiency of superpixel segmentation, insufficient feature representation ability, and difficulty in cross-domain feature migration, resulting in a decrease in detection accuracy.

Method used

The domain-adversarial adaptive multi-scale feature fusion method is adopted, and the image is converted into a graph structure through simple linear iterative clustering algorithm and multi-scale feature fusion, which enhances the feature representation ability of node features, and uses dynamic multi-head attention mechanism networks and generative adversarial networks for feature extraction and transfer learning, and optimizes the loss function to improve model performance.

Benefits of technology

It improves the performance of the steel defect detection model, reduces training time, enhances the accuracy of detection, and can better handle complex scenarios and diversified defect characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941664A_ABST
    Figure CN119941664A_ABST
Patent Text Reader

Abstract

The invention discloses a steel defect detection method adopting domain adversarial adaptive multi-scale feature fusion, and the method comprises the steps: collecting and inputting a first steel image training set, converting a training image in the first steel image training set into a graph structure through employing a simple linear iteration clustering algorithm and multi-scale feature fusion, obtaining an initial node feature of each node in the graph structure; inputting the initial node features into a dynamic multi-head attention mechanism network to enable the dynamic multi-head attention mechanism to enhance the feature representation capability of the initial node features, and outputting corresponding enhanced node features; and training through transfer learning and a generative adversarial network to obtain a steel defect detection model corresponding to the first steel image training set. The performance of the steel defect detection model can be improved, so that the steel defect detection is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning, and in particular to a steel defect detection method using domain-adversarial adaptive multi-scale feature fusion. Background Art

[0002] As an indispensable and important material, steel plays an important role in modern industrial production and is widely used in various fields such as mechanical processing. However, there is a close correlation between the quality of steel and its surface defects. Surface defects may lead to a decrease in the strength of steel and may even cause serious accidents, posing a serious threat to production safety and product quality. The steel industry occupies an important position in the global economy, especially in China, where steel production accounts for more than half of the world's total. However, with the intensification of market competition, steel companies are facing increasingly severe quality control challenges. Surface defects of steel, such as scratches, pores, cracks, etc., directly affect the quality and safety of products. Therefore, how to efficiently and accurately detect steel defects has become an urgent problem to be solved in the industry.

[0003] In recent years, the application of defect detection systems in industrial production has expanded significantly due to advances in artificial intelligence technology. Surface defect detection is critical in steel production, and the use of deep learning-based object detection technology for this purpose has shown promise. However, deep learning models used for training and testing require a lot of computing resources, which poses a challenge for defect detection when deployed on terminal devices with limited computing power in real applications. Therefore, optimizing deep learning models to reduce complexity and computational complexity while maintaining detection accuracy is a key challenge in the field of steel surface defect detection.

[0004] With the rapid development of deep learning technology, more and more researchers have begun to apply deep learning to steel defect detection. For example, the research team of Beijing Jiaotong University proposed a steel surface defect detection method based on convolutional neural network (Convolutional Neural Network), which can effectively identify different types of defects. The research team of the University of Michigan also proposed a steel surface defect detection method based on convolutional neural network, which can accurately identify different types of defects. In addition, some researchers have integrated multiple deep learning technologies to improve the accuracy and robustness of detection. However, the defect classification method based on CNN can usually only realize the recognition of defect categories in one frame of image. When there are multiple defect categories in the image at the same time, it can only identify defects with significant features and large areas, and cannot locate the defect area.

[0005] As an emerging deep learning technology, Graph Neural Network has begun to attract the attention of researchers in recent years. Studies have shown that GNN has significant advantages in processing complex data structures and diverse defect characteristics. Some domestic universities and research institutions are actively exploring the application of GNN in steel defect detection in order to improve the detection effect and efficiency. With the rise of graph neural networks, some researchers at home and abroad have begun to explore its application in steel defect detection. Some researchers have built large-scale steel defect datasets, such as NEU, GP, and WebSteel. At the same time, some evaluation criteria, such as accuracy, recall, and F1 score, have been proposed to evaluate the performance of different methods. With the rise of intelligent manufacturing, the steel industry urgently needs to introduce advanced detection technologies to improve production efficiency and product quality. Steel defects may lead to structural failure and then cause safety accidents. Therefore, timely and accurate detection of steel defects is a key link to ensure product quality and safety of use. By introducing advanced detection technologies, safety hazards can be effectively reduced and the safety of users' lives and property can be protected. The defect detection method based on GNN can not only reduce manual intervention, but also realize real-time monitoring and feedback, and promote the intelligent transformation of the steel industry.

[0006] However, the GNN-based defect detection method is still affected by problems such as the detail retention and computational efficiency of superpixel segmentation, insufficient feature representation capabilities, and difficulties in cross-domain feature migration, resulting in a decrease in detection accuracy. Summary of the invention

[0007] In view of some of the above-mentioned defects in the prior art, the technical problem to be solved by the present invention is to provide a steel defect detection method using domain-adversarial adaptive multi-scale feature fusion, aiming to improve the performance of the steel defect detection model, thereby making steel defect detection more accurate.

[0008] To achieve the above object, the present invention provides a steel defect detection method using domain-adversarial adaptive multi-scale feature fusion, the method comprising:

[0009] Step S1, collecting and inputting a first steel image training set, converting the training images in the first steel image training set into a graph structure by using a simple linear iterative clustering algorithm and multi-scale feature fusion, and obtaining initial node features of each node in the graph structure;

[0010] Step S2: inputting the initial node features into a dynamic multi-head attention mechanism network, so that the dynamic multi-head attention mechanism enhances the feature representation capability of the initial node features and outputs corresponding enhanced node features;

[0011] Step S3, determining a first source model that matches the enhanced node feature, and obtaining source domain data corresponding to the first source model, and determining the enhanced node feature as target domain data; wherein the first source model is a trained model;

[0012] Step S4: using a generative adversarial network to generate generated data that matches the source domain data and the target domain data; using transfer learning and the generative adversarial network to discriminate the generated data from the source domain data, extract common features between the source domain data and the target domain data, and obtain a corresponding first loss function;

[0013] Step S5: Based on the first loss function, a steel defect detection model corresponding to the first steel image training set is obtained by performing transfer learning on the first source model.

[0014] Optionally, the step S1 includes:

[0015] Collecting and inputting the first steel image training set, and performing superpixel segmentation on the training images in the first steel image training set using a simple linear iterative clustering algorithm; and constructing a graph structure corresponding to the training images according to the superpixel segmentation results;

[0016] The relationship between nodes in the graph structure is strengthened by the multi-scale feature fusion, thereby obtaining the initial node features of each node in the graph structure.

[0017] Optionally, the node features include three primary color features, texture features, shape features and position features; wherein the shape features include area and perimeter.

[0018] Optionally, step S2 includes:

[0019] Generate a query vector matrix, a key vector matrix, and a value vector matrix of the dynamic multi-head attention mechanism network according to the initial node features;

[0020] According to the query vector matrix, the key vector matrix and the value vector matrix, the corresponding enhanced node features are output.

[0021] Optionally, outputting the corresponding enhanced node features according to the query vector matrix, the key vector matrix, and the value vector matrix includes:

[0022] According to the query vector matrix, the key vector matrix and the value vector matrix, calculating the attention mechanism for each attention head to obtain an output vector of each attention head;

[0023] The output vectors are connected and linearly transformed to obtain the corresponding enhanced node features.

[0024] Optionally, step S3 includes:

[0025] According to the target model corresponding to the requirements of the first steel image training set, multiple source models with similar fields and already completed training are determined; from the multiple original models, the most suitable first source model is determined;

[0026] The first source model is used as an initial model for transfer learning, the data corresponding to the first source model is determined as source domain data for transfer learning, and the enhanced node features are determined as target domain data for transfer learning.

[0027] Optionally, step S4 includes:

[0028] Generate generated data that matches the source domain data and the target domain data using a generative adversarial network;

[0029] An adversarial classifier is constructed, and the adversarial classifier is trained based on the source domain data, the target domain data, and the generated data, so as to obtain common features between the source domain data and the target domain data, and obtain a corresponding first loss function.

[0030] Beneficial effects of the present invention: 1. The present invention proposes a SLIC (Simple Linear Iterative Clustering) segmentation for multi-scale feature extraction. The fine-tuned SLIC can achieve efficient superpixel segmentation through simple linear clustering, which is suitable for large-scale image processing. Secondly, the generated superpixel boundaries are accurate and can effectively retain image details. Multi-scale feature extraction is then performed on the segmented superpixels, by extracting color, texture and shape features at different scales and merging them. This fusion can enhance the model's recognition ability for diverse objects and make the segmentation results more accurate. 2. The dynamic multi-head attention mechanism adopted by the present invention allows the attention calculation method to be adjusted according to the task characteristics. Multi-head attention allows the model to calculate multiple attention heads in parallel, so that information can be extracted from different subspaces, thereby improving the diversity and richness of feature representation. Secondly, dynamic weight allocation allows each attention head to adjust its weight in real time according to contextual information, thereby focusing on important features more accurately. This mechanism enhances the ability of feature representation, allowing the model to perform better in complex scenes. 3. The present invention constructs a new domain adversarial migration framework to achieve cross-domain feature migration. By introducing an improved GAN (Generative Adversarial Network) and designing domain adversarial training technology, the feature distribution tends to be consistent, thereby enhancing the generalization ability of the model. By selecting features related to the target task and recalibrating them, the effect of transfer learning is improved, and the commonalities between the source and target domains are captured. Knowledge is transferred from multiple source domains to make full use of feature information from different domains and further improve model performance.

[0031] In summary, the present invention reduces the training time of the steel defect detection model, improves the performance of the steel defect detection model, and thus makes steel defect detection more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a flow chart of a steel defect detection method using domain-adversarial adaptive multi-scale feature fusion provided by a specific embodiment of the present invention;

[0033] Figure 2 It is an overall framework diagram of a steel defect detection method using domain-adversarial adaptive multi-scale feature fusion provided by a specific embodiment of the present invention. DETAILED DESCRIPTION

[0034] The present invention discloses a steel defect detection method using domain-adversarial adaptive multi-scale feature fusion. Those skilled in the art can refer to the content of this article and appropriately improve the technical details for implementation. It is particularly important to point out that all similar substitutions and modifications are obvious to those skilled in the art, and they are all deemed to be included in the present invention. The methods and applications of the present invention have been described through preferred embodiments, and relevant personnel can obviously modify or appropriately change and combine the methods and applications described herein without departing from the content, spirit and scope of the present invention to implement and apply the technology of the present invention.

[0035] The applicant has found that the existing steel defect detection model based on graph neural network still faces the following three challenges: (1) Detail retention and computational efficiency of superpixel segmentation: In large-scale image processing tasks, how to accurately retain image details and boundaries while ensuring computational efficiency is a major challenge. Traditional segmentation methods often find it difficult to balance processing speed and accuracy, resulting in the inability to effectively extract important features in complex images. (2) Insufficient feature representation capabilities: In complex scenes, how to effectively extract and represent diverse features is a significant challenge. The existing attention mechanism may not be able to fully utilize contextual information, resulting in the singleness and insufficiency of feature expression, limiting the performance of the model in diverse tasks. (3) Difficulty in cross-domain feature transfer: In transfer learning, the difference in feature distribution between the source domain and the target domain often leads to a decline in model performance. Effectively realizing cross-domain feature transfer, improving the generalization ability of the model, and ensuring that useful information in the source domain is captured in the target task are major challenges.

[0036] Therefore, an embodiment of the present invention provides a steel defect detection method using domain-adversarial adaptive multi-scale feature fusion, such as Figure 1 As shown, the method includes:

[0037] Step S1, collecting and inputting a first steel image training set, converting the training images in the first steel image training set into a graph structure by using a simple linear iterative clustering algorithm and multi-scale feature fusion, and obtaining initial node features of each node in the graph structure.

[0038] It should be noted that Simple Linear Iterative Clustering (SLIC) is an image segmentation algorithm, especially used to generate superpixels. Multi-scale feature fusion (MSFF) is an important technology in the field of computer vision and image processing, which aims to improve the performance of the model by fusing feature information of different scales.

[0039] In this specific embodiment, step S1 includes:

[0040] Collect and input a first steel image training set, and use a simple linear iterative clustering algorithm to perform superpixel segmentation on the training images in the first steel image training set; and construct a graph structure corresponding to the training images according to the superpixel segmentation results;

[0041] The relationship between nodes in the graph structure is strengthened by multi-scale feature fusion, and then the initial node features of each node in the graph structure are obtained.

[0042] In this specific embodiment, the node features include three primary color features, texture features, shape features, and position features; wherein the shape features include area and perimeter.

[0043] The embodiment of the present invention converts the image into a graph structure, making the definition of nodes and edges more intuitive, which is helpful for subsequent feature extraction and optimization. SLIC (Simple Linear Iterative Clustering) is a commonly used superpixel segmentation algorithm that aims to generate superpixels with good connectivity and uniformity, reducing the computational complexity of subsequent image processing. The size of the superpixel can be adjusted as needed to adapt to different types of images and application scenarios. It is widely used in medical image processing, image retrieval, target detection and other fields. However, it is sensitive to noise, which may affect the quality of superpixels. And in complex backgrounds or low-contrast images, the boundaries of superpixels may not be clear enough. To address this shortcoming, the embodiment of the present invention adopts multi-scale feature extraction based on the original SLIC method, combining feature information of different scales to capture richer image features and adapt to defects or objects of different sizes.

[0044] In this specific embodiment, the image I is converted into a graph structure G = (V, E, A, X), where: (1) V = {v1, v2, ..., v i ,…,v n} is a node set; (2) E = {e i,j =(v i ,v j )} is the edge set; (3) A∈R n×n is the adjacency matrix, where a i j represents node v i and v j The weight of the edge between them; (4) X∈R n×d is a feature matrix, where X i ∈R d Represents a given node v i The image I is divided into N parts. Each superpixel v1,v2,…,v n , with superpixels v1,v2,…,v generated under SLIC nis a node, which will be used as the feature embedding of the graph. The edge set represents the connection between two nodes. The adjacency matrix represents the connection between any two nodes, and the value of the element in the adjacency matrix is ​​0 or 1. The definition of the adjacency matrix is ​​shown in formula (1):

[0045]

[0046] Node features can be represented by superpixel visual features of the region:

[0047] X i =[R,G,B,T,S,P,x,y](2)

[0048] Among them, R, G, B are color features, namely red, green, and blue; T is texture feature; S, P are shape features, namely area and perimeter; x, y are position features.

[0049] The embodiment of the present invention combines the SLIC method and the multi-scale feature fusion (MSFF) method to construct a graph structure. The superpixels created by SLIC provide clear node representation and retain the boundary information of the image; while MSFF strengthens the relationship between nodes through multi-scale features, making the graph structure more expressive. This combination improves the operability and analysis capabilities of the graph, enhances the performance of graph algorithms in segmentation, classification and detection tasks, and can better capture details and contextual information in complex scenes.

[0050] Step S2: Input the initial node features into the dynamic multi-head attention mechanism network so that the dynamic multi-head attention mechanism enhances the feature representation capability of the initial node features and outputs the corresponding enhanced node features.

[0051] In this specific embodiment, step S2 includes:

[0052] Generate query vector matrix, key vector matrix and value vector matrix of dynamic multi-head attention mechanism network according to the initial node features;

[0053] According to the query vector matrix, the key vector matrix and the value vector matrix, the corresponding enhanced node features are output.

[0054] Furthermore, according to the query vector matrix, the key vector matrix and the value vector matrix, the corresponding enhanced node features are output, including:

[0055] According to the query vector matrix, key vector matrix and value vector matrix, the attention mechanism is calculated for each attention head to obtain the output vector of each attention head;

[0056] Connect each output vector and perform linear transformation to obtain the corresponding enhanced node features.

[0057] In a specific embodiment, a single attention mechanism may not be able to fully capture the association between different features. Therefore, a dynamic multi-head attention mechanism is introduced to adjust the calculation method of attention according to the characteristics of the task. In the attention mechanism, Q (Query), K (Key) and V (Value) are three basic representations, and their role is to determine the degree of attention to the input information through calculation. Among them, Q is the query vector, which represents the information or context that the embodiment of the present invention wants to focus on. K is the key vector, which represents the characteristics or attributes of the input information. V is the value vector, which represents the actual information associated with each key. Their matrix is ​​composed of feature X i Generated by a specific linear transformation, as follows:

[0058] Q i h =X i W h Q

[0059] K i h =X i W h K

[0060] V i h =X i W h V (3)

[0061] Where h represents the hth attention head; W h Q ,W h K ,W h V Represent the weight matrix of the h-th attention head for each vector.

[0062] For each head h, the calculation of the attention mechanism can be expressed as:

[0063]

[0064] Where Q i h ,K i h ,V i h are the query matrix, key matrix, and value matrix of the h-th attention head respectively; softmax converts the model output into a probability distribution, which is suitable for multi-class classification tasks. Indicates that the similarity between the query vector and all key vectors is calculated using the dot product; is a small context-aware neural network used to generate additional adjustments; dk is the dimension of the key, used for scaling. Concatenating the output vectors of multiple dynamic attention heads yields:

[0065]

[0066] Perform a linear transformation on the connected matrix to obtain the weighted feature information output Z:

[0067] Z=σ(MultiHead(Q i ,K i ,V i )×W)(6)

[0068] Where σ is the ReLU activation function, which is applied to the aggregated features, adding nonlinearity to the neural network and helping to learn more complex feature representations. W is a trainable weight matrix.

[0069] By dynamically calculating the number and weights of attention heads, the model can flexibly adjust key features according to the needs of the current task and improve the effectiveness of feature representation.

[0070] It should be noted that adding This adjustment item reflects the dynamic allocation mechanism of attention in the embodiment of the present invention, which enables the model to focus on different features more flexibly and allows the model to better capture the potential relationships and features in complex data.

[0071] Step S3: determine a first source model that matches the enhanced node feature, obtain source domain data corresponding to the first source model, and determine the enhanced node feature as the target domain data.

[0072] Among them, the first source model is a trained model.

[0073] In this specific embodiment, step S3 includes:

[0074] According to the target model corresponding to the requirements of the first steel image training set, multiple source models with similar fields and already completed training are determined; and the most suitable first source model is determined from multiple original models;

[0075] The first source model is used as the initial model of transfer learning, the data corresponding to the first source model is determined as the source domain data of transfer learning, and the enhanced node features are determined as the target domain data of transfer learning.

[0076] Step S4: using a generative adversarial network to generate generated data that matches the source domain data and the target domain data; using transfer learning and a generative adversarial network to discriminate the generated data and the source domain data, extract the common features between the source domain data and the target domain data, and obtain the corresponding first loss function.

[0077] In this specific embodiment, step S4 includes:

[0078] Generate data that matches the source domain data and the target domain data using a generative adversarial network;

[0079] An adversarial classifier is constructed and trained based on the source domain data, the target domain data and the generated data, so as to obtain the common features between the source domain data and the target domain data and obtain the corresponding first loss function.

[0080] It should be noted that Generative Adversarial Transfer Learning combines the ideas of Generative Adversarial Network (GAN) and transfer learning, aiming to achieve knowledge transfer and adaptation through generative adversarial methods. In transfer learning, there may be differences in the data distribution of the source domain and the target domain. Generative adversarial transfer learning uses the generator of GAN to generate data samples similar to the target domain, and uses the discriminator to distinguish between the generated data and the real data. In this process, the adversarial training of the generator and the discriminator helps to extract the common features between the source domain and the target domain and align their feature spaces.

[0081] Step S5: Based on the first loss function, a steel defect detection model corresponding to the first steel image training set is obtained by performing transfer learning on the first source model.

[0082] In this specific embodiment, in the field of deep learning, the computational overhead of model training is very high. To alleviate this problem, the embodiment of the present invention combines a generative adversarial network with a transfer learning strategy to improve the effectiveness of the model proposed in the embodiment of the present invention. Generative adversarial networks (GANs) can generate high-quality and diverse data, can perform unsupervised learning and effectively capture complex data distributions. Transfer learning can reduce training time and computing resources, improve model performance in data-scarce conditions, and can quickly adapt to new tasks. The combination of the two can achieve better results in steel defect detection and give full play to their respective advantages. The GATL (Generative Adversarial Transfer Learning) method emphasizes the key elements in the node graph and promotes the transfer of learning parameters from existing classifiers to new classifiers. In order to promote transfer learning, the embodiment of the present invention uses a domain adversarial training technique that aligns the distribution of node representations in the source graph and the target graph. This alignment is performed while ensuring that the distinctive information of the target task is retained. The embodiment of the present invention proposes an adversarial classifier that aims to distinguish the node representations of the source graph and the target graph. The domain adversarial training technique is expressed as:

[0083]

[0084] Where D represents the discriminator; G represents the generator; v s is a real sample from the source domain; G(z) is a sample generated by the generator, where z is random noise sampled from the latent space; v' is the discriminator's prediction of the target domain sample, and the prediction term helps to improve training stability, reduce mode collapse, and enhance the generalization ability of unknown data; Y is a conditional variable, and D(ZY) represents the output of the discriminator under given conditions. By introducing conditional information, the features of the generated samples are controllable, and samples of the target category can be generated according to specific labels, which improves interpretability and enhances the diversity and quality of generation. The embodiment of the present invention draws on the GAN method, and adds discriminator prediction terms and conditional variables on the basis of the former, improves training stability, and makes the features more controllable.

[0085] The key to domain adversarial transfer learning is to balance the performance of the source domain and the target domain. To this end, the embodiment of the present invention constructs a comprehensive model by jointly optimizing the loss functions of the source classifier, the domain adversarial classifier, and the target classifier, aiming to enhance the adaptability to the target domain while maintaining the performance of the source domain.

[0086] L=αL s +γ1L D +γ2L G +βL t (8)

[0087] Where L s , L D , L G , L t They are the source classifier loss, discriminator loss, generator loss, and target classifier loss. α and β are weight parameters used to adjust the loss of the source task and the target task. γ1 and γ2 are balance parameters used to adjust the loss of the discriminator and the generator. The specific introduction of each loss function is as follows.

[0088] (1) Source classifier loss. By minimizing the cross entropy loss, the model can perform classification tasks more accurately and improve the overall prediction ability.

[0089]

[0090] Where N s is the number of samples in the source dataset, y i is the true label of sample i, is the class y i The weight is set according to the number of samples in the category. i is the model's predicted probability for sample i.

[0091] (2) The goal of the discriminator D is to maximize the probability of correctly classifying real samples and generated samples. Its loss function can be expressed as:

[0092]

[0093] where v s is a real sample from the source domain; G(z) is a sample generated by the generator, z is a random noise sampled from the latent space; v' is the discriminator's prediction of the target domain sample; Y is a conditional variable, and D(ZY) represents the output of the discriminator under given conditions.

[0094] (3) The goal of the generator G is to generate samples that can deceive the discriminator, so that the discriminator thinks these samples are real. Its loss function can be expressed as:

[0095] L G =-Ε z [logD(G(z))]+λΕ v' [logD(G(f(x')))](11)

[0096] Where G(z) is the sample generated by the generator, z is the random noise sampled from the latent space; λ is the weight for adjusting the domain invariance loss; v' is the discriminator's prediction of the target domain sample; f(x') is the operation of extracting features from the target domain sample. The loss function of the generator in the embodiment of the present invention adds a prediction term. Because the generator is generated based on the detailed features of the original data, it is not suitable to use conditional variables. Therefore, the embodiment of the present invention adds weight adjustment before the prediction term.

[0097] (4) The target classifier loss evaluates the uncertainty of the model on the target task through entropy loss:

[0098]

[0099] Where N t is the number of samples of the target task, q i is the predicted probability of the model on the target task.

[0100] In a specific embodiment, the overall process of a steel defect detection method using domain-adversarial adaptive multi-scale feature fusion is as follows: Figure 2 As shown, an image is input, and superpixel segmentation is performed on the input image. After segmentation, multi-scale feature extraction is performed to obtain the graph structure and node features. The enhanced features of the node features are obtained through the dynamic multi-head attention mechanism. The loss function is determined through domain adversarial training technology and domain transfer optimization learning. Finally, the detection model is obtained. The defect just mentioned is detected through the detection model to obtain the diagnosis result. Figure 2It can be summarized that the model consists of three important parts and is jointly trained: 1) multi-scale feature extraction framework; 2) dynamic multi-head attention mechanism; 3) domain adversarial adaptive transfer.

[0101] The embodiment of the present invention proposes a SLIC (Simple Linear Iterative Clustering) segmentation for multi-scale feature extraction. The fine-tuned SLIC can achieve efficient superpixel segmentation through simple linear clustering, which is suitable for large-scale image processing. Secondly, the generated superpixel boundaries are accurate and can effectively retain image details. Multi-scale feature extraction is then performed on the segmented superpixels, by extracting color, texture and shape features at different scales and fusing them. This fusion can enhance the model's recognition ability for diverse objects, making the segmentation results more accurate.

[0102] The dynamic multi-head attention mechanism adopted in the embodiment of the present invention allows the attention calculation method to be adjusted according to the characteristics of the task. Multi-head attention allows the model to calculate multiple attention heads in parallel, so that information can be extracted from different subspaces, thereby improving the diversity and richness of feature representation. Secondly, dynamic weight allocation allows each attention head to adjust its weight in real time according to contextual information, so as to focus on important features more accurately. This mechanism enhances the ability of feature representation and enables the model to perform better in complex scenarios.

[0103] The embodiment of the present invention constructs a new domain adversarial transfer framework to achieve cross-domain feature transfer. By introducing the improved GAN (Generative Adversarial Network) and designing domain adversarial training technology, the feature distribution is made consistent, thereby enhancing the generalization ability of the model. By selecting features related to the target task and recalibrating them, the effect of transfer learning is improved, and the commonalities between the source domain and the target domain are captured. Knowledge is transferred from multiple source domains to make full use of feature information from different domains and further improve model performance.

[0104] In summary, the embodiments of the present invention reduce the training time of the steel defect detection model, improve the performance of the steel defect detection model, and thus make steel defect detection more accurate.

[0105] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0106] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0107] The above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A steel defect detection method using domain-adversarial adaptive multi-scale feature fusion, characterized in that: The method comprises: Step S1, collecting and inputting a first steel image training set, converting the training images in the first steel image training set into a graph structure by using a simple linear iterative clustering algorithm and multi-scale feature fusion, and obtaining initial node features of each node in the graph structure; Step S2: inputting the initial node features into a dynamic multi-head attention mechanism network, so that the dynamic multi-head attention mechanism enhances the feature representation capability of the initial node features and outputs corresponding enhanced node features; Step S3, determining a first source model that matches the enhanced node feature, and obtaining source domain data corresponding to the first source model, and determining the enhanced node feature as target domain data; wherein the first source model is a trained model; Step S4: using a generative adversarial network to generate generated data that matches the source domain data and the target domain data; using transfer learning and the generative adversarial network to discriminate the generated data from the source domain data, extract common features between the source domain data and the target domain data, and obtain a corresponding first loss function; Step S5: Based on the first loss function, a steel defect detection model corresponding to the first steel image training set is obtained by performing transfer learning on the first source model.

2. The steel defect detection method using domain-adversarial adaptive multi-scale feature fusion according to claim 1 is characterized in that: The step S1 comprises: Collecting and inputting the first steel image training set, and performing superpixel segmentation on the training images in the first steel image training set using a simple linear iterative clustering algorithm; and constructing a graph structure corresponding to the training images according to the superpixel segmentation results; The relationship between nodes in the graph structure is strengthened by the multi-scale feature fusion, thereby obtaining the initial node features of each node in the graph structure.

3. The steel defect detection method using domain-adversarial adaptive multi-scale feature fusion according to claim 1 is characterized in that: The node features include three primary color features, texture features, shape features and position features; The shape features include area and perimeter.

4. The steel defect detection method using domain-adversarial adaptive multi-scale feature fusion according to claim 1 is characterized in that: The step S2 comprises: Generate a query vector matrix, a key vector matrix, and a value vector matrix of the dynamic multi-head attention mechanism network according to the initial node features; According to the query vector matrix, the key vector matrix and the value vector matrix, the corresponding enhanced node features are output.

5. The steel defect detection method using domain-adversarial adaptive multi-scale feature fusion according to claim 4 is characterized in that: Outputting the corresponding enhanced node features according to the query vector matrix, the key vector matrix, and the value vector matrix includes: According to the query vector matrix, the key vector matrix and the value vector matrix, calculating the attention mechanism for each attention head to obtain an output vector of each attention head; The output vectors are connected and linearly transformed to obtain the corresponding enhanced node features.

6. The steel defect detection method using domain-adversarial adaptive multi-scale feature fusion according to claim 1 is characterized in that: The step S3 comprises: According to the target model corresponding to the requirements of the first steel image training set, multiple source models with similar fields and already completed training are determined; from the multiple original models, the most suitable first source model is determined; The first source model is used as an initial model for transfer learning, the data corresponding to the first source model is determined as source domain data for transfer learning, and the enhanced node features are determined as target domain data for transfer learning.

7. The steel defect detection method using domain-adversarial adaptive multi-scale feature fusion according to claim 1 is characterized in that: The step S4 comprises: Generate generated data that matches the source domain data and the target domain data using a generative adversarial network; An adversarial classifier is constructed, and the adversarial classifier is trained based on the source domain data, the target domain data, and the generated data, so as to obtain common features between the source domain data and the target domain data, and obtain a corresponding first loss function.

Citation Information

Cited By

  • Dislocation defect identification method and system based on image identification and deep learning network

    CN121545154A