Steel defect detection method based on synthetic data and CNN
By generating synthetic data and using ResNet-50 models, the problem of insufficient data in steel defect detection is solved, and efficient and accurate automatic visual detection is achieved.
Patent Information
- Application Number
- CN202510425853.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, steel defect detection relies on manual inspection and the automatic vision detection system lacks sufficient training data, resulting in unreliable and cumbersome detection results, and the performance of machine learning algorithms is limited.
Using a method based on synthetic data and CNN, synthetic data with steel surface defects is generated, and the steel defect detection model is constructed through the ResNet-50 model, and a manual synthetic data is trained, which solves the problem of insufficient data.
It improves the accuracy and efficiency of steel defect detection, reduces dependence on high-quality training data, and improves the performance of automatic visual detection systems.
Smart Images

Figure CN120259272A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of steel detection, and particularly relates to a steel defect detection method based on synthetic data and CNN. Background Art
[0002] Steel quality inspection has always been a recognized difficult problem in the past two decades. This process usually relies on the naked eye observation of manual inspectors for defect classification. However, due to human factors such as fatigue, this visual inspection method may result in classification errors, leading to unreliable and cumbersome detection results. In contrast, automatic visual inspection systems can provide better results than manual inspection due to their high reliability, clarity, and precision.
[0003] Machine learning algorithms are widely adopted in the field of computer vision, spanning numerous industries. A key feature of these algorithms is that they require a large amount of labeled data sets for training. The performance of machine learning-based automation systems is largely restricted by the quality of the training samples. These samples should be as reliable as possible to accurately reflect the characteristics of the research object, that is, they need to be representative. However, collecting such samples is a daunting task; it is necessary to capture as many variants of the object under investigation in different states as possible.
[0004] In most cases, developers of industrial automation systems lack sufficient production data to train machine learning algorithms. This situation may be due to companies not pre-recording the required parameters properly, or recording them incorrectly; it is difficult to automatically label production data, while manual labeling requires a high level of professional skills; data collection often takes months or even years. Therefore, these various limitations together exacerbate the complexity of implementing machine learning algorithms in technical process automation control systems. Currently, machine learning methods, as part of a steel surface inspection system, require a large number of defect images for training. This in turn increases the time required to collect and label the training data set. Summary of the Invention
[0005] Aiming at the above deficiencies in the prior art, the steel defect detection method based on synthetic data and CNN provided by the present invention solves the problem of the lack of original data in the automatic visual inspection system.
[0006] To achieve the above invention purpose, the technical solution adopted by the present invention is: a steel defect detection method based on synthetic data and CNN, comprising the following steps:
[0007] S1. Set up the environment for generating synthetic data to prepare for generating synthetic data with steel surface defects;
[0008] S2. Setting a shader of a steel model of synthetic data to generate synthetic data with steel surface defects;
[0009] S3, preprocessing the synthetic data to generate a training set and a validation set;
[0010] S4, inputting the training set and the validation set into the ResNet-50 model, training and testing the ResNet-50 model, and obtaining a steel defect detection model;
[0011] S5. Perform steel defect detection using a steel defect detection model.
[0012] Further: In S1, the method for setting the environment for generating synthetic data is specifically:
[0013] In Blender, a scene simulating the use of a camera to shoot steel is set up to simulate the shooting conditions of the steel model to get close to the actual industrial environment.
[0014] Further: in said S2, the surface defects include crack defects and rust defects, and the method for generating synthetic data with steel surface defects is specifically:
[0015] Assign corresponding shaders to the steel model according to the surface defect type, adjust the offset and noise parameters, set the surface defect type and occurrence frequency of the steel model, and obtain synthetic data of randomly distributed steel surface defects;
[0016] The shader that generates synthetic data with crack defects works as follows:
[0017] Dynamic geometric deformation is implemented in the vertex shader stage. The 3D mesh topology is reconstructed by importing the normal map displacement algorithm. At the same time, the basic noise texture is mapped into a directional elliptical crack shape by combining the non-uniform sampling technology of the spherical gradient field. The synthetic data with crack defects is generated by changing the UV coordinates and shape transformation.
[0018] The shader that generates synthetic data with rust defects works as follows:
[0019] Through the multi-band synthesis of improved Perlin noise and the physically based material attenuation model, the gradual process of the oxide layer from pitting corrosion to regional peeling is simulated to generate synthetic data with rust defects.
[0020] Further: In said S2, when generating synthetic data with steel surface defects, a corresponding mask is created, and the mask is used as a label of the classifier, wherein the method of creating the mask is specifically as follows:
[0021] The surface defect pixels of the steel model are marked as white, and the remaining pixels are saved as black.
[0022] Furthermore: Specifically, S3 is as follows:
[0023] Divide the synthetic data into synthetic data sets of cracks, rust, and defect-free according to the defect type, and divide each synthetic data set into training data and validation data according to a ratio of 7:3 to generate a training set and a validation set.
[0024] Furthermore: In S4, the ResNet-50 model includes a zero-padding block, a first-stage layer, a second-stage layer, a third-stage layer, a fourth-stage layer, a fifth-stage layer, and an output layer connected in sequence;
[0025] Among them, the structures of the first to fifth stage layers are the same, and each includes a convolutional block and an identity block connected to each other. The convolutional block is provided with three convolutional layers, and the identity block is provided with a batch normalization layer, an activation layer, and a max pooling layer connected in sequence at the connection points;
[0026] The output layer includes an average pooling layer, a flattening layer, and a fully connected layer connected in sequence.
[0027] Furthermore: The expression of the response Yi of the convolutional layer is specifically:
[0028] Yi = Wi * xi + Bi
[0029] In the formula, Wi is the weight of the i-th neuron, Bi is the bias, xi is the i-th input neuron, and Wi is updated through the following formula;
[0030] Wi(new) = Wi(old) - μ * C
[0031] In the formula, Wi(new) is the updated weight, Wi(old) is the weight before update, μ is the learning rate, and C is the cost function, and its expression is specifically:
[0032]
[0033] In the formula, y l represents the true value of the l-th sample, y l represents the predicted value of the l-th sample, and n is the number of samples.
[0034] Furthermore: In S4, the ResNet-50 model is tested through evaluation metrics, and the evaluation metrics include accuracy, recall rate, and Dice coefficient;
[0035] Among them, the expression of the accuracy Precision is specifically:
[0036]
[0037] Wherein, TP is the number of samples correctly predicted as the positive class by the ResNet-50 model, and FP is the number of samples incorrectly predicted as the positive class by the ResNet-50 model;
[0038] The expression of the recall rate is specifically:
[0039]
[0040] Wherein, TP is the number of samples correctly predicted as the positive class by the ResNet-50 model, and FN is the number of positive samples incorrectly predicted as the negative class by the ResNet-50 model;
[0041] The expression of the Dice coefficient Dice(X,Y) is specifically:
[0042]
[0043] Wherein, X is the predicted pixel set, and Y is the true meaning.
[0044] The beneficial effects of the present invention are:
[0045] (1) The present invention provides a steel defect detection method based on synthetic data and CNN, which uses artificially synthesized data for training and does not require a large amount of high-quality training data, solving the problem of lack of original data in the automatic vision detection system.
[0046] (2) The present invention uses the trained ResNet-50 model to construct a steel defect detection model, and uses the steel defect detection model to detect steel defects, which has stronger performance compared with traditional machine learning methods, improving the accuracy and efficiency of defect detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flowchart of a steel defect detection method based on synthetic data and CNN of the present invention.
[0048] Figure 2 is a schematic diagram of a cracked steel material.
[0049] Figure 3 is a schematic diagram of a rusty steel material.
[0050] Figure 4 is a schematic diagram of a defect-free steel material.
[0051] Figure 5 is a schematic diagram of the synthetic data generation process.
[0052] Figure 6 is a schematic diagram of a simple model of ResNet.
[0053] Figure 7 is a schematic diagram of the ResNet-50 model. Detailed implementation manners
[0054] The following describes the detailed implementation manners of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed implementation manners. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.
[0055] As Figure 1 shown, in an embodiment of the present invention, a steel defect detection method based on synthetic data and CNN includes the following steps:
[0056] S1. Set the environment for generating synthetic data to prepare for generating synthetic data with steel surface defects;
[0057] S2. Set the shader of the steel model for synthetic data to generate synthetic data with steel surface defects;
[0058] S3. Preprocess the synthetic data to generate a training set and a validation set;
[0059] S4. Input the training set and the validation set into the ResNet-50 model, train and test the ResNet-50 model to obtain a steel defect detection model;
[0060] S5. Perform steel defect detection through the steel defect detection model.
[0061] In the above S1, the method for setting the environment for generating synthetic data is specifically as follows:
[0062] Set a scene in Blender to simulate a camera shooting steel, and simulate the shooting conditions of the steel model to be close to the actual industrial environment.
[0063] In this embodiment, the steel model is shown as a parallelepiped with an overlay shader material. According to the input parameters passed to the shader, the texture displayed on the surface of the parallelepiped will change. Therefore, each new set of parameters can reproduce the unique manifestation of a specific type of defect. The disadvantage of this method is that a new shader needs to be created for each type of defect. Therefore, the present invention only considers three basic types of defects: cracks, rust, and no defects. Schematic diagrams of the three types of steel are as Figures 2 to 4 shown.
[0064] In a three-dimensional scene, the present invention simulates the shooting conditions of the steel model to approximate the actual industrial environment. For this purpose, a camera is installed directly above the steel model. At the same time, the light source in the scene is placed slightly above the camera, and such a layout mimics the lighting effect provided by the camera in a computer vision system.
[0065] In S2, the surface defects include crack defects and rust defects. The method for generating synthetic data with steel surface defects is specifically as follows:
[0066] Allocate corresponding shaders to the steel model according to the surface defect type, adjust the offset and noise parameters, set the surface defect type and occurrence frequency of the steel model, and obtain synthetic data with randomly distributed steel surface defects.
[0067] In this embodiment, by allocating specific material shaders to the steel model, the present invention sets its surface texture, defect types, and their occurrence frequencies. In this way, the defect types to be included in the synthetic data can be predefined before generating the dataset. The creation process of the image dataset is automated and is achieved by moving the steel model relative to the camera sequentially in the scene, which simulates the movement of the steel on a roller conveyor. In each iteration, the shader parameters are adjusted to generate defects of different shapes and positions, thus achieving a random distribution of surface defects. Finally, these surface images with defects are rendered.
[0068] The working process of the shader for generating synthetic data with crack defects is specifically as follows:
[0069] Implement dynamic geometric deformation in the vertex shader stage, perform three-dimensional mesh topology reconstruction by importing the normal map displacement algorithm, and at the same time combine the non-uniform sampling technology of the spherical gradient field to map the basic noise texture into a directional elliptical crack morphology, and generate synthetic data with crack defects by changing the UV coordinates and shape transformation. Among them, the construction method of the spherical gradient field is specifically as follows: establish an ellipsoidal coordinate system
[0070]
[0071] In the formula, a is the reference radius, ε is the control ellipticity, and ε ∈ [0.1, 0.8], θ is the polar angle, representing the angle with the positive z-axis, 0 ≤ θ ≤ π. When θ = 0 or π, it corresponds to the "north pole" or "south pole" of the ellipsoid; when θ = π / 2, it is located in the equatorial plane (xy plane), and at this time the radius reaches the maximum value. is the azimuth angle, representing the angle of rotation around the z-axis in the xy plane. In the formula The term controls the periodic change of the radius through the ellipticity ε, so that the cross-section shows elliptical characteristics in the azimuth angle direction.
[0072] The shader that generates synthetic data with rust defects works as follows:
[0073] Through the multi-band synthesis of improved Perlin noise and the physically based material attenuation model, the gradual process of the oxide layer from pitting corrosion to regional peeling is simulated to generate synthetic data with rust defects.
[0074] In this embodiment, the shader system is generated using a layered architecture design, and a dual-channel processing mechanism is built in the GPU rendering pipeline. In view of crack defects, dynamic geometric deformation is implemented in the vertex shader stage, and three-dimensional mesh topology reconstruction is achieved by importing the normal map displacement algorithm. At the same time, the non-uniform sampling technology of the spherical gradient field is combined to map the basic noise texture into a directional elliptical crack shape.
[0075] The rust generation module mainly operates at the fragment shader level, simulating the gradual process of the oxide layer from pitting corrosion to regional peeling through the multi-band synthesis of the improved Perlin noise and the physically based material attenuation model. The 3D Perlin noise algorithm is reconstructed and an anisotropic filter kernel is introduced, specifically a 3×3 Gaussian-Poisson hybrid filter.
[0076] In the process of generating surface defects of steel models, the present invention has written special shaders for two types of defects, cracks and rust. A program that sequentially transforms the original noise texture can shape various defect forms by adjusting the offset and noise parameters. For example, crack defects are generated by a programmed spherical gradient texture, which is elliptical and pre-deformed multiple times by changing the UV coordinates and the symmetry of the original pattern. The boundary of the crack will be displaced along the normal direction. The rust texture is generated by transforming the Perlin noise. The overall process is as follows: Figure 5 shown.
[0077] In S2, when generating synthetic data with steel surface defects, a corresponding mask is created, and the mask is used as a label of the classifier. The method of creating the mask is specifically as follows:
[0078] The surface defect pixels of the steel model are marked as white, and the remaining pixels are saved as black.
[0079] In this embodiment, the created mask is used as the annotation of the segmentation architecture training data, which is used for the annotation of the automated classifier, solving the annotation problem encountered when preparing real data. Since the data generation process of the data set is completely controllable, the present invention ensures the most uniform distribution between different defect types by predetermining the distribution of surface defect types within the sample, and normalizes the distribution of instances in the data set, which will improve the accuracy of defect classification and reduce the impact of the number of main instances of a single class in the sample on the overall classification result.
[0080] Specifically, S3 is as follows:
[0081] Divide the synthetic data into synthetic data sets of cracks, rust, and defect-free according to the defect type, and divide each synthetic data set into training data and validation data according to a ratio of 7:3 to generate a training set and a validation set.
[0082] In this embodiment, in order to make the input data conform to the input size of ResNet-50, adjust the size of the data and convert all colors to 224, 224, 3.
[0083] In S4, the ResNet-50 model includes a zero-padding block, a first-stage layer, a second-stage layer, a third-stage layer, a fourth-stage layer, a fifth-stage layer, and an output layer connected in sequence;
[0084] Among them, the structures of the first to fifth stage layers are the same, and each includes a convolutional block and an identity block connected to each other. The convolutional block is provided with three convolutional layers, and the identity block is provided with a batch normalization layer, an activation layer, and a max pooling layer connected in sequence at the connection points;
[0085] The output layer includes an average pooling layer, a flattening layer, and a fully connected layer connected in sequence.
[0086] In ResNet, the input of the nth l layer is directly connected to the (n l +x)th layer through a shortcut connection. In this way, even if the gradient disappears, the gradient can be passed back to earlier layers through the identity mapping. Conventional convolution operations will reduce the spatial resolution. For example, an image of 32×32 becomes 30×30 after a 3×3 convolution. In order to match the number of channels, a linear projection W is needed to expand the channels of the shortcut connection. Finally, the original input x and the learned F(x) are combined as the input of the next layer. Figure 6 This is a simple model of ResNet used in the present invention. Training a network in this form is easier than training a simple deep convolutional neural network and solves the problem of accuracy degradation.
[0087] As Figure 7 shown, the ResNet-50 model consists of five stages, and each stage contains a convolutional block and an identity block. Each convolutional block has three convolutional layers, and each identity block also has three convolutional layers. ResNet-50 has more than 23 million trainable parameters.
[0088] The specific expression of the response Yi of the convolutional layer is:
[0089] Yi = Wi * xi + Bi
[0090] Wherein, Wi is the weight of the i-th neuron, Bi is the bias, xi is the i-th input neuron, and Wi is updated through the following formula;
[0091] Wi(new) = Wi(old) - μ * C
[0092] Wherein, Wi(new) is the updated weight, Wi(old) is the weight before update, μ is the learning rate, and C is the cost function, and its expression is specifically:
[0093]
[0094] Wherein, y l represents the true value of the l-th sample, and y l represents the predicted value of the l-th sample, and n is the number of samples.
[0095] After preprocessing and training the network, SVM (Support Vector Machine) is used to classify the data, and the specific method is as follows:
[0096] In this embodiment, K(K - 1) / 2 binary support vector machine models are used, and a one-versus-one coding design is adopted, where K is the number of unique class labels. The error-correcting output coding model is used to reduce the classification problem of three or more classes to a set of binary classifiers. For each binary learner, one class is the positive class and the other is the negative class, and the rest are ignored. This design includes all combinations of class pair assignments.
[0097] In step S4, the ResNet-50 model is tested through evaluation metrics, and the evaluation metrics include precision, recall, and Dice coefficient;
[0098] Among them, the expression of precision Precision is specifically:
[0099]
[0100] Wherein, TP is the number of samples correctly predicted as the positive class by the ResNet-50 model, and FP is the number of samples incorrectly predicted as the positive class by the ResNet-50 model;
[0101] Precision is a commonly used metric in information retrieval and machine learning, which measures the proportion of samples that are actually positive among all samples predicted as positive by the model.
[0102] The expression of recall is specifically:
[0103]
[0104] Wherein, TP is the number of samples correctly predicted as positive by the ResNet-50 model, and FN is the number of positive samples wrongly predicted as negative by the ResNet-50 model;
[0105] Recall rate is an indicator to measure the performance of a classification model. Especially in binary classification problems, it represents the proportion of positive class samples correctly identified by the model among all actual positive class samples.
[0106] The expression of the Dice coefficient Dice(X,Y) is specifically:
[0107]
[0108] Wherein, X is the predicted pixel set, and Y is the true meaning.
[0109] The Dice coefficient is used to compare the pixel matching between the predicted segmentation and the corresponding ground truth.
[0110] The beneficial effects of the present invention are as follows: The present invention provides a steel defect detection method based on synthetic data and CNN. By using artificially synthesized data for training, a large amount of high-quality training data is not required, solving the problem of lack of original data in the automatic vision detection system.
[0111] The present invention uses the trained ResNet-50 model to construct a steel defect detection model, and adopts the steel defect detection model for steel defect detection. Compared with traditional machine learning methods, it has stronger performance and improves the accuracy and efficiency of defect detection.
[0112] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of technical features. Therefore, the features defined by "first", "second", "third" may explicitly or implicitly include one or more of such features.
Claims
1. A steel defect detection method based on synthetic data and CNN, characterized in that, It includes the following steps: S1. Set up the environment for generating synthetic data to prepare for generating synthetic data with steel surface defects; S2. Set the shader of the steel model for synthetic data to generate synthetic data with steel surface defects; S3. Preprocess the synthetic data to generate a training set and a validation set; S4. Input the training set and the validation set into the ResNet-50 model, train and test the ResNet-50 model to obtain a steel defect detection model; S5. Detect steel defects through the steel defect detection model.
2. The steel defect detection method based on synthetic data and CNN according to claim 1, characterized in that In S1, the method for setting up the environment for generating synthetic data is specifically as follows: Set up a scene in Blender to simulate a camera shooting steel, and simulate the shooting conditions of the steel model to be close to the actual industrial environment.
3. The steel defect detection method based on synthetic data and CNN according to claim 1, characterized in that In S2, the surface defects include crack defects and rust defects. The method for generating synthetic data with steel surface defects is specifically as follows: Allocate corresponding shaders to the steel model according to the surface defect type, adjust the offset and noise parameters, and set the surface defect type and occurrence frequency of the steel model to obtain synthetic data with randomly distributed steel surface defects; The working process of the shader for generating synthetic data with crack defects is specifically as follows: Implement dynamic geometric deformation in the vertex shader stage, perform three-dimensional mesh topology reconstruction by importing the normal map displacement algorithm, and at the same time combine the non-uniform sampling technology of the spherical gradient field to map the basic noise texture into a directional elliptical crack morphology, and generate synthetic data with crack defects by changing the UV coordinates and shape transformation; The working process of the shader for generating synthetic data with rust defects is specifically as follows: Generate synthetic data with rust defects through the multi-band synthesis of improved Perlin noise, combined with a physically based material attenuation model to simulate the progressive process of the oxide layer from pitting corrosion to regional spalling.
4. The steel defect detection method based on synthetic data and CNN according to claim 3, characterized in that In S2, when generating synthetic data with steel surface defects, create a corresponding mask, and the mask is used as the annotation of the classifier. Among them, the method for creating the mask is specifically as follows: Mark the surface defect pixels of the steel model as white, and save the remaining pixels as black.
5. The steel defect detection method based on synthetic data and CNN according to claim 1, characterized in that S3 is specifically as follows: Divide the synthetic data into synthetic data sets of cracks, rusts, and defect-free according to the defect type, and divide each synthetic data set into training data and validation data according to a ratio of 7:3 to generate a training set and a validation set.
6. The steel defect detection method based on synthetic data and CNN according to claim 1, characterized in that, In S4, the ResNet-50 model includes a zero-padding block, a first-stage layer, a second-stage layer, a third-stage layer, a fourth-stage layer, a fifth-stage layer, and an output layer connected in sequence; Among them, the structures of the first to fifth stage layers are the same, and each includes a convolutional block and an identity block connected to each other. The convolutional block is provided with three convolutional layers, and the identity block is provided with a batch normalization layer, an activation layer, and a max pooling layer connected in sequence at the connection points; The output layer includes an average pooling layer, a flattening layer, and a fully connected layer connected in sequence.
7. The steel defect detection method based on synthetic data and CNN according to claim 6, wherein, The expression of the response Yi of the convolutional layer is specifically: Yi = Wi * xi + Bi In the formula, Wi is the weight of the i-th neuron, Bi is the bias, xi is the i-th input neuron, and Wi is updated through the following formula; Wi(new) = Wi(old) - μ * C Where Wi(new) is the updated weight, Wi(old) is the weight before update, μ is the learning rate, and C is the cost function, and its specific expression is: where y l represents the true value of the \(l\)-th sample, and \(\hat{y}\) l represents the predicted value of the \(l\)-th sample, and \(n\) is the number of samples.
8. The steel defect detection method based on synthetic data and CNN according to claim 1, characterized in that In S4, the ResNet-50 model is tested through evaluation metrics, and the evaluation metrics include precision, recall, and Dice coefficient; Among them, the specific expression of precision Precision is: Where TP is the number of samples correctly predicted as positive by the ResNet-50 model, and FP is the number of samples incorrectly predicted as positive by the ResNet-50 model; The specific expression of recall is: Where TP is the number of samples correctly predicted as positive by the ResNet-50 model, and FN is the number of positive samples incorrectly predicted as negative by the ResNet-50 model; The specific expression of Dice coefficient Dice(X, Y) is: Where X is the predicted pixel set and Y is the true meaning.
Citation Information
Cited By
Industrial workpiece surface defect detection method and system based on TSK fuzzy mapping
CN122115331A