An industrial defect segmentation method based on neural network architecture search
Through the method based on neural network architecture search, the weight initialization and structure of convolutional neural networks, Transformer and multi-layer perceptron are optimized, which solves the problem of insufficient accuracy, efficiency and adaptability in industrial defect segmentation, and achieves high-precision and efficient defect detection.
Patent Information
- Application Number
- CN202411261011.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-09-10
AI Technical Summary
The prior art has problems of insufficient accuracy, efficiency and adaptability in industrial defect segmentation, especially the limited receptive field of convolutional neural networks leads to discontinuous defect segmentation, Transformer lacks local information exchange, and multi-layer perceptrons are prone to overfitting.
Using a neural network architecture search method, the weight initialization of convolutional neural network, Transformer and multi-layer perceptron is shared with each other through a cross-size weight sharing strategy, a search space is built and a hypernetwork is trained, and the optimal subnet is finally obtained for defect segmentation.
By optimizing the network structure, the accuracy of surface defect detection of industrial products is significantly improved, training time and resource consumption is reduced, and various types of defects can be accurately detected and divided.
Smart Images

Figure CN119205810B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of industrial defect segmentation, and particularly to an industrial defect segmentation method based on neural network architecture search. Background Art
[0002] Industrial product surface defect segmentation is a key task in various industrial applications. The main challenge lies in the variability of surface defects, whose shapes and sizes can vary greatly. Modern detection technologies, such as machine vision, deep learning, and automated detection systems, can significantly improve the speed and accuracy of detection. However, common deep learning-based methods, such as convolutional neural networks (CNNs), usually have difficulty modeling long-range dependencies due to their limited receptive fields, which can lead to discontinuities in defect segmentation. Transformers, on the other hand, lack a local mechanism for information exchange within local regions, which may cause local details of defects (such as edges and shapes) to be lost. In addition, the high flexibility and large capacity of multi-layer perceptrons (MLPs) make them prone to overfitting during training, especially when the number of samples in the defect dataset is small.
[0003] Given the respective advantages and disadvantages of the above architectures, no single architecture can be suitable for all defect segmentation tasks. A direct solution is to design a hybrid network architecture. In recent years, neural network architecture search (NAS) methods have been used to automatically design the most effective architectures. However, due to the characteristics of different operators, most network architecture search-based methods fail to efficiently combine these operators, and as the number of operators increases, the search efficiency drops significantly, which is impractical in actual applications. Another difficulty is how to fuse multi-level features. Research shows that multi-level features are crucial for surface defect segmentation. High-level features capture discriminative information, while low-level features contain rich texture information. Therefore, it is necessary to develop an effective neural network architecture search-based method for defect detection on industrial product surfaces. Summary of the Invention
[0004] Based on this, it is necessary to provide an industrial defect segmentation method based on neural network architecture search, which includes:
[0005] S1: Preprocess the industrial product surface defect dataset with pixel-level annotations;
[0006] S2: Make the convolutional neural network, Transformer, and multi-layer perceptron with weight initialization share weights with each other at different sizes through a cross-size weight sharing strategy; construct a search space, and put the convolutional neural network, Transformer, and multi-layer perceptron after weight sharing processing into the search space; the search space is used to search for the three networks simultaneously;
[0007] S3: Construct a super network based on all combinations of the three networks in the search space and train the super network;
[0008] S4: Perform evolutionary search on the trained super network to obtain the optimal sub-network;
[0009] S5: Train the optimal sub-network based on the preprocessed industrial product surface defect dataset, and construct a defect segmentation model based on the trained optimal sub-network; Input the image to be detected into the defect segmentation model to output the pixel-level segmentation result.
[0010] Beneficial effects: By optimizing the network structure, this method reduces the training time and resource consumption, and at the same time significantly improves the accuracy of industrial product surface defect detection, and can accurately detect and segment various types of defects; This method solves the deficiencies of traditional detection methods in terms of accuracy, efficiency and adaptability, and provides a more effective and reliable solution for industrial product surface defects. Description of the Drawings
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1 It is a flowchart of the industrial defect segmentation method based on neural network architecture search according to the embodiment of the present application.
[0013] Figure 2 It is a flowchart of the operation of the defect segmentation model according to the embodiment of the present application. Detailed Embodiments
[0014] In order to make the above-mentioned objects, features and advantages of the present application more obvious and understandable, the following will make a detailed description of the specific embodiments of the present application in conjunction with the drawings. Many specific details are set forth in the following description in order to fully understand the present application. However, the present application can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.
[0015] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0016] like Figure 1 As shown, this embodiment provides an industrial defect segmentation method based on neural network architecture search, the method comprising:
[0017] S1: Preprocessing of the pixel-level annotated industrial product surface defect dataset.
[0018] Specifically, the step includes:
[0019] S1.1: Divide the industrial product surface defect dataset into a training set (60%), a validation set (20%) and a test set (20%);
[0020] S1.2: Perform tensor transformation and normalization on the training set, validation set, and test set in turn to obtain the preprocessed training set, preprocessed validation set, and preprocessed test set, respectively.
[0021] In this embodiment, the pixel-level annotated industrial product surface defect dataset uses the magnetic tile surface defect dataset MTDD, which contains a total of 1344 images. The dataset contains five types of defects: blowhole, break, crack, fray, and unevenness. For training and testing, this example only selects defect images from the data for experiments (a total of 782 images), and resizes all images to 320×320 pixels.
[0022] S2: Through the cross-size weight sharing strategy, the weight-initialized convolutional neural network, Transformer, and multi-layer perceptron share weights with each other at different sizes; construct a search space, and put the convolutional neural network, Transformer, and multi-layer perceptron processed with weight sharing into the search space; the search space is used to search the three networks simultaneously.
[0023] Specifically, in the search space, given a standard convolution, it can be described as Where k represents the size of the convolution kernel. Assume that the input feature map Output feature map Where h and w represent the width and height of the feature map respectively, and the calculation process of the convolutional neural network is expressed as:
[0024]
[0025] S(x i,j ,Δm,Δn)=x i+Δm,j+Δn ;
[0026] Among them, y i,j' represents the value of the first output feature map at (i, j); a, b represent the convolutional kernel at (a, b), where a, b ∈ {0, 1, …, k - 1} and k represents the size of the convolutional kernel; S(·) represents the Shift function; k (a,b) represents the convolutional kernel weight at (a, b); represents the feature obtained by passing the value of the input feature map at (i, j) through the convolutional kernel at (a, b); x i,j represents the value of the input feature map at (i, j); Δm represents the horizontal displacement; Δn represents the vertical displacement; x i+Δm,j+Δn represents the feature value after two steps of displacement of Δm and Δn at (i, j) in the input feature map;
[0027] There are H heads in the Transformer (self-attention). Let the input feature map output feature map where h and w represent the width and height of the feature map respectively. The calculation process of the Transformer is expressed as:
[0028] y i,j ″ = Concat(x i,j head (1) , …, x i,j head (H) )W O ;
[0029]
[0030] where, y i,j ″ represents the value of the second output feature map at (i, j); Concat(·) represents the concatenation operation; W O represents the first linear projection parameter; x i,j head (1) represents the feature obtained by passing the value of the input feature map at (i, j) through the first head; x i,j head (H) represents the feature obtained by passing the value of the input feature map at (i, j) through the H-th head; x i,j head (l) represents the feature obtained by passing the value of the input feature map at (i, j) through the l-th head; Attention(·) represents the self-attention mechanism; represents the second linear projection parameter; represents the third linear projection parameter; represents the fourth linear projection parameter; Q represents the query; K represents the key; V represents the value; softmax(·) represents the softmax function; T represents the transpose; d K represents the feature dimension of the key;
[0031] Let the input feature map Output feature map Where h and w represent the width and height of the feature map respectively. Using the spatial shift paradigm in S2 MLPv2 (an advanced MLP structure), the calculation process of the multi-layer perceptron is expressed as:
[0032]
[0033] Among them, y i,j ″′ represents the value of the third output feature map at (i, j); MLP(·) represents the multi-layer perceptron; SA(·) represents recalibrating different branches through the multi-layer perceptron; Indicates shifting each element in the value of the third output feature map at (i, j) down by one position; Indicates shifting each element in the value of the third output feature map at (i, j) down by two positions; Indicates shifting each element in the value of the third output feature map at (i, j) down by three positions; X represents the input feature map; Represents the input feature map calibrated by the multi-layer perceptron.
[0034] In this embodiment, after converting the above three networks into the same format, they are put into a unified weight-sharing search space. The three networks first share weights through three 1×1 convolutions, and then the encoder in the defect segmentation model selects different network combinations according to the paradigm. The structure of the search space is shown in Table 1;
[0035] Table 1 is the structure table of the search space;
[0036]
[0037] As shown in Table 1, the search space includes four stages. Each stage samples the three networks respectively, and the depth, number of channels, feedforward neural network ratio, and model size of the three networks in each stage are different.
[0038] Furthermore, the cross-size weight-sharing strategy includes:
[0039] For the convolutional neural network, in each model size of each stage, the remaining convolutional kernels except the largest convolutional kernel inherit the weights of the largest convolutional kernel; The first inheritance relationship is expressed as:
[0040]
[0041] Among them, Represents the weight of the convolutional kernel S in the i-th layer of the convolutional neural network; Denote the weight of the largest convolutional kernel \(L\) in the \(i\)-th layer of the convolutional neural network; \((p, q)\) represents the starting position of the inheritance process; \(k\) S Denote the size of the convolutional kernel \(S\); \(k\) L Denote the size of the largest convolutional kernel;
[0042] For Transformer and multi-layer perceptron, in each model size at each stage, make the remaining blocks except the largest block inherit the weight of the largest block; the second inheritance relationship is expressed as:
[0043]
[0044] Among them, Denote the weight of the block \(S'\) in the \(i\)-th layer of the Transformer or multi-layer perceptron; Denote the weight of the largest block \(L'\) in the \(i\)-th layer of the Transformer or multi-layer perceptron; \(c\) S, Denote the embedding dimension of the block \(S'\).
[0045] S3: Construct a super network based on all combinations of the three networks in the search space, and train the super network.
[0046] Specifically, take each combination of the three networks in the search space as a subnet in the super network; the training of the super network includes:
[0047] Step 1: Use AdamW as the optimizer, set the initial first learning rate to 0.0004, the first learning rate adopts the cosine decay strategy, perform data augmentation operations on the images in the preprocessed training set, and set the training period to 2000 epochs; the data augmentation operations include random horizontal flipping, random resizing, and random rotation;
[0048] Step 2: Randomly sample the subnets, and only update the weights of the sampled subnets during each training;
[0049] Step 3: For the convolutional neural network in the subnet, only train the largest convolutional kernel, and for the Transformer and multi-layer perceptron in the subnet, only train the largest block;
[0050] Step 4: During the training process, use the preprocessed validation set to evaluate the performance of the sampled subnets.
[0051] S4: Perform evolutionary search on the trained super network to obtain the optimal subnetwork.
[0052] Specifically, in S4, the evolutionary search on the trained super network includes:
[0053] Step 1: Obtain a trained supernetwork, where the trained supernetwork includes multiple subnets; the search space includes parameter combinations of multiple subnets;
[0054] Step 2: Randomly initialize a population, where each individual represents a parameter combination of a subnet in the search space;
[0055] Step 3: In each round of iteration, select the top k subnets with the average intersection over union ratio as the parents;
[0056] When the preset crossover probability is satisfied, perform a crossover operation on these parents to generate new offspring;
[0057] When the preset mutation probability is satisfied, perform a mutation operation on the offspring of these parents to generate new offspring;
[0058] Step 4: The generated offspring will be added to the population;
[0059] Step 5: Repeat Steps 3 - 4 until a predetermined number of iterations is reached, and select the parameter combination of the subnet with the largest average intersection over union ratio in the population as the optimal solution;
[0060] Step 6: Extract a subnet that meets the parameter requirements from the supernetwork according to the searched optimal solution as the optimal subnetwork.
[0061] S5: Train the optimal subnetwork based on the pre - processed industrial product surface defect dataset, and construct a defect segmentation model based on the trained optimal subnetwork; input the image to be detected into the defect segmentation model to output a pixel - level segmentation result.
[0062] In this embodiment, training the optimal subnetwork based on the pre - processed industrial product surface defect dataset includes:
[0063] Based on the pre - processed training set, use Adam as the optimizer, set the initial second learning rate to 0.0001, and use a multi - learning rate scheduler, and set the training period to 500 epochs;
[0064] After training, perform testing based on the pre - processed test set.
[0065] Furthermore, constructing a defect segmentation model based on the trained optimal subnetwork includes:
[0066] Step 1: Use the trained optimal subnetwork as the encoder of the defect segmentation model;
[0067] Step 2: Design a multi - level feature aggregation module as the decoder of the defect segmentation model;
[0068] The multi-level feature aggregation module refines the multi-level features extracted by the encoder using a variety of feature extraction operations; the refinement process is a directed acyclic graph including N F nodes and N E edges, where each node represents a feature map and each directed edge represents a candidate operation; the types of candidate operations include 1×1 convolution, 3×3 convolution, 5×5 convolution, 3×3 dilated convolution, and 3×3 depthwise separable convolution; after refinement, a pyramid pooling module is adopted, and the pyramid pooling module is used to fuse the refined multi-level features and output them.
[0069] In this embodiment, the defect segmentation model includes four stages, and the parameters in the defect segmentation model include: the depth of each stage, the expansion ratio of the feed-forward neural network in each layer of each stage, the number of heads of the feature extraction in each layer of each stage, the network type adopted in each layer of each stage, the input channel number of each stage, the type of candidate operation in the multi-level feature aggregation module of each stage, and the scale of the pooling layer in the pyramid pooling module.
[0070] Further, the specific description of the parameters is as follows:
[0071] 1. Depth: It represents the depth of each stage in the network, that is, the number of layers included in each stage. Specifically, the network includes four stages, with 3, 5, 6, and 4 layers respectively in each stage;
[0072] 2. FFN_ratio: It represents the expansion ratio of the feed-forward neural network in each layer of each stage in the network. Specifically, the network includes four stages. The first stage has three layers, and the FFN ratios are 7.5, 7.5, and 8.0 respectively; the second stage has five layers, and the FFN ratios are 8.0, 8.0, 8.5, 8.0, and 7.5 respectively; the third stage has six layers, and the FFN ratios are 4.5, 4.5, 3.5, 3.5, 3.5, and 4.5 respectively; the fourth stage has four layers, and the FFN ratios are 4.5, 4.0, 3.5, and 4.0 respectively;
[0073] 3. Num_heads: It represents the number of heads in the feature extraction block in each layer of each stage in the network. Specifically, the network includes four stages. The first stage has three layers, and the number of heads is 1, 1, and 1 respectively; the second stage has five layers, and the number of heads is 3, 3, 2, 2, and 2 respectively; the third stage has six layers, and the number of heads is 6, 5, 5, 5, 6, and 5 respectively; the fourth stage has four layers, and the number of heads is 8, 7, 7, and 8 respectively;
[0074] 4. Operation: Represents the type of operation adopted by each layer in each stage of the network. Among them, 0 represents the MLP operation, 1 represents the Transformer operation, 3 represents the 3×3 convolution, and 5 represents the 5×5 convolution. Specifically, the network consists of four stages. The first stage has three layers, and the operation types are 1, 0, and 1 respectively; the second stage has five layers, and the operation types are 0, 0, 1, 1, and 0 respectively; the third stage has six layers, and the operation types are 0, 3, 1, 3, 0, and 5 respectively; the fourth stage has four layers, and the operation types are 1, 0, 5, and 1 respectively;
[0075] 5. Channel: Represents the number of input channels in each stage. Specifically, the network consists of four stages. The first stage has 64 input channels, the second stage has 128 input channels, the third stage has 384 input channels, and the fourth stage has 448 input channels;
[0076] 6. Conv_choice: Represents the candidate operations in the multi-level feature aggregation module. Among them, 0 represents the 1×1 convolution, 1 represents the 3×3 convolution, 2 represents the 5×5 convolution, 3 represents the 3×3 dilated convolution, and 4 represents the 3×3 depthwise separable convolution. Specifically, the network consists of four stages. Among them, the operation types in the first stage are 2 and 1, the operation types in the second stage are 1 and 3, the operation types in the third stage are 4 and 2, and the operation types in the fourth stage are 0, 1, 1, and 2;
[0077] 7. Pool_scale: Corresponds to the scale of the pooling layer in the pyramid pooling module PPM, which can be a pooling operation of 1×1, 2×2, 3×3, or 6×6.
[0078] As Figure 2 shown, this embodiment shows the operation process of the defect segmentation model, including:
[0079] Step 1: The encoder obtains the image to be detected. The encoder internally performs linear projection processing on the image to be detected, and then performs Transformer operations, convolution operations, and MLP operations respectively. The features after the three operations pass through the feed-forward neural network to obtain multi-level features;
[0080] Step 2: The decoder obtains the multi-level features. The multi-level feature aggregation module uses various feature extraction operations to refine the multi-level features. The pyramid pooling module cascades and upsamples the refined multi-level features in sequence, and outputs the pixel-level segmentation result after upsampling.
[0081] This industrial defect segmentation method based on neural network architecture search provided by this embodiment has the following
[0082] beneficial effects:
[0083] This method optimizes the network structure, reduces the training time and resource consumption, and at the same time significantly improves the accuracy of industrial product surface defect detection, and can accurately detect and segment various types of defects; this method solves the deficiencies of traditional detection methods in terms of accuracy, efficiency, and adaptability, and provides a more effective and reliable solution for industrial product surface defects.
[0084] The network architecture of this method searches for an end-to-end segmentation network architecture by jointly searching for CNN, Transformer, and MLP operators, and can adapt to the detection and segmentation of various defect types.
[0085] CNN is good at processing local features in images, such as edges, textures, and local patterns, which makes it excellent in detecting the nuances in industrial defects; Transformer can capture global information and long-range dependencies in images through its self-attention mechanism, giving it a significant advantage in dealing with complex global features and long-range dependency problems; MLP provides a simple and efficient way to process numerical features and some linear relationships. By combining these three operators, this method can not only effectively process various local and global features in images, but also flexibly adapt to multiple defect types in different industrial scenarios.
[0086] At the same time, the network directly uses the segmentation accuracy as an indicator during the search process. After training is completed, there is no need to adapt to other segmentation frameworks, and the segmentation result can be obtained by inputting the detection image. This comprehensiveness and adaptability enable this method to maintain high efficiency and accuracy in various industrial defect detection tasks, thus improving the overall detection performance and reliability.
[0087] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0088] The above-described embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. An industrial defect segmentation method based on neural network architecture search, characterized in that: include: S1: Preprocessing of the pixel-level annotated industrial product surface defect dataset; S2: Through the cross-size weight sharing strategy, the weight-initialized convolutional neural network, Transformer, and multi-layer perceptron share weights with each other at different sizes; construct a search space and put the convolutional neural network, Transformer, and multi-layer perceptron processed with weight sharing into the search space; The search space is used to search three networks simultaneously; The cross-size weight sharing strategy includes: For convolutional neural networks, in each model size at each stage, the remaining convolution kernels except the largest convolution kernel inherit the weight of the largest convolution kernel; the first inheritance relationship is expressed as: ; in, Represents the convolutional neural network i The weight of the convolution kernel S in the layer; Represents the convolutional neural network i The weight of the largest convolution kernel L in the layer; ( p , q ) indicates the starting position of the inheritance process; Represents the size of the convolution kernel S; Indicates the size of the maximum convolution kernel; For Transformer and multi-layer perceptron, in each model size at each stage, all blocks except the largest block inherit the weight of the largest block; the second inheritance relationship is expressed as: ; in, Represents Transformer or Multilayer Perceptron i Block in layer The weight of Represents Transformer or Multilayer Perceptron i Largest block in the layer The weight of Representation Block The embedding dimension of S3: construct a super network based on all combinations of the three networks in the search space and train the super network; S4: performing evolutionary search on the trained super network to obtain the optimal sub-network; S5: training the optimal sub-network based on the preprocessed industrial product surface defect data set, and constructing a defect segmentation model based on the trained optimal sub-network; inputting the image to be detected into the defect segmentation model to output a pixel-level segmentation result.
2. The industrial defect segmentation method based on neural network architecture search according to claim 1 is characterized in that: S1 includes: S1.1: Divide the industrial product surface defect dataset into a training set, a validation set and a test set; S1.2: Perform tensor transformation and normalization on the training set, validation set, and test set in turn to obtain the preprocessed training set, preprocessed validation set, and preprocessed test set, respectively.
3. The industrial defect segmentation method based on neural network architecture search according to claim 1 is characterized in that: In the search space, the calculation process of the convolutional neural network is expressed as: ; ; in, Indicates that the first output feature map is in ( i , j) The value at represents the convolution kernel at (a,b), , k Indicates the size of the convolution kernel; Represents the Shift function; Represents the convolution kernel weight at (a, b); Indicates that the input feature map is in ( i , j) The value at is passed through the convolution kernel at (a, b) to obtain the feature; Indicates that the input feature map is in ( i , j ) at the value; represents horizontal displacement; represents vertical displacement; Represented in the input feature map ( i , j )go through and Eigenvalues after two-step shift; The calculation process of Transformer is expressed as: ; ; ; in, Indicates that the second output feature map is in ( i , j) The value at Represents a splicing operation; represents the first linear projection parameters; Indicates that the input feature map is in ( i , j) The value at is the feature obtained by passing through the first head; Indicates that the input feature map is in ( i , j) The value at is the feature obtained by passing through the Hth head; Indicates that the input feature map is in ( i , j) The value at l Characteristics obtained by individual head; Represents the self-attention mechanism; represents the second linear projection parameters; represents the third linear projection parameters; represents the fourth linear projection parameter; Q represents query; K represents key; V represents value; represents the softmax function; T represents transposition; The characteristic dimension of the representation key; The calculation process of the multilayer perceptron is expressed as: ; ; in, Indicates that the third output feature map is in ( i , j) The value at represents a multi-layer perceptron; Represents recalibration of different branches through multi-layer perceptron; Indicates that the third output feature map is in ( i , j) Each element in the value at is shifted down one position; Indicates that the third output feature map is in ( i , j) Each element in the value at is shifted down two positions; Indicates that the third output feature map is in ( i , j) Each element in the value at is shifted down three positions; Represents the input feature map; Represents the input feature map after multi-layer perceptron calibration.
4. The industrial defect segmentation method based on neural network architecture search according to claim 1 is characterized in that: The search space includes four stages, each stage samples three networks respectively, and the depth, number of channels, feedforward neural network ratio and model size of the three networks in each stage are different.
5. The industrial defect segmentation method based on neural network architecture search according to claim 2 is characterized in that: In S3, each combination of the three networks in the search space is used as a subnet in the super network; the training super network includes: Step 1: AdamW is used as the optimizer, the initial first learning rate is set to 0.0004, the first learning rate adopts a cosine decay strategy, and data enhancement operations are performed on the images in the preprocessed training set. The training cycle is set to 2000 epochs; the data enhancement operations include random horizontal flipping, random resizing, and random rotation; Step 2: Randomly sample the subnetwork, and only update the weight of the sampled subnetwork during each training; Step 3: For the convolutional neural network in the subnet, only the largest convolution kernel is trained; for the Transformer and multi-layer perceptron in the subnet, only the largest block is trained; Step 4: During the training process, the performance of the sampled subnetwork is evaluated using the preprocessed validation set.
6. The industrial defect segmentation method based on neural network architecture search according to claim 2 is characterized in that: In S4, performing evolutionary search on the trained super network includes: Step 1: Obtain a trained super network, wherein the trained super network includes multiple sub-networks; the search space includes parameter combinations of the multiple sub-networks; Step 2: Randomly initialize a population, where each individual represents a parameter combination of a subnetwork in the search space; Step 3: In each iteration, select the top k subnetworks with the highest average intersection-to-union ratio as the parent network; When the preset crossover probability is met, these parents are crossovered to generate new offspring; When the preset mutation probability is met, mutation operations are performed on the offspring of these parents to generate new offspring; Step 4: The generated offspring will be added to the population; Step 5: Repeat steps 3-4 until the predetermined number of iterations is reached, and select the parameter combination of the subnet with the largest average intersection-to-union ratio in the population as the optimal solution; Step 6: extracting a subnet that meets the parameter requirements from the super network according to the searched optimal solution as the optimal subnet.
7. The industrial defect segmentation method based on neural network architecture search according to claim 6 is characterized in that: The training of the optimal sub-network based on the pre-processed industrial product surface defect data set includes: Based on the preprocessed training set, Adam is used as the optimizer, the initial second learning rate is set to 0.0001, and a multi-learning rate scheduler is used, and the training cycle is set to 500 epochs; After training, testing is performed based on the preprocessed test set.
8. The industrial defect segmentation method based on neural network architecture search according to claim 7 is characterized in that: In S5, constructing a defect segmentation model based on the trained optimal sub-network includes: Step 1: Use the trained optimal sub-network as the encoder of the defect segmentation model; Step 2: Design a multi-level feature aggregation module as the decoder of the defect segmentation model; The multi-level feature aggregation module uses multiple feature extraction operations to refine the multi-level features extracted by the encoder; the refinement process includes nodes and The invention provides a directed acyclic graph with edges, wherein each node represents a feature graph, and each directed edge represents a candidate operation; the types of candidate operations include 1×1 convolution, 3×3 convolution, 5×5 convolution, 3×3 dilated convolution, and 3×3 depth-separable convolution; after refinement, a pyramid pooling module is used to fuse and output the refined multi-level features.
9. The industrial defect segmentation method based on neural network architecture search according to claim 8 is characterized in that: The defect segmentation model includes four stages, and the parameters in the defect segmentation model include: the depth of each stage, the expansion ratio of the feedforward neural network of each layer in each stage, the number of feature extraction heads of each layer in each stage, the network type used in each layer in each stage, the number of input channels of each stage, the type of candidate operations in the multi-level feature aggregation module of each stage, and the scale of the pooling layer in the pyramid pooling module.
Citation Information
Patent Citations
Surface defect detection method based on differentiable neural architecture search
CN117173091A
Neural network structure searching method and device and storage medium
CN117688984A