An automated blind detection target dynamic classification method

By introducing a differentiable incremental branch module and a differentiable search strategy in automated blind inspection, the problem of inefficiency in training of CNN when dynamic categories increase is solved, and rapid adaptation and efficient classification are achieved.

CN115908901BActive Publication Date: 2025-07-18GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211355872.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-01
Publication Date
2025-07-18
Estimated Expiration
2042-11-01

AI Technical Summary

Technical Problem

Existing convolutional neural networks (CNNs) are difficult to adapt to the dynamically increased target categories in automated blind inspection, resulting in the need for retraining and inefficiency.

Method used

A differentiable incremental branch (DIB) neural network module is designed, combining differentiable search strategy, and dynamically adjusts the network structure by searching and pruning to adapt to the changes in the data set.

Benefits of technology

It quickly adapts to the new target categories without retraining, improves the efficiency and accuracy of automated blind inspection, and achieves performance indicators comparable to traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115908901B_ABST
    Figure CN115908901B_ABST
Patent Text Reader

Abstract

The present invention provides an automated blind inspection target dynamic classification method, which includes the following steps: S1: Obtain an automated material blind inspection engineering data set; S2: Construct a neural network architecture, and the neural network structure outputs a classification result according to the automated material blind inspection engineering data set, and initialize the backbone of the neural network architecture; S3: When training the backbone of the neural network architecture, if the network does not meet the convergence requirement, grow the backbone of the neural network architecture in two directions of branching and depth according to a preset growth strategy, and the growth continues until the network meets the convergence requirement to obtain a grown neural network; S4: Use the grown neural network to classify the image data of the automated material blind inspection engineering data set. The present invention realizes the dynamic growth, search, and pruning of the network, and solves the problem of re-training caused by continuously adding new categories to the data set in the project.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automated blind inspection and classification, and more specifically, to an automated blind inspection target dynamic classification method. Background Art

[0002] Automated blind inspection technology refers to using automated technology to block the information of producers on materials before material quality inspection and then conducting quality inspection. Compared with the low reliability and high cost of manual blind inspection, automated blind inspection technology reduces manual intervention in the blind inspection process, ensures the standardization of the blind inspection process while reducing labor costs, and can be widely applied in various scenarios. Automated blind inspection technology mainly uses machine vision technology to complete information occlusion, collects target samples through cameras, and often uses Convolutional Neural Network (CNN) for target recognition to assist robots in performing corresponding operations on the targets.

[0003] Since the CNN structure used is carefully designed by hand, the CNN structure is limited by design experience and cannot well adapt to different data sets; in addition, the target categories in the project will increase dynamically, and at this time, the CNN must be retrained, resulting in reduced efficiency. In order to overcome the problems existing in the manually designed neural network to obtain better performance, Neural Architecture Search (NAS) is dedicated to automatically searching for the optimal network structure. NAS is a part of Automated Machine Learning (AutoML) and aims to solve the problem of limited structure in manually designed neural networks. In the early NAS, mainly reinforcement learning methods

[17] and other discretization algorithms were used for spatial search. Although the discretized search method has searched for high-performance CNNs in the target classification tasks on CIFAR-10 and ImageNet, the computational complexity of the algorithm is large. To solve this problem, DAS first proposed a continuous search strategy, mainly for searching hyperparameters, and the search structure is limited. DARTS proposed a method of using a continuous search strategy to search for candidate operations, solving the problem of limited structure. Summary of the Invention

[0004] The present invention provides an automated blind inspection target dynamic classification method to obtain a CNN that adapts to the engineering data set in automated blind inspection and improve the degree of automation of blind inspection.

[0005] To solve the above technical problems, the technical solution of the present invention is as follows:

[0006] An automated blind inspection target dynamic classification method includes the following steps:

[0007] S1: Obtain an automated material blind inspection engineering data set;

[0008] S2: Construct a neural network architecture, design an ordinary incremental branch and a differentiable incremental branch, and initialize the backbone of the neural network architecture based on the ordinary incremental branch;

[0009] When training the backbone of the neural network architecture, if the network does not meet the convergence requirement, according to the preset growth strategy, use the differentiable incremental branch to grow the backbone of the neural network architecture in two directions: branching and depth. The growth continues until the network meets the convergence requirement, and the grown neural network is obtained;

[0010] S4: Use the grown neural network to classify the image data of the automated material blind inspection engineering dataset.

[0011] Preferably, the automated material blind inspection engineering dataset in step S1 includes 6 categories, namely nameplate (plate), bottom of nameplate (platebottom), top of nameplate (plate top), locking nail (buckle), locking nail point (buckle point), and packaging crossbeam (trunk), with a total of 13,500 samples. The first three categories are multi-label mixtures.

[0012] Preferably, constructing the neural network architecture in step S2 is specifically as follows:

[0013] The neural network architecture consists of a feature extractor and a classifier. The feature extractor, which is the backbone of the neural network architecture, is composed of a convolution and ReLU layer (CU, Convolution and relu layer) and five normal incremental branches (NIB, Normal incremental branch) connected in sequence. The last normal incremental branch is connected to the classifier, and the classifier is a CGAP unit (CGAP, Convolution and global average pooling layer).

[0014] Preferably, the convolution and ReLU layer is a sequence container containing a convolutional layer, a ReLU activation layer, and a BN layer, and the normal incremental branch is composed of a convolutional layer, a ReLU activation layer, a max pooling layer, and a BN layer.

[0015] Preferably, growing the backbone of the neural network architecture in two directions: branching and depth according to the preset growth strategy in step S3 is specifically as follows:

[0016] Let the neural network connected to the backbone of the neural network architecture be a branch. When growing in depth, a Differentiable Incremental Branch (DIB) is added after the last layer of the neural network architecture backbone. The depth of the branch connected to the normal incremental branch of the backbone or the differentiable incremental branch of the backbone is called the depth of the branch, and the number of layers of the branch is equal to the depth.

[0017] Start growing the branch from the third layer of the neural network architecture backbone. Give priority to growing the branch first, and then grow the depth.

[0018] When growing the branch, set the kernel size, stride of the differentiable incremental branch or convolution and ReLU layer inside, and the stride of the pooling layer, so as to keep the shape of the feature map consistent with the layer at the same depth in the neural network architecture backbone.

[0019] When growing the branch, connect it to the neural network architecture backbone in the order from shallow to deep. When the deepest layer in the neural network architecture backbone has been connected with the branch, then grow in the depth direction.

[0020] When growing the depth, it is only necessary to connect a differentiable incremental branch that does not change the shape of the feature map to the backbone. Next time, the branch can be grown on the basis of the differentiable incremental branch, and this cycle continues until the convergence condition is met or the preset number of iterations is reached, and no further growth is performed.

[0021] Preferably, the branch is specifically:

[0022] The branch is composed of a convolution and ReLU layer, several differentiable incremental branches, and a convolution and ReLU layer connected in sequence. Among them, the first growing branch has one differentiable incremental branch, and the number of differentiable incremental branches of the subsequent growing branches increases sequentially.

[0023] Preferably, the differentiable incremental branch is specifically:

[0024] The differentiable incremental branch is composed of a convolution layer, a mixed operation, a max pooling layer, and a BN layer. Among them, the input features of the differentiable incremental branch are fused using the Add operation, the mixed operation provides a differentiable search function, and the differentiable incremental branch also uses a residual structure, and the residual structure straddles the mixed operation.

[0025] Preferably, the mixed operation is specifically:

[0026] Continuous values are used in the mixed operation to represent discrete candidate operations, which are composed of the following 7 candidate operations:

[0027] Identity, which means outputting without any processing of the input;

[0028] Conv 3x3 represents a convolution with a 3x3 convolution kernel;

[0029] Conv 5x5 represents a convolution with a 5x5 convolution kernel;

[0030] DiC 3x3, Dilation = 2 represents a dilated convolution with a 3x3 convolution kernel and a dilation factor of 2;

[0031] DiC 5x5, Dilation = 2 represents a dilated convolution with a 5x5 convolution kernel and a dilation factor of 2;

[0032] DiC 5x5, Dilation = 1 represents a dilated convolution with a 5x5 convolution kernel and a dilation factor of 1;

[0033] Conv 5x5 + Dropout represents a 5x5 convolution kernel followed by a Dropout layer;

[0034] Each candidate operation is attached with a structural weight α between 0 and 1. During the forward propagation of the differentiable incremental branch, the softmax transformation is used to convert the structural weight α into a sampling probability and then weighted sum, so that the candidate operations originally represented by discrete values can be represented by the continuous structural weight α. This is the premise for converting discrete search into differentiable search. Therefore, the output of the hybrid operation is as follows:

[0035]

[0036] In the formula, i and j are the indices of the candidate operations, and op i is the candidate operation with index i.

[0037] Preferably, it further includes step S5: when the categories of the automated material blind inspection engineering dataset increase dynamically, first initialize the neural network architecture, and then start searching on the initial dataset with the number of categories being c. If there is a differentiable incremental branch in the neural network architecture, first optimize the structure weight α using the validation set, and then optimize the neural network architecture weight w using the training set; when optimizing w, adopt the weight sharing strategy, that is, when optimizing w, if the neural network architecture has grown, use a larger learning rate η1 for the weight w1 of the units obtained by growth, and use a smaller learning rate η2 for the weight w2 of the units that existed before growth; after training for k epochs, perform branch growth according to the branch growth strategy. The variable record is used to record the number of iterations of model search during the growth process, and record is initialized to 0; the variable growing is used to determine the model growth timing, and growing is initialized to k; when record is equal to growing, growth is performed. When growing the nth branch, the number of DIBs owned by the branch is n - 2, then growing is incremented by k(n - 2), which means growth will occur after another k(n - 2) iterations; when performing growth in the depth direction, since only one differentiable incremental branch is grown, then growing is incremented by k, so each differentiable incremental branch is evenly allocated k iterations for search; after growing the branch, to prevent the number of differentiable incremental branches from growing too fast and reducing the search efficiency, before each branch growth, prune the existing differentiable incremental branches in the network. When growing in depth, a new differentiable incremental branch will be obtained, so the existing DIBs are not pruned. After pruning, continue the training of the next epoch; when reaching the maximum epoch, the number of categories in the dataset increases by 1, and continue a new round of training on the dataset with the number of categories being c + 1: to prevent the newly added categories from affecting the α of the trained model, first prune all differentiable incremental branches, and the α and w of the trained model are retained for weight sharing; after the training of the dataset with the newly added categories is completed, add another category for the next round of training until all newly added categories are trained.

[0038] Preferably, when pruning the differentiable incremental branch, retain the one with the largest structure weight α in the candidate operations, set it to 1 and then eliminate other candidate operations, and no longer perform the forward propagation of other candidate operations. The output of the mixed operation after pruning is:

[0039]

[0040] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0041] 1) The present invention designs a Differentiable Incremental Branch (DIB) neural network module. A differentiable search strategy is used in DIB, which can dynamically access CNNs or perform pruning during the search process.

[0042] 2) A search algorithm suitable for dynamically growing branches in CNNs is designed. After the search is completed, there is no need to retrain. By pruning while searching and utilizing the trained weights, the convergence process is accelerated. The search method proposed in this paper only takes 19 hours and obtains mAP and acc metrics similar to those of DARTS, while the DARTS method takes a total of 41 hours for searching and training on this dataset.

[0043] 3) This paper solves the problem that new target categories need to be retrained in engineering. In the 6-class dataset of the automated material blind inspection project, the differentiable search-growing branch CNN proposed in this paper achieves mAP = 96.4%, exceeding classical CNNs such as the manually designed AlexNet, ResNet-18, and ResNet-50. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a schematic flowchart of the method of the present invention.

[0045] Figure 2 It is a schematic flowchart of the differentiable search-growing branch neural network architecture search method (DIB-NAS) provided in the embodiment.

[0046] Figure 3 It is a schematic diagram of the searchable growing branch neural network structure provided in the embodiment.

[0047] Figure 4 It is a schematic diagram of the NIB structure provided in the embodiment.

[0048] Figure 5 It is a schematic diagram of the DIB structure provided in the embodiment.

[0049] Figure 6 They are the search results of DIB-NAS on 4 datasets respectively.

[0050] Figure 7 It is a schematic diagram of the acc and mAP metrics evaluated on the validation set and test set during the search process provided in the embodiment.

[0051] Figure 8 It is a comparison schematic diagram between IDB-NAS and DARTS during the search and training processes DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0053] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, which do not represent the dimensions of the actual product;

[0054] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0055] The technical solutions of the present invention will be further described below in conjunction with the drawings and embodiments.

[0056] Embodiment 1

[0057] This embodiment provides an automated blind inspection target dynamic classification method, as Figure 1 shown, including the following steps:

[0058] S1: Obtain an automated material blind inspection engineering data set;

[0059] S2: Construct a neural network architecture, design an ordinary incremental branch and a differentiable incremental branch, and initialize the backbone of the neural network architecture based on the ordinary incremental branch;

[0060] S3: When training the backbone of the neural network architecture, if the network does not meet the convergence requirement, according to the preset growth strategy, use the differentiable incremental branch to grow the backbone of the neural network architecture in both the branch and depth directions until the network meets the convergence requirement to obtain the grown neural network;

[0061] S4: Use the grown neural network to classify the image data in the automated material blind inspection engineering data set.

[0062] Embodiment 2

[0063] Based on Embodiment 1, this embodiment further discloses the following content:

[0064] The automated material blind inspection engineering data set in step S1 includes 6 categories: nameplate (plate), bottom of nameplate (plate bottom), top of nameplate (plate top), lock nail (buckle), lock nail key point (buckle point), and packaging crossbeam (trunk), with a total of 13,500 samples, among which the first three categories are multi-label mixtures.

[0065] The DIB-NAS process is as Figure 3 shown, mainly divided into 4 steps: DIB design, initialization of the neural architecture backbone, growth strategy, and adaptive search for category addition.

[0066] In step S2, the construction of the neural network architecture is specifically as follows:

[0067] The neural network architecture consists of a feature extractor and a classifier. When data passes through the initialized backbone, the number of feature map channels doubles and the resolution halves after each layer. At initialization, the feature extractor part is as Figure 3 shown. The feature extractor is the backbone of the neural network architecture, which is sequentially connected by a convolutional layer and a ReLU layer and five normal incremental branches. The last normal incremental branch is connected to the classifier, and the classifier is a CGAP unit.

[0068] The convolutional layer and ReLU layer is a sequence container containing a convolutional layer, a ReLU activation layer, and a BN layer, which mainly processes features before the input branch and after the branch output, and does not have search and access functions. The normal incremental branch consists of a convolutional layer, a ReLU activation layer, a max pooling layer, and a BN layer, as Figure 4 shown. When there is access to NIB, in order to keep the features of each unit in the same layer Figure 1 consistent, the two inputs first go through addition and then through the network in NIB. NIB uses a convolutional layer to obtain potential features; the ReLU activation layer alleviates the problem of gradient disappearance of the Sigmoid activation layer; the max pooling layer is used for downsampling; the BN (Batch Normaization, BN) layer is used to accelerate the convergence speed and alleviate the problem of gradient explosion. NIB has no mixing operation and is not searchable. When the performance of the initialized network is sufficient to meet the task requirements, there is no need for search at this time. Therefore, NIB is used in the backbone instead of DIB to avoid unnecessary time costs.

[0069] The CGAP unit contains a convolutional layer, an average pooling layer, and a fully connected layer, and predicts labels based on the input features. The convolutional layer of the CGAP layer doubles the number of channels and keeps the resolution unchanged. Then, it goes through the average pooling layer for adaptive sampling, and the vectorized features are input into the fully connected layer for prediction.

[0070] In step S3, the backbone of the neural network architecture is grown in two directions: branches and depth according to the preset growth strategy, as Figure 3 shown. Specifically:

[0071] Let the neural network connected to the backbone of the neural network architecture be a branch, and the depth of connecting the branch to the backbone NIB or DIB is called the depth of the branch. When the network does not meet the convergence requirements, growth begins. DIB-NAS has two growth directions: branches and depth during the search process. The growth of branches depends on the network backbone, and the purpose is to provide diverse potential features for the backbone. The DIB that grows in the depth direction provides access for the branches.

[0072] Since the role of the branch is to provide diverse features for the backbone, branches with too shallow depth have fewer layers and it is difficult to achieve the goal. Therefore, during the search process, branches start growing from the third layer of the network backbone. Compared with a single DIB growing in the depth direction, branches can provide richer features for the backbone. Therefore, branches are preferentially grown first, and then the depth is grown.

[0073] When growing in depth, a differentiable incremental branch is added after the last layer of the backbone of the neural network architecture. The depth of the branch connected to the normal incremental branch or the differentiable incremental branch of the backbone is called the depth of the branch, and the number of layers of the branch is equal to the depth.

[0074] Branches start growing from the third layer of the backbone of the neural network architecture. Branches are preferentially grown first, and then the depth is grown.

[0075] When growing branches, by setting the kernel size, stride of the differentiable incremental branch or convolution and ReLU layer inside, and the stride of the pooling layer, the shape of the feature map can be kept consistent with the layer at the same depth in the backbone of the neural network architecture, ensuring that it can be connected to the backbone.

[0076] When growing branches, they are connected to the backbone of the neural network architecture in order from shallow to deep. When the deepest layer in the backbone of the neural network architecture has been connected with branches, then grow in the depth direction.

[0077] When growing depth, it is sufficient to connect a differentiable incremental branch that does not change the shape of the feature map to the backbone. Next time, branches can be grown on the basis of the differentiable incremental branch. This cycle continues until the convergence condition is met or the preset number of iterations is reached, and no further growth is performed.

[0078] The specific form of the branch is as follows:

[0079] The branch is composed of a convolution and ReLU layer, several differentiable incremental branches, and a convolution and ReLU layer connected in sequence. Among them, the first growing branch has one differentiable incremental branch, and the number of differentiable incremental branches of the subsequent growing branches increases sequentially.

[0080] The specific form of the differentiable incremental branch is as follows:

[0081] As Figure 5 shown, the differentiable incremental branch is composed of a convolution layer, a mixed operation, a max pooling layer, and a BN layer. Among them, the input features of the differentiable incremental branch are fused using the Add operation. The mixed operation provides a differentiable search function. The differentiable incremental branch also uses a residual structure, and the residual structure straddles the mixed operation.

[0082] The specific form of the mixed operation is as follows:

[0083] In the mixed operation, continuous values are used to represent discrete candidate operations, which are composed of the following 7 candidate operations:

[0084] Table 1 shows seven candidate operations in the mixing operation.

[0085]

[0086] In the table, Identity means output without any processing on the input;

[0087] Conv 3x3 means convolution with a 3x3 convolution kernel;

[0088] Conv 5x5 means convolution with a 5x5 convolution kernel;

[0089] DiC 3x3, Dilation = 2 means dilated convolution with a 3x3 convolution kernel and a dilation factor of 2;

[0090] DiC 5x5, Dilation = 2 means dilated convolution with a 5x5 convolution kernel and a dilation factor of 2;

[0091] DiC 5x5, Dilation = 1 means dilated convolution with a 5x5 convolution kernel and a dilation factor of 1;

[0092] Conv 5x5 + Dropout (abbreviated as DrC) means a 5x5 convolution kernel followed by a Dropout layer;

[0093] Each candidate operation is attached with a structural weight α between 0 and 1. During the forward propagation of the differentiable incremental branch, the softmax transformation is used to convert the structural weight α into a sampling probability and then weighted summation is performed, so that the candidate operations originally represented by discrete values can be represented by the continuous structural weight α. This is the premise for converting discrete search into differentiable search. Therefore, the mixing operation output is as follows:

[0094]

[0095] In the formula, i and j are the indices of the candidate operations, op i is the candidate operation with index i. After the feature map passes through the mixing operation, its shape remains unchanged.

[0096] Max pooling layer is used for downsampling in DIB, which can select the maximum value feature and allow the residual structure to cross the mixing operation. After passing through the pooling layer, the length and width of the feature map are halved. In addition, the BN layer in DIB can slow down the gradient explosion and accelerate the network convergence speed.

[0097] Example 3

[0098] Based on Example 1 and Example 2, this example further discloses the following content:

[0099] It also includes step S5: When the categories of the automated material blind inspection engineering data set increase dynamically, first initialize the neural network architecture, and then start searching on the initial data set with the number of categories being c. If the neural network architecture contains a differentiable incremental branch, first optimize the structure weight α using the validation set, and then optimize the neural network architecture weight w using the training set; when optimizing w, adopt the weight sharing strategy, that is, when optimizing w, if the neural network architecture has grown, use a larger learning rate η1 for the weight w1 of the units obtained by growth, and use a smaller learning rate η2 for the weight w2 of the units that existed before growth, because w1 has not been optimized yet and is noise for the model. These noises will interfere with the gradient calculation of w2. To prevent excessive fluctuations during the optimization of w2, a smaller learning rate η2 is used. At this time, it is equivalent to the model after growth sharing all the weights of the model before growth; after training for k epochs, perform branch growth according to the branch growth strategy. The variable record is used to record the number of iterations of model search during the growth process, and record is initialized to 0; the variable growing is used to determine the model growth timing, and growing is initialized to k; when record is equal to growing, growth occurs. When growing the nth branch, the number of DIBs owned by the branch is n - 2, then growing is incremented by k(n - 2), meaning growth will occur after another k(n - 2) iterations; when depth-wise growth occurs, since only one differentiable incremental branch is grown, then growing is incremented by k, so each differentiable incremental branch is evenly allocated k iterations for search; after growing the branch, to prevent the number of differentiable incremental branches from growing too fast and reducing the search efficiency, before each branch growth, prune the existing differentiable incremental branches in the network. When growing the depth, a new differentiable incremental branch will be obtained, so the existing DIBs are not pruned. After pruning, continue with the training of the next epoch; when the maximum epoch is reached, the number of categories in the data set is increased by 1 classification, and continue with a new round of training on the data set with the number of categories being c + 1: To prevent the newly added category from affecting the α of the trained model, first prune all the differentiable incremental branches, and the α and w of the trained model are retained for weight sharing, so that both the trained model can extract the features of the trained categories and the growing branches can quickly adapt to the new categories; after the data set with the newly added classification is trained, add another category for the next round of training until all the newly added categories are trained, as follows.

[0100]

[0101]

[0102] When the differentiable incremental branch is pruned, the one with the largest structural weight α in the candidate operations is retained, and the other candidate operations are eliminated after it is set to 1, and the forward propagation of other candidate operations is no longer performed. The output of the mixed operation after pruning is:

[0103]

[0104] In the specific implementation process, the experimental environment Ubuntu 20.04 was used, and the deep learning framework was pytorch1.11+cu113. The hardware configuration used by the server is: the CPU uses Intel i5, the GPU is NVIDIA RTX3090, and the memory is 16G. The engineering dataset was made into 4 incremental datasets to simulate the scenario of adding new target categories. mAP was used as the evaluation indicator in the experiment. DIB-NAS and DARTS were compared in NAS. In terms of manual networks, AlexNet, ResNet-18, and ResNet-50 were selected for comparison. The important parameters in the experiment remained consistent.

[0105] This paper uses an image classification dataset from a blind material inspection project. The sample resolution of this dataset is 1080x1920, and there are six categories: P (nameplate), PB (bottom of nameplate), PT (top of nameplate), B (lock nail), BP (lock nail card point), and T (packaging beam). Among them, the three categories of nameplate have multiple labels, so the dataset has a total of 10 labels. The number of original samples and the number of enhanced samples in the dataset are shown in Table 2. The original samples are divided into a test set, a validation set, and a training set according to 2:1:7. Then the number of samples of each label in the test set is enhanced to 270, the test set is enhanced to 135, and the training set is enhanced to 945. Considering the influence of light environment and viewing angle, three methods are used when enhancing the data: horizontal flip (0 or 1), rotation (-10°~10°), and brightness adjustment (-0.5~0.5). The enhanced values of each method are randomly generated, and the three methods are nested for enhancement. In order to speed up the search, the sample resolution is compressed to 280x480.

[0106] Table 2. Distribution of original samples and enhanced samples of the dataset.

[0107]

[0108] The network backbone is initialized with 6 layers and a maximum depth of 12. In Algorithm 1, the parameter k=4, that is, each DIB is assigned 4 iterations for search. When rotating the dataset, the number of searches for each dataset is e=32. During the search process, the Adam optimizer is used to optimize α and w, and the batch size is 20.

[0109] Search results for each dataset Figure 6As shown in the figure. Starting from dataset 1 for search, the initialized network backbone has 6 layers. After 10 training iterations, branch 1 starts to grow. Branch 1 is connected to the 3rd layer of the backbone and contains 1 DIB. Therefore, according to Algorithm 1, the growth of branch 2 is carried out every 4 iterations. Before growing branch 2, the DIB in branch 1 is pruned first. Branch 2 is connected to the 4th layer of the backbone and contains 2 DIBs. After growing to branch 4 according to Algorithm 1, since the initialized network backbone only has 6 layers and it is impossible to grow branch 5 to be connected to the 7th layer of the backbone, after growing branch 4, the 7th layer of the backbone is grown next, and the next growth is postponed by 4 iterations. After growing the 7th layer of the backbone, branch 5 is grown and connected to the 7th layer. And so on until the maximum depth is reached and the corresponding branches are grown. After pruning branches 5, 7, and 9, the dataset is rotated. In the search results of the first dataset, the candidate operations tend to select dilated convolutions and ordinary convolutions with a kernel size of 3, which is more in line with the experience of manually designed networks, that is, mainly using 3x3 kernels to obtain latent layer features. Dataset 2 adds a type of rivet, and its background has a relatively large difference from that of the nameplate. In the search results, it tends to select ordinary convolutions with a kernel size of 5x5 to focus on background information. Dataset 3 adds a type of rivet clamping point, and the area of the rivet clamping point in the image is relatively small. The search results tend to select ordinary convolutions with a kernel size of 3x3 to focus on local information. Dataset 4 adds a type of packaging crossbeam, and the search results tend to select ordinary convolutions with a kernel size of 3x3 with a dropout layer, which may be the result of overfitting in the previous categories.

[0110] The area of the rivet clamping point is relatively small in the image, and the search results tend to select 3x3 ordinary convolutions to focus on local information. Dataset 4 adds a type of packaging crossbeam, and the search results tend to select 3x3 ordinary convolutions with a dropout layer, which may be the result of overfitting in the previous categories.

[0111] Table 3 Main parameters of the optimizer used in the search process. Other parameters are default values of the torch optimizer.

[0112]

[0113] The prediction accuracy of the network is evaluated using the mAP (mean average precision) metric in multi-label tasks. The formula for calculating mAP is:

[0114]

[0115] where N is the number of categories in the dataset, and AP is the average precision of the i-th category. The calculation of AP adopts the standard used in PASCAL VOC 2010, that is, AP can be considered as the area under the precision-recall (P-R) curve. Therefore, the calculation method of AP is:

[0116]

[0117] where the calculation methods of precision and recall are respectively:

[0118] P = TP / (TP + FN)

[0119] R = TP / (TP + FP)

[0120] TP represents the number of correctly predicted positive samples, FN represents the number of incorrectly predicted negative samples, and FP represents the number of incorrectly predicted positive samples.

[0121] In the search experiment, the network was evaluated every 4 iterations. mAP was used to evaluate the test set and the validation set in the evaluation, and the evaluation results are as Figure 7 shown. The metrics on the test set are generally higher than those on the validation set and are relatively stable compared to the validation set. DIB-NAS incremental search can handle the situation of adding new target classes.

[0122] Figure 8 The mAP curve in shows that when the dataset rotation occurs during the search process of DARTS, the mAP metric will drop significantly, while the mAP of DIB-NAS is relatively stable when the dataset rotation occurs. In the DARTS search experiment, each dataset was searched for 50 iterations. When training, the 4 datasets were trained for 50, 64, 100, and 200 iterations respectively. The total time for DARTS search and training is 41 hours. DIB-NAS does not need to be retrained after the search, and each dataset is searched for 65 iterations.

[0123] In the comparative experiment, AlexNet, ResNet-18, and ResNet-50 were selected as comparison objects. The 3 networks were searched for 65 iterations on each dataset, evaluated every 4 iterations and when the dataset rotation occurred, and the highest metrics were retained for comparison. In addition to mAP, the experiment also compared the floating-point operations per second, the number of parameters, the storage size, and the total duration of the experiment. The evaluation metrics of DIB-NAS on the 4 datasets exceeded those of CNNs such as AlexNet, ResNet-18, and ResNet-50, and the floating-point operations per second were relatively lower, and the number of parameters was lower than that of ResNet-50. DIB-NAS does not need to retrain the network after the search and achieved results similar to those of DARTS in only half the time of DARTS.

[0124] The same or similar reference numerals correspond to the same or similar components;

[0125] The terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0126] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. An automated blind inspection target dynamic classification method, characterized in that, It includes the following steps: S1: Obtain the automated material blind inspection engineering dataset; S2: Construct a neural network architecture, design an ordinary incremental branch and a differentiable incremental branch, and initialize the backbone of the neural network architecture based on the ordinary incremental branch; S3: When training the backbone of the neural network architecture, if the network does not meet the convergence requirement, according to the preset growth strategy, use the differentiable incremental branch to grow the backbone of the neural network architecture in both the branch and depth directions. The growth continues until the network meets the convergence requirement, and the grown neural network is obtained; S4: Use the grown neural network to classify the image data in the automated material blind inspection engineering dataset; In step S2, constructing the neural network architecture specifically is: The neural network architecture consists of a feature extractor and a classifier. The feature extractor is the backbone of the neural network architecture, which is sequentially connected by a convolution and ReLU layer and five normal incremental branches. The last normal incremental branch is connected to the classifier, and the classifier is a CGAP unit; The convolution and ReLU layer is a sequence container containing a convolutional layer, a ReLU activation layer, and a BN layer. The normal incremental branch consists of a convolutional layer, a ReLU activation layer, a max pooling layer, and a BN layer; In step S3, growing the backbone of the neural network architecture in both the branch and depth directions according to the preset growth strategy specifically is: Let the neural network connected to the backbone of the neural network architecture be a branch. When growing in depth, add a differentiable incremental branch after the last layer of the backbone of the neural network architecture. The depth at which the branch is connected to the normal incremental branch or the differentiable incremental branch of the backbone is called the depth of the branch, and the number of layers of the branch is equal to the depth; Start growing branches from the third layer of the backbone of the neural network architecture. Give priority to growing branches and then growing in depth; When growing branches, set the kernel size, stride of the differentiable incremental branch or the convolution and ReLU layer inside, and the stride of the pooling layer so as to keep the shape of the feature map consistent with the layers at the same depth in the backbone of the neural network architecture; When growing branches, connect them to the backbone of the neural network architecture in the order from shallow to deep. When the deepest layer in the backbone of the neural network architecture has been connected with a branch, grow in the depth direction; When growing in depth, just connect a differentiable incremental branch that does not change the shape of the feature map to the backbone. Next, branches can be grown on the basis of the differentiable incremental branch, and so on until the convergence condition is met or the preset number of iterations is reached, and no more growth is performed; The branch specifically is: The branch is sequentially connected by a convolution and ReLU layer, several differentiable incremental branches, and a convolution and ReLU layer. Among them, the first growing branch has one differentiable incremental branch, and the number of differentiable incremental branches of the subsequent growing branches increases sequentially; The differentiable incremental branch specifically is: The differentiable incremental branch consists of a convolutional layer, a mixed operation, a max pooling layer, and a BN layer. Among them, the input features of the differentiable incremental branch are fused using the Add operation. The mixed operation provides a differentiable search function. The differentiable incremental branch also uses a residual structure, and the residual structure spans the mixed operation.

2. The automated blind inspection target dynamic classification method according to claim 1, wherein The automated material blind inspection engineering data set described in step S1 includes six categories, namely, nameplate, plate bottom, plate top, buckle, buckle point, and packaging crossbeam, with a total of 13,500 samples, of which the first three categories are multi-label mixed.

3. The automated blind inspection target dynamic classification method according to claim 2, characterized in that, The mixing operation is specifically as follows: The mixed operation uses continuous values to represent discrete candidate operations, which are divided into the following 7 candidate operations: composition: Identity means outputting without any processing of the input; Conv 3x3, indicating a convolution with a convolution kernel of 3x3; Conv 5x5, indicating a convolution with a convolution kernel of 5x5; DiC 3x3, Dilation = 2, indicating a dilated convolution with a kernel of 3x3 and a dilation factor of 2; DiC 5x5, Dilation=2, which means a dilated convolution with a kernel of 5x5 and a dilation factor of 2; DiC 5x5, Dilation = 1, indicating a dilated convolution with a kernel of 5x5 and a dilation factor of 1; Conv 5x5+Dropout means the convolution kernel is 5x5 convolution followed by a Dropout layer; Each candidate operation is accompanied by a structural weight α between 0 and 1. When the differentiable incremental branch is forward propagated, the structural weight α is converted into a sampling probability using softmax transformation and then weighted summed, so that the candidate operation originally represented by a discrete value can be represented by a continuous structural weight α. This is the premise for converting discrete search to differentiable search. Therefore, the output of the hybrid operation is as follows: where i and j are indices of candidate operations, and op i is the candidate operation with index i.

4. The automated blind inspection target dynamic classification method according to claim 3, wherein It also includes step S5: when the categories of the automated material blind inspection engineering data set are dynamically increased, the neural network architecture is first initialized, and then the search is started on the initial data set with a category number of c. If the neural network architecture contains differentiable incremental branches, the structural weight α is first optimized using the verification set, and then the neural network architecture weight w is optimized using the training set; when optimizing w, a weight sharing strategy is adopted, that is, when optimizing w, if the neural network architecture has been grown, a larger learning rate η1 is used for the weight w1 of the grown unit, and a smaller learning rate η2 is used for the weight w2 of the unit existing before the growth; after training k epochs, branch growth is performed according to the branch growth strategy, and the variable record is used to record the number of iterations of the model search during the growth process, and record is initialized to 0; Use the variable growing to determine the model growth timing. growing is initialized to k. When record is equal to growing, growth occurs. When growing the nth branch, the number of DIBs the branch has is n - 2, then growing is incremented by k(n - 2), meaning growth will occur after another k(n - 2) iterations. When growth occurs in the depth direction, since only one differentiable incremental branch is grown, then growing is incremented by k. Thus, each differentiable incremental branch is evenly allocated k iterations for search. After growing a branch, to prevent the number of differentiable incremental branches from growing too fast and reducing the search efficiency, before growing each branch, prune the existing differentiable incremental branches in the network. When growing in depth, a new differentiable incremental branch will be obtained, then do not prune the existing DIBs. After pruning, continue with the training of the next epoch. When the maximum epoch is reached, one more classification is added to the dataset categories, and continue with a new round of training on the dataset with the number of categories c + 1. To prevent the newly added category from affecting the α of the already trained model, first prune all the differentiable incremental branches, and the α and w of the already trained model are retained for weight sharing. After the training of the dataset with the newly added category is completed, add another category for the next round of training until all the newly added categories are trained.

5. The automated blind inspection target dynamic classification method according to claim 4, wherein When pruning the differentiable incremental branches, retain the one with the largest structural weight α in the candidate operations, set it to 1 and eliminate the other candidate operations, and no longer perform the forward propagation of the other candidate operations. The output of the mixed operation after pruning is: