Small sample pipeline DR defect detection method
By introducing a smooth variational autoencoder to generate new class sample features and adaptively fine-tune the backbone network, the problem of low accuracy of new class defect classification in traditional pipeline detection methods is solved, and efficient detection under small sample conditions is achieved.
Patent Information
- Application Number
- CN202510515948.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-05
AI Technical Summary
Traditional pipeline defect detection methods rely on a large amount of labeled data, which is cost-effective and time-consuming to obtain, and the accuracy of classification and positioning of new defects is low.
A two-stage training method based on smooth variational autoencoder and adaptive fine-tuning is adopted to improve the defect detection performance of small sample class by generating new class sample features and adaptive fine-tuning backbone network parameters.
The accuracy of defect detection is significantly improved under small sample conditions and is suitable for health monitoring and safety assessment of industrial pipelines.
Smart Images

Figure CN120431315A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pipeline detection and relates to a small-sample pipeline digital radiography (DR) defect detection method. Specifically, it relates to a small-sample pipeline DR defect detection method based on a smoothed variational autoencoder (S-VAE) and adaptive fine-tuning, which can be applied to defect detection of pipelines such as gas pipelines. Background Art
[0002] The safe operation of gas pipelines is crucial to social stability and the safety of residents. However, traditional pipeline defect detection methods typically rely on large amounts of labeled data, which is expensive and time-consuming to obtain. Therefore, small-sample learning techniques are of great significance in pipeline defect detection. Summary of the Invention
[0003] To address the low accuracy of traditional detection methods in classifying and locating new types of defects, the present invention provides a small-sample pipeline DR defect detection method. This method performs defect detection based on a two-stage training method using a smoothed variational autoencoder (S-VAE) and adaptive fine-tuning. By generating new feature samples and adaptively fine-tuning backbone network parameters, it improves the detection performance of small-sample defects.
[0004] The present invention discloses a small sample pipeline DR defect detection method, comprising:
[0005] Step 1: Obtain the pipeline DR defect dataset X and the corresponding label Y, and split the dataset; wherein, X is divided into the training set X according to the ratio of N1:N2 train and the test set X test , N1 and N2 are both greater than 0 and N1>N2; all categories C corresponding to the dataset are divided into base classes and new classes according to the ratio of M1:M2, M1 and M2 are both greater than 0 and M1>M2; each new class considers a small sample of annotated training samples within K in turn;
[0006] Step 2: Build a small sample target detection model based on Faster RCNN, which consists of three parts: backbone network, smooth variational autoencoder and target detection module;
[0007] Step 3: Construct a loss function, which includes RPN binary cross entropy loss, bounding box classification loss, bounding box regression loss, and smooth variational autoencoder loss;
[0008] Step 4: Adaptively fine-tune the small sample target detection model. The L2 norm is used to automatically select the top k important convolution kernels in the backbone network to update the network. The adaptive fine-tuning ratio is 20% < 30%, so that the backbone network can extract the features of the new class of samples.
[0009] Step 5: Test the trained small-sample object detection network based on smooth variational autoencoder and adaptive fine-tuning on the test set, and calculate the average accuracy of different numbers of samples on different new class sets, the average accuracy of the base class, and the average accuracy of the new class.
[0010] As a further improvement of the present invention, N1:N2 is preferably 7:3, M1:M2 is preferably 3:1, K≤10, preferably K={1, 2, 3, 5, 10}.
[0011] As a further improvement of the present invention, the step 2 specifically includes:
[0012] Step 2.1. Build the backbone network: Use the ImageNet pre-trained Resnet-101 as the backbone network, with a mini-block size of 16 and stochastic gradient descent (SGD) as the optimization method; in the base training phase, the initial learning rate is set to 0.001, the momentum is 0.9, and the weight decay is 1e-4; in the few-shot fine-tuning phase, the learning rate is set to 0.0004.
[0013] Step 2.2: Build a smooth variational autoencoder:
[0014] First, the smooth variational autoencoder uses an encoder to map the input image features to a latent space, generating a set of latent variables, including the mean and variance of the latent variables;
[0015] Then, the loss functions of log-cosh and smooth L1 are introduced. Log-cosh and smooth L1 are expressed as:
[0016]
[0017] Among them, h is the feature of the real data, represents the features reconstructed by S-VAE, and t represents the error between the actual value and the predicted value z is the latent variable, φ is the parameter of the encoder, q φ (z) is the distribution of the latent variable z, p(z) is the prior distribution, and a is a positive hyperparameter used to adjust the sensitivity of the loss function to the error. When the error |t| is large, log-cosh and smooth L1 are close to the L1 norm, showing robustness to large errors. When the error |t| is small, it is close to the L2 norm, and log-cosh and smooth L1 show the advantage of smoothness.
[0018] The reconstruction loss is expressed as:
[0019] L recon (|t|)=αL log-cosh (|t|)+(1-α)L sm L1 (|t|) (3)
[0020] Among them, α is a hyperparameter between 0 and 1, which is used to control the contribution of the two losses;
[0021] Finally, the distance between latent variables is used as a regularization term to decouple the latent variables. The regularization term is expressed as:
[0022]
[0023] in, Represents the sum of the off-diagonal elements of the covariance matrix, which is used to penalize the off-diagonal elements of the covariance matrix. Represents the sum of the squares of the differences between the diagonal elements of the covariance matrix and 1, which is used to adjust the variance matrix of the latent variables to ensure that the variance of each latent variable is close to 1 and maintain an appropriate degree of decoupling in the latent space; od and λ d is a hyperparameter;
[0024] To sum up, the total loss of S-VAE is:
[0025] L S-VAe =L recon (|t|)+D KL (q φ(z|x) ||p(z))+βD(q φ (z)||p(z)) (5)
[0026] Among them, L recon (|t|) represents the reconstruction loss calculated using log-cosh and smooth L1 loss, D KL (q φ(z|x) ||p(z)) represents the output q of the encoder φ(z|x) KL divergence between the prior distribution p(z), D(q φ (z)||p(z)) represents the decoupled regularization term, and β is D(q φ The weight of the term (z)||p(z));
[0027] Step 2.3: Build the target detection module:
[0028] First, RPN is introduced to generate candidate boxes. RPN slides on the feature map to generate multiple candidate boxes.
[0029] Then, the candidate regions are uniformly mapped to a fixed-size feature map through the RoI pooling layer;
[0030] Finally, the classification head is used to determine the category of the candidate box, and the regression network adjusts the boundary of the candidate box to make it more accurate.
[0031] As a further improvement of the present invention, the step 3 specifically includes:
[0032] Step 3.1: Construct the RPN binary cross entropy loss. The RPN binary cross entropy loss is used to optimize the model's ability to distinguish foreground from background, and is expressed as:
[0033]
[0034] Among them, y i represents the true label of the i-th sample, is the predicted foreground probability, N represents the number of samples involved in training;
[0035] Step 3.2: Construct the bounding box classification loss. The bounding box classification loss is used to classify different targets and optimize the model's ability to distinguish different categories. It is expressed as:
[0036]
[0037] Among them, y ic represents the true category label of the i-th sample, Represents the category probability predicted by the model, and C represents the number of categories involved in training;
[0038] Step 3.3: Construct a smooth L1 loss for bounding box regression. The smooth L1 loss for bounding box regression is used for accurate positioning of the bounding box and is expressed as:
[0039]
[0040] Among them, y ib represents the true bounding box parameters, represents the predicted bounding box parameters;
[0041] Step 3.4: Construct the loss of the smooth variational autoencoder. The loss of the smooth variational autoencoder is used to constrain the encoded latent variables so that they are close to the standard normal distribution, which is expressed as:
[0042] L s-VAE =L recon +D KL (q φ(z|x) ||p(z))+βD(q φ (z)||p(z)) (9)
[0043] Among them, L recon represents the reconstruction loss calculated using log-cosh and smooth L1 loss, D KL (q φ(z|x) ||p(z)) represents the output q of the encoder φ(z|x) KL divergence between the prior distribution p(z), D(q φ (z)||p(z)) represents the decoupled regularization term, and β is D(q φ(z|x) The weight of the term ||p(z));
[0044] Finally, the total loss function of the training phase is expressed as:
[0045] L base =L RPN +L cls +L reg +aL S-VAE (10)
[0046] Among them, a is used to balance the weights of S-VAE loss and other loss terms of S-EDH-Faster RCNN.
[0047] As a further improvement of the present invention, step 4 specifically includes:
[0048] The L2 norm is selected as the importance index of the convolution kernel, and the calculation method is:
[0049]
[0050] Among them, W i represents the i-th convolution kernel, w ij Represents the value of the jth parameter in the i-th convolution kernel;
[0051] After obtaining the L2 norm of the convolution kernel, the convolution kernels are sorted and the top k convolution kernels are selected for network update. Adaptive fine-tuning can effectively improve the generalization ability of the model under small sample conditions; among them, the proportion of adaptive fine-tuning is preferably 26%.
[0052] Compared with the prior art, the present invention has the following beneficial effects:
[0053] The present invention introduces a smooth variational autoencoder to generate new class sample features, increase the diversity of training data, and improve the generalization ability of the model. At the same time, an adaptive strategy is used to update the key convolution kernels in the backbone network during the fine-tuning stage, so that the model can better learn the new class sample features and improve the detection accuracy in the case of few samples. This detection method can improve the accuracy of defect detection in small sample situations and is widely used in industrial pipeline health monitoring and safety assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 This is a flow chart of the small sample pipeline DR defect detection method disclosed in the present invention;
[0055] Figure 2 This is a schematic diagram of the overall network architecture disclosed in the present invention. DETAILED DESCRIPTION
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0057] The present invention will be described in further detail below with reference to the accompanying drawings:
[0058] like Figure 1 、 2 As shown, the present invention provides a small sample pipeline DR defect detection method based on smooth variational autoencoder (S-VAE) and adaptive fine-tuning. It generates new class sample features by introducing a smooth variational autoencoder, increases the diversity of training data, and improves the generalization ability of the model. At the same time, an adaptive strategy is used to update the key convolution kernels in the backbone network during the fine-tuning stage, so that the model can better learn the new class sample features and improve the detection accuracy in the case of few samples. The method specifically includes:
[0059] Step 1: Obtain the pipeline DR defect dataset X and the corresponding label Y, and segment the dataset;
[0060] Specifically include:
[0061] Step 11: Obtain images and corresponding labels from the gas pipeline DR defect dataset, which contains different types of pipeline defects;
[0062] Step 12: Divide the entire dataset into training and test sets in a ratio of 7:3, that is, 70% of the data is used for model training and 30% of the data is used for testing and verifying model performance.
[0063] Step 13: Divide all defect categories into base categories and new categories in a ratio of 3:1.
[0064] Step 14: To simulate small-shot learning, consider 1, 2, 3, 5, and 10 annotated training samples in each new class for subsequent model fine-tuning and testing.
[0065] Step 2: Build a small sample target detection model based on Faster RCNN, which consists of three parts: backbone network, smooth variational autoencoder and target detection module;
[0066] Specifically include:
[0067] Step 2.1. Build the backbone network: ResNet-101, pre-trained on ImageNet, is used as the model's backbone network. ResNet-101 uses deep residual learning to extract multi-scale image features, enabling it to handle complex visual tasks. During the base training phase, a large number of base class samples are used, and stochastic gradient descent (SGD) is used as the optimization method. The initial learning rate is set to 0.001, the momentum is 0.9, and the weight decay is 1e-4. When processing a small number of new class samples, the learning rate is reduced to 0.0004 to better adapt to small sample conditions and avoid overfitting.
[0068] Step 2.2: Construct a Smoothed Variational Autoencoder (S-VAE): Image features are mapped via an encoder, and Log-Cosh and Smooth L1 loss functions are used to handle large and small errors, ensuring decoupling of latent variables. The S-VAE module first uses an encoder to map image features into a latent space, generating the mean and variance of the latent variables. To enhance the model's robustness to large errors and smoothness to small errors, Log-Cosh and Smooth L1 are introduced as reconstruction loss functions. Log-Cosh and Smooth L1 are closer to the L1 norm when handling large errors and to the L2 norm when handling small errors, respectively. To ensure decoupling in the latent space, the distance between latent variables is used as a regularization term to constrain their correlation and make them more independent. Finally, the S-VAE's overall loss function includes the reconstruction loss, KL divergence, and the decoupling regularization term, optimizing both the diversity and accuracy of the generated features.
[0069] Step 2.3: Build the object detection module: Use the RPN to generate candidate boxes, the RoI pooling layer to process the candidate regions, the classification head to determine the category, and the regression network to adjust the bounding box. First, the RPN module uses the region proposal network (RPN) to slide on the feature map to generate multiple candidate boxes. Then, the RoI pooling layer maps the candidate boxes to a fixed-size feature map through the RoI pooling layer, ensuring that the dimensions of candidate boxes of different sizes are consistent during subsequent processing. Finally, the classification and regression heads are responsible for determining the category contained in the candidate region and accurately locating the target position, respectively.
[0070] Step 3: Construct a loss function, which includes RPN binary cross entropy loss, bounding box classification loss, bounding box regression loss, and smooth variational autoencoder loss;
[0071] Specifically include:
[0072] Step 3.1: Construct RPN binary cross entropy loss to optimize the classification of foreground and background.
[0073] Step 3.2: Construct bounding box classification loss to optimize target classification ability.
[0074] Step 3.3: Construct a smooth L1 loss for bounding box regression to accurately locate the bounding box.
[0075] Step 3.4. Construct the S-VAE loss, including reconstruction loss, KL divergence, and decoupling regularization terms.
[0076] Step 4: Adaptive fine-tuning training; in the fine-tuning stage, the L2 norm is used to select the first k important convolution kernels in the backbone network for fine-tuning, and the fine-tuning ratio is set to 26% to improve the feature extraction ability of new class samples.
[0077] Step 5: Test the trained small-shot object detection network based on smoothed variational autoencoders and adaptive fine-tuning on the test set. Calculate the average accuracy of the three new class sets (split1, split2, and split3) with different numbers of samples (1, 2, 3, 5, and 10 shots), the average accuracy of the base class, and the average accuracy of the new class.
[0078] Example:
[0079] The present invention uses the pipeline defect image dataset (PIP-DET) for experiments. PIP-DET contains 20 defect categories, with a total of 6,010 samples, of which the number of training set samples is 4,258 and the number of test set samples is 1,752. The image labeling tool is LabelImg. The present invention follows previous work in the data set division method to ensure fair comparison. Specifically, for the PIP-DET dataset, the present invention randomly selects 15 classes as base classes and the remaining 5 classes as new classes. Each new class has K = {1, 2, 3, 5, 10} annotated training samples to simulate the few-sample scenario. Three random groupings are considered in the present invention and named split1, split2, and split3. During the training and testing process, the image size in the dataset is uniformly adjusted to 256×256, and the number of channels is 3.
[0080] S1: Obtain the pipeline DR defect dataset X and the corresponding label Y, and split the dataset; divide X into the training set X and the training set Y in a ratio of 7:3. train and the test set X test; Split category C into base class and new class in a ratio of 3:1. Each new class considers K = {1, 2, 3, 5, 10} annotated training samples in turn.
[0081] S2: Build a small-sample target detection model based on Faster RCNN. The model consists of three parts: the backbone network, the smoothed variational autoencoder, and the target detection module. Specifically:
[0082] S2.1: Build the backbone network. Specifically, use the ImageNet pre-trained Resnet-101 as the backbone network, with a mini-block size of 16, and use stochastic gradient descent (SGD) as the optimization method. During the base training phase, the initial learning rate is set to 0.001, the momentum is 0.9, and the weight decay is 1e-4. During the few-shot fine-tuning phase, the learning rate is set to 0.0004.
[0083] S2.2: Construct a smooth variational autoencoder. First, the smooth variational autoencoder uses an encoder to map the input image features to a latent space, generating a set of latent variables, including the mean and variance of the latent variables. Then, the Log-Cosh and Smooth L1 loss functions are introduced. Log-Cosh and Smooth L1 can be expressed as:
[0084]
[0085]
[0086] Among them, h is the feature of the real data, represents the features reconstructed by S-VAE, and t represents the error between the actual value and the predicted value z is the latent variable, φ is the parameter of the encoder, q φ (z) is the distribution of the latent variable z, p(z) is the prior distribution, and a is a positive hyperparameter used to adjust the sensitivity of the loss function to the error |t|. When the error |t| is large, log-cosh and smooth L1 approach the L1 norm, demonstrating robustness to large errors. When the error |t| is small, it approaches the L2 norm, with log-cosh and smooth L1 exhibiting the advantage of smoothness.
[0087] The reconstruction loss is expressed as:
[0088] L recon (|t|)=αL log-cosh (|t|)+(1-α)L smooth L1 (|t|) (3)
[0089] Among them, α is a hyperparameter between 0 and 1, which is used to control the contribution of the two losses.
[0090] Finally, the distance between latent variables is used as a regularization term to decouple the latent variables. This regularization term can be expressed as:
[0091]
[0092] in, Represents the sum of the off-diagonal elements of the covariance matrix, which is used to penalize the off-diagonal elements of the covariance matrix. It represents the sum of the squares of the differences between the diagonal elements of the covariance matrix and 1, and is used to adjust the variance matrix of the latent variables to ensure that the variance of each latent variable is close to 1 and maintain an appropriate degree of decoupling in the latent space. od and λ d is a hyperparameter.
[0093] To sum up, the total loss of S-VAE is:
[0094] L S-VA =L recon (|t|)+D KL (q φ(z|x) ||p(z))+βD(q φ (z)||p(z)) (5)
[0095] Among them, L recon (|t|) represents the reconstruction loss calculated using log-cosh and smooth L1 loss, D KL (q φ(z|x) ||p(z)) represents the output q of the encoder φ(z|x) KL divergence between the prior distribution p(z), D(q φ (z)||p(z)) represents the decoupled regularization term, and β is D(q φ (z)||p(z)) item’s weight.
[0096] S2.3: Build the target detection module;
[0097] First, an RPN is introduced to generate candidate boxes. The RPN slides across the feature map to generate multiple candidate boxes. The candidate regions are then mapped to a fixed-size feature map using a RoI pooling layer. Finally, the classification head determines the category of the candidate box. Simultaneously, the regression network adjusts the candidate box boundaries for greater precision.
[0098] S3: Construct loss function. The loss function includes RPN binary cross entropy loss, bounding box classification loss, bounding box regression loss and smooth variational autoencoder loss, specifically:
[0099] S3.1: Construct RPN binary cross entropy loss. RPN binary cross entropy loss is used to optimize the model's ability to distinguish foreground from background, which can be expressed as:
[0100]
[0101] Among them, y i represents the true label of the i-th sample, is the predicted foreground probability, and N represents the number of samples involved in training.
[0102] S3.2: Construct bounding box classification loss. The bounding box classification loss is used to classify different targets and optimize the model's ability to distinguish different categories. It can be expressed as:
[0103]
[0104] Among them, y ic represents the true category label of the i-th sample, Represents the category probability predicted by the model, and C represents the number of categories involved in training.
[0105] S3.3: Construct a smooth L1 loss for bounding box regression. The smooth L1 loss for bounding box regression is used for accurate positioning of the bounding box and can be expressed as:
[0106]
[0107] Among them, y ib represents the true bounding box parameters, Represents the predicted bounding box parameters.
[0108] S3.4: Construct the loss of the smooth variational autoencoder. The loss of the smooth variational autoencoder is used to constrain the latent variables of the encoding so that it approaches the standard normal distribution, which can be expressed as:
[0109] L s-VAE =L recon +D KL (q φ(z|x) ||p(z))+βD(q φ (z)||p(z)) (9)
[0110] Among them, L recon represents the reconstruction loss calculated using log-cosh and smooth L1 loss, D KL (q φ(z|x) ||p(z)) represents the output q of the encoder φ(z|x) KL divergence between the prior distribution p(z), D(q φ (z)||p(z)) represents the decoupled regularization term, and β is D(qφ(z|x) The weight of the term ||p(z)).
[0111] Finally, the total loss function during the training phase can be expressed as:
[0112] L base =L RPN +L cls +L reg +aL S-VAE (10)
[0113] Among them, a is used to balance the weights of S-VAE loss and other loss terms of S-EDH-Faster RCNN.
[0114] S4: Perform adaptive fine-tuning training. The adaptive fine-tuning training method automatically selects the top k important convolution kernels in the backbone network to update the network. The adaptive fine-tuning ratio is set to 26%, allowing the backbone network to extract features of new class samples.
[0115] The L2 norm is selected as the importance index of the convolution kernel, and the calculation method is:
[0116]
[0117] Among them, W i represents the i-th convolution kernel, w ij Represents the value of the jth parameter in the i-th convolution kernel.
[0118] After obtaining the L2 norm of the convolution kernel, the convolution kernels are sorted and the top k convolution kernels are selected for update. Adaptive fine-tuning can effectively improve the generalization ability of the model under small sample conditions.
[0119] S5: Test the trained few-shot object detection network based on the smooth variational autoencoder and adaptive fine-tuning on the test set. Calculate the average accuracy of the three new class sets (split1, split2, and split3) for different numbers of samples (1, 2, 3, 5, and 10 shots), the average accuracy of the base class, and the average accuracy of the new class. The average accuracy (%) of the model with and without the smooth variational autoencoder and adaptive fine-tuning on split1 is shown in Table 1, the average accuracy (%) of the model with and without the smooth variational autoencoder and adaptive fine-tuning on split2 is shown in Table 2, and the average accuracy (%) of the model with and without the smooth variational autoencoder and adaptive fine-tuning on split3 is shown in Table 3.
[0120] Table 1
[0121]
[0122] Table 2
[0123]
[0124] Table 3
[0125]
[0126] in conclusion:
[0127] For split1, the small sample pipeline DR defect detection method based on smooth variational autoencoder and adaptive fine-tuning has an average detection accuracy of 43.28%, 49.56%, 58.09%, 62.67% and 68.07% at 1shot, 2shots, 3shots, 5shots and 10shots, respectively.
[0128] For split2, the small-sample pipeline DR defect detection method based on smooth variational autoencoder and adaptive fine-tuning achieved average accuracies of 43.77%, 50.77%, 51.24%, 55.92% and 65.46% when the number of small-sample category images was 1, 2, 3, 4, 5 and 10 shots, respectively.
[0129] For split3, when the number of samples in the small sample category is 2, 3, 5 and 10, the small sample pipeline DR defect detection method based on smooth variational autoencoder and adaptive fine-tuning achieved an average accuracy of 46.93%, 50.17%, 57.43% and 67.80%, respectively.
[0130] Experimental results show that this method significantly improves the accuracy of pipeline defect detection in small sample conditions and achieves high detection accuracy on multiple test sets; this method has broad industrial application prospects, especially for safety monitoring and health assessment of gas pipelines.
[0131] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A small sample pipeline DR defect detection method, characterized in that: include: Step 1: Obtain the pipeline DR defect dataset X and the corresponding label Y, and split the dataset; wherein, X is divided into the training set X according to the ratio of N1:N2 train and the test set X test , N1 and N2 are both greater than 0 and N1>N2; all categories C corresponding to the dataset are divided into base classes and new classes according to the ratio of M1:M2, M1 and M2 are both greater than 0 and M1>M2; each new class considers a small sample of annotated training samples within K in turn; Step 2: Build a small sample target detection model based on Faster RCNN, which consists of three parts: backbone network, smooth variational autoencoder and target detection module; Step 3: Construct a loss function, which includes RPN binary cross entropy loss, bounding box classification loss, bounding box regression loss, and smooth variational autoencoder loss; Step 4: Adaptively fine-tune the small sample target detection model. The L2 norm is used to automatically select the top k important convolution kernels in the backbone network to update the network. The adaptive fine-tuning ratio is 20% < 30%, so that the backbone network can extract the features of the new class of samples. Step 5: Test the trained small-sample object detection network based on smooth variational autoencoder and adaptive fine-tuning on the test set, and calculate the average accuracy of different numbers of samples on different new class sets, the average accuracy of the base class, and the average accuracy of the new class.
2. The small sample pipeline DR defect detection method according to claim 1, characterized in that: N1:N2 is 7:3, M1:M2 is 3:1, K≤10.
3. The small sample pipeline DR defect detection method according to claim 1, characterized in that: The step 2 specifically includes: Step 2.
1. Build the backbone network: Use ImageNet pre-trained Resnet-101 as the backbone network and adopt stochastic gradient descent as the optimization method; Step 2.2: Build a smooth variational autoencoder: First, the smooth variational autoencoder uses an encoder to map the input image features to a latent space, generating a set of latent variables, including the mean and variance of the latent variables; Then, the loss functions of log-cosh and smooth L1 are introduced. Log-cosh and smooth L1 are expressed as: Among them, h is the feature of the real data, represents the features reconstructed by S-VAE, and t represents the error between the actual value and the predicted value z is the latent variable, φ is the parameter of the encoder, q φ (z) is the distribution of the latent variable z, p(z) is the prior distribution, and a is a positive hyperparameter used to adjust the sensitivity of the loss function to the error; The reconstruction loss is expressed as: L recon (|t|)=αL log-cosh (|t|)+(1-α)L smootL1 (|t|) (3) Among them, α is a hyperparameter between 0 and 1, which is used to control the contribution of the two losses; Finally, the distance between latent variables is used as a regularization term to decouple the latent variables. The regularization term is expressed as: in, Represents the sum of the off-diagonal elements of the covariance matrix, which is used to penalize the off-diagonal elements of the covariance matrix, ∑ i ([Cov qφ(z) [z]] ii -1) 2 Represents the sum of the squares of the differences between the diagonal elements of the covariance matrix and 1, which is used to adjust the variance matrix of the latent variables to ensure that the variance of each latent variable is close to 1 and maintain an appropriate degree of decoupling in the latent space; od and λ d is a hyperparameter; To sum up, the total loss of S-VAE is: L S-VAe =L recon (|t|)+D KL (q φ(z|x) ||p(z))+βD(q φ (z)||p(z)) (5) Among them, L recon (|t|) represents the reconstruction loss calculated using log-cosh and smooth L1 loss, D KL (q φ(z|x) ||p(z)) represents the output q of the encoder φ(z|x) KL divergence between the prior distribution p(z), D(q φ (z)||p(z)) represents the decoupled regularization term, and β is D(q φ The weight of the term (z)||p(z)); Step 2.3: Build the target detection module: First, RPN is introduced to generate candidate boxes. RPN slides on the feature map to generate multiple candidate boxes. Then, the candidate regions are uniformly mapped to a fixed-size feature map through the RoI pooling layer; Finally, the classification head is used to determine the category of the candidate box, and the regression network is used to adjust the boundary of the candidate box.
4. The small sample pipeline DR defect detection method according to claim 1, characterized in that: The step 3 specifically includes: Step 3.1: Construct the RPN binary cross entropy loss. The RPN binary cross entropy loss is used to optimize the model's ability to distinguish foreground from background, and is expressed as: Among them, y i represents the true label of the i-th sample, is the predicted foreground probability, N represents the number of samples involved in training; Step 3.2: Construct the bounding box classification loss. The bounding box classification loss is used to classify different targets and optimize the model's ability to distinguish different categories. It is expressed as: Among them, y ic represents the true category label of the i-th sample, Represents the category probability predicted by the model, and C represents the number of categories involved in training; Step 3.3: Construct a smooth L1 loss for bounding box regression. The smooth L1 loss for bounding box regression is used for accurate positioning of the bounding box and is expressed as: Among them, y ib represents the true bounding box parameters, represents the predicted bounding box parameters; Step 3.4: Construct the loss of the smooth variational autoencoder. The loss of the smooth variational autoencoder is used to constrain the encoded latent variables so that they are close to the standard normal distribution, which is expressed as: L S-VA =L recon +D KL (q φ(z|x) ||p(z))+βD(q φ (z)||p(z)) (9) Among them, L recon represents the reconstruction loss calculated using log-cosh and smooth L1 loss, D KL (q φ(z|x) ||p(z)) represents the output q of the encoder φ(z|x) KL divergence between the prior distribution p(z), D(q φ (z)||p(z)) represents the decoupled regularization term, and β is D(q φ(z|x) The weight of the term ||p(z)); Finally, the total loss function of the training phase is expressed as: L base =L RPN +L cls +L reg +aL S-VAE (10) Among them, a is used to balance the weights of S-VAE loss and other loss terms of S-EDH-Faster RCNN.
5. The small sample pipeline DR defect detection method according to claim 1, characterized in that: The step 4 specifically includes: The L2 norm is selected as the importance index of the convolution kernel, and the calculation method is: Among them, W i represents the i-th convolution kernel, w ij Represents the value of the jth parameter in the i-th convolution kernel; After obtaining the L2 norm of the convolution kernel, the convolution kernels are sorted and the top k convolution kernels are selected for network update.