A controllable generation method for bolt defects based on perspective and attribute guidance
Through the generative adversarial network guided by perspective and attributes, the problem of insufficient bolt defect data in the transmission line is solved, and high-quality bolt defect images are generated, which improves the recognition accuracy and meets the data needs of the power system.
Patent Information
- Application Number
- CN202411510992.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-10-28
AI Technical Summary
The scarcity of bolt defect samples in the transmission line leads to insufficient training of the generative adversarial network and deviation of the generated image from the actual scene, affecting the accuracy of bolt defect recognition.
Introduce viewing angle and attribute information, use residual cross-layer excitation and U-Net self-supervised reconstruction discriminator to build a generative adversarial network, generate bolt defect images of specific viewing angles and attributes, and enhance feature extraction and texture generation capabilities through global and local reconstruction.
High-quality bolt defect images can be controlledly generated under the condition of few samples, which improves the accuracy of the bolt defect identification network and meets the data needs of the power system.
Smart Images

Figure CN119540713B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image analysis, and particularly to a controllable generation method for bolt defects guided by perspective and attributes. Background Art
[0002] As an important part of the power system, the stable, safe and efficient operation of transmission lines is crucial for ensuring the continuity and reliability of power supply. Among the numerous components of transmission lines, bolts play a vital role in fixing and supporting various equipment and structures of the lines. However, due to the usually harsh operating conditions of transmission lines, bolts are prone to defects due to environmental and other factors. Therefore, timely and accurate detection of bolt defects has become the key to ensuring the safe operation of the lines.
[0003] In recent years, the intelligent inspection technology of unmanned aerial vehicles has been widely applied in transmission lines, with the advantages of high safety and high efficiency, and can be combined with object detection and intelligent recognition algorithms of deep learning to achieve intelligent processing. However, in actual transmission lines, due to various reasons, the defect samples of bolts are often very scarce. This few-shot problem poses a huge challenge to the automatic detection and classification of bolt defects. Traditional machine learning and deep learning methods usually perform well in the face of a large number of samples, but in few-shot scenarios, their performance is often severely limited. Therefore, how to improve the model performance under few-shot conditions is one of the urgent problems to be solved currently.
[0004] In the case of limited data, how to make full use of and expand samples has become the focus of research. Traditional data augmentation methods, such as geometric transformation, color transformation and pixel transformation, expand the dataset and optimize the image quality through image processing techniques. However, these methods are not applicable to all data augmentation operations and face great application limitations in some cases. Generative adversarial network is an excellent data augmentation strategy. Goodfellow first proposed the generative adversarial network (GAN) in 2014. GAN conducts adversarial training by constructing a set of generators and discriminators and reaches the Nash equilibrium in the game process. As one of the fastest-growing models in recent years, GAN and its variants have achieved great success in image generation, style transfer, image super-resolution, data augmentation, etc. However, there are still great challenges in the generation of bolt defects in transmission lines by GAN.
[0005] 1) Insufficient training set of GAN: The training of the GAN model requires a large number of real bolt defect images. However, in actual transmission lines, there are too few bolts with defects. Therefore, it is easy to lose information during the training of GAN, resulting in the model being prone to collapse during training;
[0006] 2) There are significant visual differences in bolts from different perspectives: Usually, during the generation process, GAN classifies bolts with the same attribute into the same category, ignoring the visual differences between bolt images, which may lead to deviations between the generated images and the actual application scenarios. Due to the differences in the installation positions, environmental factors, and shooting perspectives of bolts, there are often significant visual differences in bolt images in transmission lines, which further increases the complexity of the generation task, affecting the generation trend and resulting in the generation of incorrect samples.
[0007] Therefore, in the above context, introducing the generative adversarial network into the field of power systems to solve the problem of insufficient data on bolt defects in current transmission lines, expanding the bolt defect dataset while assisting in improving the accuracy of bolt defect recognition is of utmost importance to meet the current industrialization requirements in the field of power systems. Summary of the Invention
[0008] The object of the present invention is to provide a controllable generation method for bolt defects guided by perspective and attribute, to solve problems such as insufficient data on bolt defects in transmission lines and poor effects of downstream tasks caused by it. By introducing perspective and attribute information, and proposing a residual cross-layer excitation and a U-Net self-supervised reconstruction discriminator USRD, a novel generative adversarial network is designed to controllably generate bolt defect images with specific perspectives and attributes under few-shot conditions, thereby improving the accuracy of bolt defect recognition.
[0009] To achieve the above object, the present invention provides the following solution:
[0010] A controllable generation method for bolt defects guided by perspective and attribute, comprising the following steps:
[0011] S1, construct a dataset Dataset 1 for training the bolt perspective label generator, including the front and side perspectives of normal bolts; construct a few-shot bolt defect dataset Dataset 2, including bolt defect images of three categories, and the three categories include normal bolts, missing pins, and missing nuts.
[0012] S2, adopt the contrastive language-image pre-training CLIP model, propose a bolt perspective label generator based on CLIP fine-tuning, automatically assign perspective labels to the few-shot bolt defect dataset Dataset 2, and construct a bolt defect dataset Dataset 3 containing perspective and attribute labels to provide data support for subsequent image generation tasks.
[0013] S3. Adopt the architecture of the conditional generative adversarial network, introduce perspective and attributes as additional conditional inputs in the generator, and propose the residual cross-layer excitation ResSLE to construct a generator with stronger gradient information flow and higher generation controllability. Among them, ResSLE introduces cross-layer connections on the basis of the cross-layer excitation SLE to strengthen the gradient signal between layers, and controls the intensity of the excitation through the Lambda layer to enhance the gradient flow of the generator.
[0014] S4. Adopt the encoder-decoder architecture to expand the discriminator to form a U-Net, propose the self-supervised reconstruction discriminator USRD based on the U-Net, and design a unique local cropping method Shape-crop for T-bolts. The feature extraction and detailed texture generation capabilities of the discriminator are enhanced through global and local reconstructions.
[0015] Among them, S2 specifically includes:
[0016] First, to generate more accurate bolt images, it is proposed to introduce perspective information based on bolt defect attributes in the GAN, and divide the bolt perspective into two categories with the largest visual differences: the front perspective of 0° and the side perspective of 90°, to ensure the feasibility and accuracy of the task.
[0017] Next, a bolt perspective label generator based on CLIP fine-tuning is proposed and trained on Dataset 1 to learn the features of the two perspectives. The weight space integration method WiSE-FT is used to fine-tune CLIP, and the formula is as follows:
[0018] wse(x,α)=f(x,(1-α)·θ0+α·θ1) (1)
[0019] Among them, x represents the input data, f is the classification prediction function of the model, θ0 represents the original model parameters, θ1 represents the model parameters obtained through standard fine-tuning, α is the mixing coefficient for controlling the weights, and the image encoder E of the fine-tuned CLIP is extracted as the bolt perspective label generator.
[0020] Finally, the bolt defect dataset Dataset 2 with attribute labels is input into the bolt perspective label generator to assign perspective labels to each image. By combining these perspective labels with the attribute labels of the bolts, a perspective-attribute bolt dataset Dataset 3 is constructed, and this process is defined as follows:
[0021] Label(D3)=E(D2)+Attribute(D2) (2)
[0022] Among them, E represents the bolt perspective label generator, Attribute is the attribute label of Dataset 2, and the categories of Dataset 3 include 6 categories: front - normal bolt, side - normal bolt, front - pin missing, side - pin missing, front - nut missing, and side - nut missing. The images in Dataset 2 and Dataset 3 are the same, only the category labels are different. The bolt perspective label generator reduces human bias and improves label consistency.
[0023] Among them, the architecture of the conditional generative adversarial network is adopted. Perspectives and attributes are introduced as additional conditional inputs in the generator, and the residual cross - layer excitation ResSLE is proposed to construct a generator with stronger gradient information flow and higher generation controllability, which specifically includes:
[0024] First, an n - dimensional conditional vector is added to the input N - dimensional random noise vector. After concatenating the two, they are input into the generator. The input is defined as shown in the following formula:
[0025] Z = concat(z, y) (3)
[0026] Among them, Z is the input vector of the generator, z is the noise vector, and y is the conditional vector. Labels containing perspectives and attributes are used as conditional information to assist in generation, increasing the prior knowledge of the generator and at the same time improving the controllability of GAN generation. Another reason for introducing conditions is that different bolt defect images usually have certain commonalities. Although the attributes or perspectives of bolts are different, the main bodies of bolts are similar, which is conducive to the learning of the generator.
[0027] In addition, the generator of the few - shot generation model FastGAN is used as the basic model, and the architecture of the generator is improved. The residual cross - layer excitation ResSLE is proposed. ResSLE introduces cross - layer connections on the basis of SLE to strengthen the gradient signal between layers and controls the intensity of excitation through the Lambda layer, effectively enhancing the gradient flow of the generator. The definition of ResSLE is as follows:
[0028] ResSLE = SLE(x)+λ·U(x) (4)
[0029] Among them, x represents the input feature, U() represents upsampling, and λ is used to control the excitation intensity. The improvement of the generator increases its own information content, aims to alleviate the unequal confrontation between the generator and the discriminator, and helps to improve the generation performance index and training stability.
[0030] Among them, the discriminator is extended using an encoder-decoder architecture to form a U-Net, a self-supervised reconstruction discriminator USRD based on the U-Net is proposed, and a unique local cropping method Shape-crop is designed for T-bolts. The feature extraction and detailed texture generation capabilities of the discriminator are enhanced through global and local reconstructions, specifically including:
[0031] In the training of GANs, the information between the generator and the discriminator is asymmetric, which often leads to the discriminator relying too much on non-target features to judge true or false, affecting the generation effect and falling into a vicious cycle. To solve this problem, a self-supervised reconstruction discriminator USRD based on the U-Net is proposed. Using an encoder-decoder architecture, USRD takes the feature extraction network of the discriminator as the encoder and constructs the decoder by reversing its structure to form a U-Net structure, including the encoder D enc the global decoder D dec1 and the local decoder D dec2 .
[0032] First, for the underlying architecture encoder D enc , a traditional CNN network is used to construct it. The upper layer architecture of D enc consists of a traditional adversarial head and a conditional classification head. Specifically, the underlying architecture of the discriminator is used as a feature extractor to extract the features of the input image, which are respectively sent to the adversarial head and the classification head. Similar to a general GAN, the adversarial head uses a sigmoid activation function to judge the true or false of the discriminant input image; the classification head uses a softmax activation function to classify the category of the input image and compare it with the input conditions of the generator, expecting the two to belong to the same category.
[0033] Second, the global decoder D dec1 takes the bottleneck layer output by the encoder D enc as the input and has skip connections with the feature maps of the encoder to ensure that the generated image is effectively aligned with the original data distribution globally. This means that the generated image is similar in overall structure and layout. Through global reconstruction, the model can learn to capture the diversity and accuracy of the data distribution from a global perspective. The architecture of the local decoder D dec2 is similar to that of the global decoder D dec1 . The difference is that the input of D dec2 is the feature cropped from the final output of the encoder D enc , rather than directly input from the bottleneck layer. It should be noted that the encoder D enc and the local decoder D dec2The jump connections between them all need to go through the corresponding cropping operations. The local decoder mainly focuses on generating the local regions of the image to ensure that these regions have high-quality textures and details. Another important reason for using local reconstruction is that the colors of the bolts and their background fittings are similar, and it is sometimes difficult to distinguish them well under lighting conditions. By reconstructing local pixels, the shape characteristics of the bolts can be highlighted, which is beneficial for the generator to accurately generate the bolt shapes. The global and local reconstruction losses are defined as follows:
[0034]
[0035] Among them, I represents the input image, L global Using the L2 norm, the significant differences are optimized in global reconstruction to ensure that the overall image is close to the original image; L local Using the L1 norm, the error of each pixel is focused on in local reconstruction to emphasize the accuracy of details and textures.
[0036] Among them, it also includes: aiming at the shape characteristics of T-shaped bolts, a shape cropping method Shape-crop is proposed. The bolt image is cropped in a cross-centered manner, and four squares are intercepted from the middle in the horizontal and vertical directions. At the same time, the central region containing the most bolt information is increased. The areas of the regions cropped by the two methods are the same, but the number of bolt pixels extracted by Shape-crop is significantly more than that of the random cropping method to improve the reconstruction quality of D dec2 of D.
[0037] Among them, it also includes: using the bolt defect controllable generation method to assist in improving the accuracy of the bolt defect recognition network and further verifying the effectiveness of the bolt defect controllable generation method, specifically including:
[0038] To verify the effectiveness in generating few-shot bolt defect images, experiments were conducted on the classification task. The experimental process is divided into three steps: First, use the generation method to generate a large number of bolt defect images; then, use these generated images to augment the dataset; finally, use the augmented dataset to train the bolt defect classification model and evaluate the classification accuracy of the model on the test set. The final result significantly improves the accuracy of the bolt defect classification model, proving the effectiveness of the generation method.
[0039] This application also provides a bolt defect controllable generation device based on view and attribute guidance. The device includes:
[0040] The S1 module is used to construct the dataset Dataset 1 for training the bolt view label generator, which includes the front and side views of normal bolts; construct the few-shot bolt defect dataset Dataset 2, which includes three categories of bolt defect images, and the three categories include normal bolts, pin missing, and nut missing.
[0041] The S2 module is used to adopt the contrastive language-image pre-training CLIP model, propose a bolt perspective label generator based on CLIP fine-tuning, automatically assign perspective labels to the few-shot bolt defect dataset Dataset 2, construct a bolt defect dataset Dataset 3 containing perspective and attribute labels, and provide data support for subsequent image generation tasks.
[0042] The S3 module is used to adopt the architecture of a conditional generative adversarial network, introduce perspective and attributes as additional conditional inputs in the generator, and propose a residual cross-layer excitation ResSLE to construct a generator with stronger gradient information flow and higher generation controllability; among them, ResSLE introduces cross-layer connections on the basis of the cross-layer excitation SLE to strengthen the gradient signal between layers, and controls the intensity of the excitation through the Lambda layer to enhance the gradient flow of the generator.
[0043] The S4 module is used to expand the discriminator using the encoder-decoder architecture to form a U-Net, propose a self-supervised reconstruction discriminator USRD based on the U-Net, and design a unique local cropping method Shape-crop for T-shaped bolts, enhancing the discriminator's feature extraction and detail texture generation capabilities through global and local reconstructions.
[0044] Among them, the S2 module specifically includes:
[0045] First, to generate more accurate bolt images, it is proposed to introduce perspective information based on bolt defect attributes into the GAN, divide the bolt perspective into two categories with the largest visual differences: the front perspective 0° and the side perspective 90°, to ensure the feasibility and accuracy of the task.
[0046] Next, a bolt perspective label generator based on CLIP fine-tuning is proposed and trained on Dataset 1 to learn the features of the two perspectives. The WiSE-FT method for fine-tuning CLIP in the weight space is adopted, and the formula is as follows:
[0047] wse(x,α)=f(x,(1-α)·θ0+α·θ1) (1)
[0048] Where x represents the input data, f is the classification prediction function of the model, θ0 represents the original model parameters, θ1 represents the model parameters obtained through standard fine-tuning, α is the mixing coefficient for controlling the weights, and the image encoder E of the fine-tuned CLIP is extracted as the bolt perspective label generator.
[0049] Finally, the bolt defect dataset Dataset 2 with attribute labels is input into the bolt perspective label generator to assign perspective labels to each image. By combining these perspective labels with the attribute labels of the bolts, a perspective-attribute bolt dataset Dataset 3 is constructed. This process is defined as follows:
[0050] Label(D3) = E(D2) + Attribute(D2) (2)
[0051] Among them, E represents the bolt perspective label generator, Attribute is the attribute label of Dataset 2. The categories of Dataset 3 include 6 categories: front-normal bolt, side-normal bolt, front-pin missing, side-pin missing, front-nut missing, and side-nut missing. The images in Dataset 2 and Dataset 3 are the same, only the category labels are different. The bolt perspective label generator reduces human bias and improves label consistency.
[0052] This application also provides an electronic device, which includes: at least one memory and at least one processor; the memory stores a program, and the processor calls the program stored in the memory, and the program is used to implement the controllable generation method of bolt defects guided by perspective and attribute.
[0053] This application also provides a storage medium, in which computer-executable instructions are stored, and the computer-executable instructions are used to execute the controllable generation method of bolt defects guided by perspective and attribute.
[0054] The present invention discloses the following technical effects: A method, device, electronic device, and storage medium for controllably generating bolt defects based on perspective and attribute guidance. Specifically, a bolt perspective label generator based on contrastive language-image pre-training is designed to automatically assign perspective labels to the bolt dataset, thereby alleviating the visual difference problem of bolt images and providing data support for the image generation task. A generative adversarial network is constructed, and the generator of FastGAN is selected as the basic model for the generation part. Perspective and attribute information are introduced as conditions in the generator to guide the generation, and a residual cross-layer excitation ResSLE is proposed to strengthen the gradient information flow and increase the controllability of the generation. In the discriminant part, a U-Net self-supervised reconstruction discriminator USRD is proposed, and a unique local cropping method Shape-crop is designed for T-shaped bolts, thereby enhancing the generation and feature extraction capabilities of bolt texture details. This bolt defect generation method can controllably generate bolt defect images with specific perspectives and attributes under the condition of a small number of samples, and assist in improving the accuracy of the bolt defect recognition network from the data level. The present invention applies the generative adversarial network to the generation of few-shot bolt defect images, and effectively improves the quality and controllability of the generated bolt defect images by combining perspectives and attributes as conditions to guide the generation of the GAN, meeting the data requirements of downstream tasks for few-shot bolt defects. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0056] Figure 1 is a schematic flow chart of a method for controllably generating bolt defects based on perspective and attribute guidance according to an embodiment of the present invention;
[0057] Figure 2 is a schematic structural diagram of a bolt perspective label generator according to an embodiment of the present invention;
[0058] Figure 3 is a schematic structural diagram of a residual cross-layer excitation according to an embodiment of the present invention;
[0059] Figure 4 is a schematic diagram of a shape cropping method according to an embodiment of the present invention;
[0060] Figure 5 is a schematic overall structural diagram according to an embodiment of the present invention;
[0061] Figure 6 is an effect diagram of generating bolt defects according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0063] The purpose of the present invention is to provide a controllable generation method for bolt defects guided by perspective and attributes, which solves problems such as insufficient bolt defect data in transmission lines and poor effects of downstream tasks caused by it. By introducing perspective and attribute information, and proposing residual cross-layer excitation and the U-Net self-supervised reconstruction discriminator USRD, a novel generative adversarial network is designed to controllably generate bolt defect images with specific perspectives and attributes under few-shot conditions, so as to achieve the purpose of improving bolt defect recognition.
[0064] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0065] As Figure 1 shown, a controllable generation method for bolt defects guided by perspective and attributes provided by the present invention includes the following steps:
[0066] S1. Construct a dataset Dataset 1 for training a bolt perspective label generator, which only includes the front and side views of normal bolts; construct a few-shot bolt defect dataset Dataset 2, which includes three categories of bolt defect images, and the three categories include normal bolts, missing pins, and missing nuts.
[0067] S2. Adopt contrastive language-image pre-training, propose a bolt perspective label generator based on CLIP fine-tuning, automatically assign perspective labels to the few-shot bolt defect dataset Dataset 2, and construct a bolt defect dataset Dataset 3 containing perspective and attribute labels to provide data support for subsequent image generation tasks.
[0068] S3. Adopt the architecture of a conditional generative adversarial network, introduce perspective and attributes as additional conditional inputs into the generator, and propose residual cross-layer excitation ResSLE to construct a generator with stronger gradient information flow and higher generation controllability; among them, ResSLE introduces cross-layer connections on the basis of cross-layer excitation SLE to strengthen the gradient signal between layers, and controls the intensity of the excitation through the Lambda layer to enhance the gradient flow of the generator.
[0069] In S4, the discriminator is extended to form a U-Net using an encoder-decoder architecture, a self-supervised reconstruction discriminator based on the U-Net is proposed, and a unique local cropping method Shape-crop is designed for T-bolts, enhancing the discriminator's feature extraction and detailed texture generation capabilities through global and local reconstructions;
[0070] In S5, a bolt defect controllable generation method is used to assist in improving the accuracy of the bolt defect recognition network, further verifying the effectiveness of the bolt defect controllable generation method.
[0071] The basic flowchart of the present invention is as Figure 1 shown.
[0072] When a deep learning model is trained, a large number of dataset image samples are required as support. Since the drone aerial inspection images are usually global images of transmission lines, they need to be cropped and preprocessed to unify the size of the dataset, and the defect categories of the bolts need to be further divided and labeled. In addition, for different tasks in the present invention, different datasets need to be constructed. Therefore, in the step S1, a dataset Dataset 1 for training the bolt perspective label generator is constructed, which only contains the front and side views of normal bolts; a few-shot bolt defect dataset Dataset 2 is constructed, which contains bolt defect images of different categories, specifically including:
[0073] Collect drone aerial inspection images of transmission lines, crop the areas containing bolts, clean them, and select the images with clear images, more types and quantities of bolt defects. Perform data preprocessing to unify the scale of the bolt images and label their defect types. Then, datasets are constructed for different tasks: 1) Dataset 1 contains 1000 normal bolt images with front view (0°) and 1000 normal bolt images with side view (90°) for training the bolt perspective label generator; 2) Dataset 2 contains 2260 images of 3 categories including normal bolts, missing pins, and missing nuts.
[0074] The structural schematic diagram of the bolt perspective label generator in the present invention is as Figure 2 shown.
[0075] In the present invention, considering the significant visual morphological differences of bolts in transmission lines, it is necessary to introduce the perspective and attribute information of bolts into the generative adversarial network. At the same time, due to the subjective deviation in the artificial perspective division, a model-based method is required to automatically divide the bolt perspectives. This method proposes a bolt perspective label generator based on CLIP fine-tuning to reduce human bias and improve label consistency. Among them, in the step S2, a contrastive language-image pre-training model is adopted, and a bolt perspective label generator based on CLIP fine-tuning is proposed to automatically assign perspective labels to the few-shot bolt defect dataset Dataset 2, and construct a bolt defect dataset Dataset 3 containing perspective and attribute labels to provide data support for subsequent image generation tasks, specifically including:
[0076] First of all, in transmission lines, due to the different shooting angles of drones and the positions of bolts, the visual morphologies of bolts are diverse and are often misclassified as the same visual words, which affects the feature extraction performance, especially has a significant impact on GAN. Most of the existing methods rely on bolt defect attribute classification and ignore the perspective factor. To generate more accurate bolt images, it is proposed to introduce perspective information based on bolt defect attributes into GAN. Due to the lack of a standardized inspection dataset, the bolt perspectives in the present invention are divided into two categories with the largest visual differences: the front perspective (0°) and the side perspective (90°) to ensure the feasibility and accuracy of the task.
[0077] Next, since there is a subjective deviation in the artificial division of the middle perspective, a bolt perspective label generator based on CLIP fine-tuning is proposed and trained on Dataset 1 to learn the features of the two perspectives. The present invention uses the Weight Space Integration method WiSE-FT to fine-tune CLIP, and the formula is as follows:
[0078] wse(x,α)=f(x,(1-α)·θ0+α·θ1) (1)
[0079] where x represents the input data, f is the classification prediction function of the model, θ0 represents the original model parameters, θ1 represents the model parameters obtained through standard fine-tuning, and α is the mixing coefficient that controls the weights. The image encoder E of the fine-tuned CLIP is extracted as the bolt perspective label generator.
[0080] Finally, the bolt defect dataset Dataset 2 with attribute labels is input into the bolt perspective label generator to assign perspective labels to each image. By combining these perspective labels with the attribute labels of the bolts, a perspective-attribute bolt dataset Dataset 3 is constructed. This process is defined as follows:
[0081] Label(D3)=E(D2)+Attribute(D2) (2)
[0082] Among them, E represents the bolt perspective label generator, and Attribute is the attribute label of Dataset 2. The categories of Dataset 3 include 6 categories: front - normal bolt, side - normal bolt, front - pin missing, side - pin missing, front - nut missing, and side - nut missing. Note that the images in Dataset 2 and Dataset 3 are the same, only the category labels are different. The bolt perspective label generator reduces human bias and improves label consistency, which provides reliable data support for subsequent generation tasks.
[0083] The structural schematic diagram of the residual cross - layer excitation in the present invention is as Figure 3 shown.
[0084] In step S3, the architecture of the conditional generative adversarial network is adopted. Perspectives and attributes are introduced as additional conditional inputs into the generator, and a residual cross - layer excitation ResSLE is proposed to construct a generator with stronger gradient information flow and higher generation controllability, specifically including:
[0085] First, for the problem of generating few - shot bolt defects, more useful information is input into the generator. An n - dimensional conditional vector is added to the input N - dimensional random noise vector, and the two are concatenated and then input into the generator. The input is defined as shown in the following formula:
[0086] Z = concat(z, y) (3)
[0087] Among them, Z is the input vector of the generator, z is the noise vector, and y is the conditional vector. Using labels containing perspectives and attributes as conditional information to assist in generation increases the prior knowledge of the generator and improves the controllability of GAN generation. Another reason for introducing the condition is that different bolt defect images usually have certain commonalities. Although the attributes or perspectives of the bolts are different, the main bodies of the bolts are similar, which is beneficial for the learning of the generator.
[0088] In addition, the present invention uses the generator of the few-shot generation model FastGAN as the basic model and improves the architecture of the generator. FastGAN uses the model lightweight design of SLE in the generator, which can alleviate the problem of gradient vanishing, but there are still some problems. SLE exchanges the reduction of computational cost at the expense of certain information loss, which is necessary for the generation of high-resolution images of 1024×1024. However, it is not optimal for the task of bolt defect generation. Bolt images usually have a lower resolution, require less computational resources, and have higher requirements for gradient information. Based on this discovery, residual cross-layer excitation is proposed. ResSLE introduces cross-layer connections on the basis of SLE to strengthen the gradient signal between layers and controls the intensity of excitation through the Lambda layer, effectively improving the gradient flow of the generator. The definition of ResSLE is as follows:
[0089] ResSLE = SLE(x) + λ·U(x) (4)
[0090] Where x represents the input feature, U() represents upsampling, and λ is used to control the intensity of excitation. The improvement of the generator increases its own information content, aims to alleviate the unequal confrontation between the generator and the discriminator, and helps to improve the generation performance index and training stability.
[0091] The schematic diagram of the shape cropping method in the present invention is as Figure 4 shown.
[0092] In step S4, the discriminator is extended using an encoder-decoder architecture to form a U-Net, a self-supervised reconstruction discriminator USRD based on the U-Net is proposed, and a unique local cropping method Shape-crop is designed for T-shaped bolts, enhancing the discriminator's feature extraction and detailed texture generation capabilities through global and local reconstructions, specifically including:
[0093] In the training of GAN, the information between the generator and the discriminator is asymmetric, often leading to the discriminator relying too much on non-target features to judge true or false, affecting the generation effect and falling into a vicious cycle. To solve this problem, the present invention proposes a self-supervised reconstruction discriminator USRD based on the U-Net, using an encoder-decoder architecture. USRD uses the discriminator's feature extraction network as the encoder and constructs the decoder by reversing its structure to form a U-Net structure, including the encoder D enc 、the global decoder D dec1 and the local decoder D dec2 .
[0094] First, for the underlying architecture encoder D enc , a traditional CNN network is used to construct. D encThe upper - layer architecture consists of a traditional adversarial head and a conditional classification head. Specifically, the underlying architecture of the discriminator acts as a feature extractor to extract the features of the input image, which are then sent to the adversarial head and the classification head respectively. Similar to a general GAN, the adversarial head uses the sigmoid activation function to distinguish the authenticity of the input image; the classification head uses the softmax activation function to classify the category of the input image and compares it with the input condition of the generator, expecting them to belong to the same category.
[0095] Secondly, the global decoder D dec1 takes the bottleneck layer output by the encoder D enc as the input and has skip connections with the feature maps of the encoder. This is used to ensure that the generated image is effectively aligned with the original data distribution globally, which means that the generated image is similar in overall structure and layout. Through global reconstruction, the model can learn to capture the diversity and accuracy of the data distribution from a global perspective. The local decoder D dec2 has a similar architecture to the global decoder D dec1 , the difference is that the input of D dec2 is the feature cropped from the final output of the encoder D enc , rather than directly input from the bottleneck layer. It should be noted that the skip connections between the encoder D enc and the local decoder D dec2 both require corresponding cropping operations. The local decoder mainly focuses on the local regions of the generated image to ensure that these regions have high - quality textures and details. Another important reason for using local reconstruction is that the colors of the bolts and their background fittings are similar, and it is sometimes difficult to distinguish them well under lighting conditions. By reconstructing local pixels, the shape characteristics of the bolts can be highlighted, which is beneficial for the generator to accurately generate the shape of the bolts. The global and local reconstruction losses are defined as follows:
[0096]
[0097] where I represents the input image. L global uses the L2 norm to optimize the significant differences in global reconstruction, ensuring that the overall image is close to the original image; L local uses the L1 norm to focus on the error of each pixel in local reconstruction, emphasizing the accuracy of details and textures.
[0098] Finally, in the local reconstruction part of the discriminator, traditional random cropping methods often intercept regions containing a large amount of useless background, affecting the extraction of local information of bolt defect images, especially when the bolt size is small and the key information is concentrated in the center. To solve this problem, aiming at the shape characteristics of T-shaped bolts, a shape cropping method, Shape-crop, is proposed. The bolt image is cropped in a cross-centered manner, and four squares are intercepted from the middle in the horizontal and vertical directions, while the central region containing the most bolt information is increased. The areas of the regions cropped by the two methods are the same, but the number of bolt pixels extracted by Shape-crop is significantly more than that of random cropping, thus improving the dec2 reconstruction quality of D.
[0099] USRD reconstructs real images in a self-supervised manner. The encoder is forced to learn more useful information about the data distribution of real images to improve the overall performance of the GAN. At the same time, during the training process of the GAN, the reconstruction of real images by the decoder can be used as an additional regularization term, which helps to stabilize the training.
[0100] In step S5, a bolt defect controllable generation method is used to assist in improving the accuracy of the bolt defect recognition network, and further verify the effectiveness of the bolt defect controllable generation method, specifically including:
[0101] To verify the effectiveness of the present invention in generating few-shot bolt defect images, experiments on the classification task were carried out. The experimental process is divided into three steps: First, a large number of bolt defect images are generated using the method proposed in this study; then, these generated images are used to augment the dataset; finally, the augmented dataset is used to train a bolt defect classification model, and the classification accuracy of the model is evaluated on the test set. The final results significantly improve the accuracy of the bolt defect classification model, demonstrating the effectiveness of the method proposed in the present invention.
[0102] A bolt defect controllable generation method based on view and attribute guidance according to the present invention, the network structure of the generation method is as Figure 5 shown.
[0103] The bolt defect generation effect of the method of the present invention is as Figure 6As shown in the figure. The present invention designs a bolt perspective label generator based on contrastive language-image pre-training, which automatically assigns perspective labels to the bolt dataset, thereby alleviating the visual difference problem of bolt images and providing data support for image generation tasks. A generative adversarial network is constructed, and the generator of FastGAN is selected as the basic model for the generation part. Perspective and attribute information are introduced as conditions in the generator to guide the generation, and a residual cross-layer excitation ResSLE is proposed to strengthen the gradient information flow and increase the controllability of the generation. In the discriminant part, a U-Net self-supervised reconstruction discriminator USRD is proposed, and a unique local cropping method Shape-crop is designed for T-shaped bolts, thereby enhancing the generation and feature extraction capabilities of bolt texture details. This bolt defect generation method can controllably generate bolt defect images with specific perspectives and attributes under the condition of a small number of samples, and assist in improving the accuracy of the bolt defect recognition network from the data level. The present invention applies the generative adversarial network to the generation of few-shot bolt defect images, effectively improves the quality and controllability of the bolt defect image generation by combining perspectives and attributes as conditions to guide the generation of GAN, and meets the data requirements of downstream tasks for few-shot bolt defects.
[0104] Based on the method for controllable generation of bolt defects guided by perspective and attribute provided in the above embodiments, the embodiments of the present invention correspondingly provide a device for controllable generation of bolt defects guided by perspective and attribute, and the device includes:
[0105] The S1 module is used to construct a dataset Dataset 1 for training the bolt perspective label generator, including two perspectives, the front and side of normal bolts; construct a few-shot bolt defect dataset Dataset 2, including three categories of bolt defect images, and the three categories include normal bolts, missing pins, and missing nuts.
[0106] The S2 module is used to adopt a contrastive language-image pre-training CLIP model, propose a bolt perspective label generator based on CLIP fine-tuning, automatically assign perspective labels to the few-shot bolt defect dataset Dataset 2, and construct a bolt defect dataset Dataset 3 including perspective and attribute labels to provide data support for subsequent image generation tasks.
[0107] The S3 module is used to adopt the architecture of a conditional generative adversarial network, introduce perspectives and attributes as additional conditional inputs in the generator, and propose a residual cross-layer excitation ResSLE to construct a generator with stronger gradient information flow and higher generation controllability; among them, ResSLE introduces cross-layer connections on the basis of the cross-layer excitation SLE to strengthen the gradient signal between layers, and controls the intensity of the excitation through the Lambda layer to enhance the gradient flow of the generator.
[0108] The S4 module is used to expand the discriminator using an encoder-decoder architecture to form a U-Net, propose a self-supervised reconstruction discriminator USRD based on U-Net, and design a unique local cropping method Shape-crop for T-bolts, enhancing the discriminator's feature extraction and detailed texture generation capabilities through global and local reconstructions.
[0109] Among them, the S2 module specifically includes:
[0110] First, to generate more accurate bolt images, it is proposed to introduce perspective information based on bolt defect attributes into the GAN, and divide the bolt perspectives into two categories with the largest visual differences: the front perspective of 0° and the side perspective of 90°, to ensure the feasibility and accuracy of the task;
[0111] Next, a bolt perspective label generator based on CLIP fine-tuning is proposed and trained on Dataset 1 to learn the features of the two perspectives. The weight space integration method WiSE-FT is used to fine-tune CLIP, and the formula is as follows:
[0112] wse(x,α)=f(x,(1-α)·θ0+α·θ1) (1)
[0113] Where x represents the input data, f is the classification prediction function of the model, θ0 represents the original model parameters, θ1 represents the model parameters obtained through standard fine-tuning, α is the mixing coefficient controlling the weights, and the image encoder E of the fine-tuned CLIP is extracted as the bolt perspective label generator.
[0114] Finally, the bolt defect dataset Dataset 2 with attribute labels is input into the bolt perspective label generator to assign perspective labels to each image. By combining these perspective labels with the attribute labels of the bolts, a perspective-attribute bolt dataset Dataset 3 is constructed. This process is defined as follows:
[0115] Label(D3)=E(D2)+Attribute(D2) (2)
[0116] Where E represents the bolt perspective label generator, Attribute is the attribute label of Dataset 2. The categories of Dataset 3 include 6 categories: front-normal bolt, side-normal bolt, front-pin missing, side-pin missing, front-nut missing, and side-nut missing. The images in Dataset 2 and Dataset 3 are the same, only the category labels are different. The bolt perspective label generator reduces human bias and improves label consistency.
[0117] Based on the method for controllable generation of bolt defects guided by perspective and attributes provided in the above embodiments, an embodiment of the present invention correspondingly provides an electronic device, which includes: at least one memory and at least one processor; the memory stores a program, and the processor calls the program stored in the memory, and the program is used to implement the above method for controllable generation of bolt defects guided by perspective and attributes.
[0118] Based on the method for controllable generation of bolt defects guided by perspective and attributes provided in the above embodiments, an embodiment of the present invention correspondingly provides a storage medium, in which computer-executable instructions are stored, and the computer-executable instructions are used to execute the above method for controllable generation of bolt defects guided by perspective and attributes.
[0119] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
[0120] It should be noted that the embodiments in this specification are all described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions in the method part.
[0121] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements inherent to these process, method, article or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0122] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A controllable generation method for bolt defects based on perspective and attribute guidance, characterized in that It includes the following steps: S1. Construct a dataset Dataset 1 for training a bolt perspective label generator, which includes two perspectives, the front and side views of normal bolts; construct a few-shot bolt defect dataset Dataset 2, which includes three categories of bolt defect images, and the three categories are normal bolts, missing pins, and missing nuts; S2. Adopt the contrastive language-image pre-training CLIP model, propose a bolt perspective label generator based on CLIP fine-tuning, automatically assign perspective labels to the few-shot bolt defect dataset Dataset 2, and construct a bolt defect dataset Dataset 3 containing perspective and attribute labels to provide data support for subsequent image generation tasks; S3. Adopt the architecture of a conditional generative adversarial network, introduce perspective and attribute as additional conditional inputs into the generator, and propose a residual cross-layer excitation ResSLE to construct a generator with stronger gradient information flow and higher generation controllability; among them, ResSLE introduces cross-layer connections on the basis of the cross-layer excitation SLE to strengthen the gradient signal between layers, and controls the intensity of the excitation through the Lambda layer to enhance the gradient flow of the generator; S4. Adopt the encoder-decoder architecture to expand the discriminator to form a U-Net, propose a self-supervised reconstruction discriminator USRD based on U-Net, and design a unique local cropping method Shape-crop for T-shaped bolts, and enhance the feature extraction and detail texture generation capabilities of the discriminator through global and local reconstructions; Among them, S2 specifically includes: First, to generate more accurate bolt images, it is proposed to introduce perspective information based on bolt defect attributes into the GAN, and divide the bolt perspective into two categories with the largest visual differences: the front perspective 0° and the side perspective 90°, to ensure the feasibility and accuracy of the task; Next, a bolt perspective label generator based on CLIP fine-tuning is proposed and trained on Dataset 1 to learn the features of the two perspectives, and the CLIP is fine-tuned using the weight space integration method WiSE-FT, and the formula is as follows: wse(x,α)=f(x,(1-α)·θ0+α·θ1) (1) where x represents the input data, f is the classification prediction function of the model, θ0 represents the original model parameters, θ1 represents the model parameters obtained through standard fine-tuning, α is the mixing coefficient for controlling the weight, and the image encoder E of the fine-tuned CLIP is extracted as the bolt perspective label generator; Finally, the bolt defect dataset Dataset 2 with attribute labels is input into the bolt perspective label generator to assign perspective labels to each image, and by combining these perspective labels with the attribute labels of the bolts, a perspective-attribute bolt dataset Dataset 3 is constructed, and this process is defined as follows: Label(D3)=E(D2)+Attribute(D2) (2) Among them, E represents the bolt perspective label generator, Attribute is the attribute label of Dataset 2, and the categories of Dataset 3 include 6 categories: front - normal bolt, side - normal bolt, front - pin missing, side - pin missing, front - nut missing, and side - nut missing. The images in Dataset 2 and Dataset 3 are the same, only the category labels are different. The bolt perspective label generator reduces human bias and improves label consistency.
2. The generation method according to claim 1, wherein The architecture of the conditional generative adversarial network is adopted. Perspectives and attributes are introduced as additional conditional inputs in the generator, and the residual cross - layer excitation ResSLE is proposed to construct a generator with stronger gradient information flow and higher generation controllability. Specifically, it includes: First, an n - dimensional conditional vector is added to the input N - dimensional random noise vector. After concatenating the two, they are input into the generator. The input is defined as shown in the following formula: Z = concat(z, y) (3) Among them, Z is the input vector of the generator, z is the noise vector, and y is the conditional vector. Using the labels containing perspectives and attributes as conditional information to assist in generation increases the prior knowledge of the generator and improves the controllability of GAN generation. Another reason for introducing the condition is that different bolt defect images usually have certain commonalities. Although the attributes or perspectives of bolts are different, the main bodies of bolts are similar, which is conducive to the learning of the generator; In addition, the generator of the few - shot generation model FastGAN is used as the basic model, and the architecture of the generator is improved. The residual cross - layer excitation ResSLE is proposed. ResSLE introduces cross - layer connections on the basis of SLE to strengthen the gradient signal between layers and controls the intensity of excitation through the Lambda layer, effectively enhancing the gradient flow of the generator. The definition of ResSLE is as follows: ResSLE = SLE(x)+λ·U(x) (4) Among them, x represents the input feature, U() represents upsampling, and λ is used to control the excitation intensity. The improvement of the generator increases its own information content, aiming to alleviate the unequal confrontation between the generator and the discriminator, and helps to improve the generation performance index and training stability.
3. The generation method according to claim 1, wherein The discriminator is extended using the encoder - decoder architecture to form a U - Net. The self - supervised reconstruction discriminator USRD based on the U - Net is proposed, and a unique local cropping method Shape - crop is designed for T - type bolts. The feature extraction and detailed texture generation capabilities of the discriminator are enhanced through global and local reconstructions. Specifically, it includes: In the training of GAN, the information between the generator and the discriminator is asymmetric, which often leads to the discriminator relying too much on non-target features to judge true or false, affecting the generation effect and falling into a vicious cycle. To solve this problem, a self-supervised reconstruction discriminator USRD based on U-Net is proposed. Using an encoder-decoder architecture, USRD takes the feature extraction network of the discriminator as the encoder and reversely constructs the decoder with its structure to form a U-Net structure, including the encoder D enc , the global decoder D dec1 and the local decoder D dec2 ; First, for the underlying architecture encoder D enc , it is constructed using a traditional CNN network. The upper layer architecture of D enc consists of a traditional adversarial head and a conditional classification head. Specifically, the underlying architecture of the discriminator is used as a feature extractor to extract the features of the input image, which are sent to the adversarial head and the classification head respectively. Similar to a general GAN, the adversarial head uses a sigmoid activation function to judge the authenticity of the input image; the classification head uses a softmax activation function to classify the category of the input image and compare it with the input condition of the generator, expecting the two to belong to the same category; Secondly, the global decoder D dec1 takes the bottleneck layer output by the encoder D enc as input and has skip connections with the feature maps of the encoder to ensure that the generated image is effectively aligned with the original data distribution globally. This means that the generated image is similar in overall structure and layout. Through global reconstruction, the model can learn to capture the diversity and accuracy of the data distribution from a global perspective. The architecture of the local decoder D dec2 is similar to that of the global decoder D dec1 , except that the input of D dec2 is the feature cropped from the final output of the encoder D enc , rather than directly input from the bottleneck layer. It should be noted that the skip connections between the encoder D enc and the local decoder D dec2 both require corresponding cropping operations. The local decoder mainly focuses on the local regions of the generated image to ensure that these regions have high-quality textures and details. Another important reason for using local reconstruction is that the colors of the bolts and their background fittings are similar, and it is sometimes difficult to distinguish them well under lighting conditions. By reconstructing local pixels, the shape characteristics of the bolts can be highlighted, which is beneficial for the generator to accurately generate the shape of the bolts. The global and local reconstruction losses are defined as follows: Among them, I represents the input image, and L global Using the L2 norm, significant differences are optimized in global reconstruction to ensure that the overall image is close to the original image; L local Using the L1 norm, the error of each pixel is focused on in local reconstruction, emphasizing the accuracy of details and textures.
4. The generation method according to claim 3, wherein It also includes: A shape cropping method, Shape-crop, is proposed for the shape characteristics of T-bolts. The bolt image is cropped in a cross-centralized manner, and four squares are intercepted from the middle in the horizontal and vertical directions. At the same time, the central area containing the most bolt information is increased. The areas of the regions cropped by the two methods are the same, but the number of bolt pixels extracted by Shape-crop is significantly more than that of the random cropping method, so as to improve the reconstruction quality of D dec2 .
5. The generation method according to claim 1, wherein It also includes: Using the bolt defect controllable generation method to assist in improving the accuracy of the bolt defect recognition network and further verifying the effectiveness of the bolt defect controllable generation method. Specifically, it includes: To verify the effectiveness in few-shot bolt defect image generation, experiments were conducted on the classification task. The experimental process was divided into three steps: First, a large number of bolt defect images were generated using the proposed generation method. Then, the generated images were used to augment the dataset. Finally, the augmented dataset was used to train a bolt defect classification model, and the classification accuracy of the model was evaluated on the test set. The final results significantly improved the accuracy of the bolt defect classification model, demonstrating the effectiveness of the proposed generation method.
6. A bolt defect controllable generation device based on perspective and attribute guidance, characterized in that, The device includes: An S1 module for constructing a dataset Dataset 1 for training a bolt perspective label generator, including two perspectives of the front and side of normal bolts; constructing a few-shot bolt defect dataset Dataset 2, including bolt defect images of three categories, and the three categories include normal bolts, missing pins, and missing nuts; An S2 module for adopting a contrastive language-image pre-training CLIP model, proposing a bolt perspective label generator based on CLIP fine-tuning, automatically assigning perspective labels to the few-shot bolt defect dataset Dataset 2, and constructing a bolt defect dataset Dataset 3 containing perspective and attribute labels to provide data support for subsequent image generation tasks; An S3 module for adopting the architecture of a conditional generative adversarial network, introducing perspective and attributes as additional conditional inputs into the generator, and proposing a residual cross-layer excitation ResSLE to construct a generator with stronger gradient information flow and higher generation controllability; among them, ResSLE introduces cross-layer connections on the basis of the cross-layer excitation SLE to strengthen the gradient signal between layers, and controls the intensity of the excitation through the Lambda layer to enhance the gradient flow of the generator; An S4 module for expanding the discriminator using the encoder-decoder architecture to form a U-Net, proposing a self-supervised reconstruction discriminator USRD based on the U-Net, and designing a unique local cropping method Shape-crop for T-shaped bolts, enhancing the discriminator's feature extraction and detailed texture generation capabilities through global and local reconstructions; Among them, the S2 module specifically includes: First, to generate more accurate bolt images, it is proposed to introduce perspective information based on bolt defect attributes into the GAN, and divide the bolt perspective into two categories with the largest visual differences: the front perspective of 0° and the side perspective of 90°, to ensure the feasibility and accuracy of the task; Next, a bolt perspective label generator based on CLIP fine-tuning is proposed and trained on Dataset 1 to learn the features of the two perspectives. The CLIP is fine-tuned using the weight space integration method WiSE-FT, and the formula is as follows: wse(x,α)=f(x,(1-α)·θ0+α·θ1) (1) where x represents the input data, f is the classification prediction function of the model, θ0 represents the original model parameters, θ1 represents the model parameters obtained through standard fine-tuning, α is the mixing coefficient for controlling the weights, and the image encoder E of the fine-tuned CLIP is extracted as the bolt perspective label generator; Finally, the bolt defect dataset Dataset 2 with attribute labels is input into the bolt perspective label generator to assign perspective labels to each image. By combining these perspective labels with the attribute labels of the bolts, a perspective-attribute bolt dataset Dataset 3 is constructed. This process is defined as follows: Label(D3)=E(D2)+Attribute(D2) (2) Where E represents the bolt perspective label generator, Attribute is the attribute label of Dataset 2. The categories of Dataset 3 include 6 categories: front-normal bolt, side-normal bolt, front-pin missing, side-pin missing, front-nut missing, and side-nut missing. The images in Dataset 2 and Dataset 3 are the same, only the category labels are different. The bolt perspective label generator reduces human bias and improves label consistency.
7. An electronic device, characterized in that, The electronic device includes: at least one memory and at least one processor; the memory stores a program, and the processor calls the program stored in the memory, and the program is used to implement the perspective and attribute-guided controllable bolt defect generation method according to any one of claims 1-5.
8. A storage medium, characterized in that, The computer-executable instructions are stored in the storage medium, and the computer-executable instructions are used to execute the perspective and attribute-guided controllable bolt defect generation method according to any one of claims 1-5.
Citation Information
Patent Citations
Bolt missing detection method, device and equipment and storage medium
CN112419299A
Visually inseparable bolt defect detection method based on bolt attributes and positions
CN115311648A