Lightweight transformer defect detection method and device based on knowledge distillation, vehicle and electronic equipment
By constructing a lightweight transformer defect detection method based on knowledge distillation and using a CNN model to guide the optimization of transformer model parameters, the problem of low efficiency and poor accuracy in part defect detection is solved, and efficient and accurate part defect detection is achieved on edge devices.
Patent Information
- Application Number
- CN202511162348.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies suffer from low efficiency and poor accuracy in detecting component defects. In particular, lightweight networks cannot accurately detect component defects when edge devices have limited memory and insufficient defect data.
A lightweight transformer defect detection method based on knowledge distillation is constructed. A convolutional neural network (CNN) model is used as the teacher model, and the image feature information extracted by the CNN model is used to guide the transformer network model to optimize the student model parameters. The model is trained by combining prediction loss and distillation loss, which compresses the model and retains the advantages of local feature extraction of the CNN model.
It achieves efficient and accurate detection of part defects on edge devices, combining the local feature extraction capabilities of CNN models with the global feature extraction advantages of transformer models, thereby improving detection performance.
Smart Images

Figure CN121032976A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a lightweight transformer defect detection method and device based on knowledge distillation, a vehicle and an electronic device. BACKGROUND
[0002] In the manufacturing industry, the detection of part surface defects is an important link to ensure the quality of part products. The traditional manual inspection of part defects is low in efficiency and cannot meet the needs of large-scale production. The automatic visual detection technology realizes the detection of part defects through computer vision and image processing technology. However, due to the limited memory and low computing power of the edge deployment device, online calculation of large models is not supported, and only lightweight networks can be deployed. Moreover, due to the difficulty in obtaining defect data, the number of effective defect samples collected is small. Therefore, based on a small amount of defect data, the part defects cannot be accurately detected.
[0003] In related technologies, a knowledge distillation network is established by a teacher network and a student network, and the feature information of the intermediate network layer of the teacher model is used to make the student model repeatedly approach the feature information of the corresponding intermediate network layer of the teacher model. Then, the student model network is trained using the soft label of the teacher model. Finally, the sum of the difference losses between the output feature information of the intermediate network layer and the output feature information of the intermediate network layer of the teacher model is taken as the student model loss, and the sum of the student model loss and the distillation loss multiplied by two weight parameters is taken as the total loss. Although this technology can compress the model, directly summing the losses of high-level network features and low-level network features as the total loss will introduce a large amount of noise in the bottom features into the student network, slow down the convergence of the student model, and require a large number of defects for learning. Therefore, the detection efficiency is low and the detection accuracy is poor when detecting part defects. SUMMARY
[0004] The present application provides a lightweight transformer defect detection method and device based on knowledge distillation, a vehicle and an electronic device to at least solve the technical problems of low detection efficiency and poor detection accuracy when detecting part defects in related technologies. The technical solutions of the present application are as follows:
[0005] According to a first aspect of the present application, a lightweight transformer defect detection method based on knowledge distillation is provided, the method comprising: constructing a defect detection dataset, wherein the defect detection dataset comprises defect images of industrial parts; the defect images are labeled with actual defect information; training a student defect detection model based on the defect detection dataset and a teacher defect detection model; wherein the teacher defect detection model is a convolutional neural network (CNN) model for detecting defects of industrial parts; the student defect detection model is a transformer network model; the student defect detection model is optimized by using the feature difference between the first feature information and the second feature information; the first feature information is the image features of the defect images extracted by the teacher defect detection model, and the second feature information is the image features of the defect images extracted by the student defect detection model.
[0006] According to the above technical means, the defect images of the industrial parts labeled with actual defect information are used as the defect detection dataset, and then the convolutional neural network (CNN) model is used as the teacher defect detection model, which can accurately extract the image features of the defect images in the defect detection dataset. Therefore, based on the image features extracted by the teacher defect detection model, the image features extracted from the defect images in the defect detection dataset by the lightweight transformer network model as the student defect detection model are judged. Furthermore, based on the feature difference between the image features extracted by the two models, the model parameters of the student defect detection model are optimized to obtain the student defect detection model after optimization training. In this way, the knowledge distillation compression model is used, and the advantages of the convolutional neural network (CNN) model in extracting local details, edges, textures and other low-level features are introduced into the transformer network model, so that the lightweight transformer network model not only retains the advantage of extracting global features, but also has the advantage of the convolutional neural network (CNN) model in extracting global features. The lightweight transformer network model can achieve high detection performance based on sample data. Therefore, based on the student defect detection model after optimization training, the defects of the industrial parts can be detected, and the defects of the industrial parts can be detected efficiently and accurately.
[0007] In a possible implementation, the training of the student defect detection model based on the defect detection dataset and the teacher defect detection model comprises: determining a distillation loss between the first feature information and the second feature information based on the feature difference; inputting the defect images into the student defect detection model to obtain predicted defect information of the defect images; determining a prediction loss based on the deviation between the predicted defect information and the actual defect information; and optimizing the student defect detection model based on the distillation loss and the prediction loss.
[0008] According to the technical means, the student defect detection model is used to predict the defect information of the defect image, and then the predicted defect information is combined with the actual defect information to determine the deviation between the two as a prediction loss. The determined first feature information and second feature information are used to determine a distillation loss, and then the distillation loss and the prediction loss are used to optimize the student defect detection model to obtain an optimized student defect detection model. In this way, the prediction loss and the distillation loss are combined to consider various losses of the model for training the model, and the model is efficiently and accurately optimized.
[0009] In a possible implementation, the first feature extraction layer of the convolutional neural network CNN model includes a plurality of first feature extraction sub-layers, and the second feature extraction layer of the transformer network model includes a plurality of second feature extraction sub-layers; the first feature information is feature information corresponding to part of the first feature extraction sub-layers; and the second feature information is feature information corresponding to part of the second feature extraction sub-layers.
[0010] According to the technical means, the feature extraction layer of the convolutional neural network CNN model includes feature information corresponding to part of the feature extraction sub-layers, and the feature extraction layer of the transformer network model includes part of the feature extraction sub-layers. In this way, the student defect detection model can be optimized and trained based on the combination of the feature information extracted by part of the feature extraction sub-layers of the convolutional neural network CNN model and the feature information extracted by part of the feature extraction sub-layers of the transformer network model, without considering the feature information extracted by all the feature extraction sub-layers of the convolutional neural network CNN model and the transformer network model, so that the optimized student defect detection model can be obtained through part of the feature information.
[0011] In a possible implementation, the first feature extraction layer of the convolutional neural network (CNN) model includes a plurality of first feature extraction sub-layers, and the second feature extraction layer of the transformer network model includes a plurality of second feature extraction sub-layers; the first feature information includes first global feature information corresponding to the first feature extraction layer and first local feature information corresponding to part of the first feature extraction sub-layers in the plurality of first feature extraction sub-layers; the second feature information includes second global feature information corresponding to the second feature extraction layer and second local feature information corresponding to part of the second feature extraction sub-layers in the plurality of second feature extraction sub-layers; and the determining the distillation loss between the first feature information and the second feature information based on the feature difference includes: determining a global feature loss based on a feature difference between the first global feature information and the second global feature information; determining a local feature loss based on a feature difference between the first local feature information and the second local feature information; and determining the distillation loss based on the global feature loss and the local feature loss.
[0012] According to the technical means, the global feature loss can be determined based on the global feature information extracted by all the feature extraction sub-layers of the convolutional neural network (CNN) model and the global feature information extracted by all the feature extraction sub-layers of the transformer network model. In addition, the local feature loss can be determined based on the local feature information extracted by part of the feature extraction sub-layers of the convolutional neural network (CNN) model and the local feature information extracted by part of the feature extraction sub-layers of the transformer network model. Thus, the distillation loss can be accurately determined based on the global feature loss and the local feature loss. Furthermore, the model can be trained more efficiently and accurately based on the distillation loss and the prediction loss, and an optimized model can be obtained.
[0013] In a possible implementation, the prediction defect information includes a prediction probability of a defect type being a preset defect type, a confidence of the prediction probability, and a predicted defect position; the actual defect information includes an actual defect type and an actual defect position; and the determining the prediction loss based on a deviation between the prediction defect information and the actual defect information includes: determining a classification prediction loss based on the prediction probability and the actual defect type; determining a position prediction loss based on the predicted defect position and the actual defect position; determining a confidence loss based on the confidence of the prediction probability; and determining the prediction loss based on the classification prediction loss, the position prediction loss, and the confidence loss.
[0014] Based on the aforementioned technical means, this application can comprehensively consider the predicted probability of a defect type being a preset defect type, the confidence level of the predicted probability, and the predicted defect location, combined with the actual defect type and actual defect location, to further determine the classification prediction loss, location prediction loss, and confidence loss. By further summing and considering the classification prediction loss, location prediction loss, and confidence loss, the prediction loss can be accurately determined. Furthermore, by combining the prediction loss with the distillation loss, the model's loss can be considered more comprehensively, allowing for efficient and accurate model training to obtain an optimized model.
[0015] According to the second aspect of this application, a lightweight transformer defect detection method based on knowledge distillation is provided. The method includes: acquiring a defect image to be detected, wherein the defect image to be detected is an image of an industrial part to be detected; inputting the defect image to be detected into a student defect detection model to determine the defect type and defect location of the industrial part to be detected.
[0016] According to a third aspect provided in this application, a lightweight transformer-based defect detection device based on knowledge distillation is provided. The device includes: a construction module and a training module; the construction module is used to construct a defect detection dataset, wherein the defect detection dataset includes defect images of industrial parts; the defect images are labeled with actual defect information; the training module is used to train a student defect detection model based on the defect detection dataset and a teacher defect detection model; wherein the teacher defect detection model is a convolutional neural network (CNN) model for detecting defects in industrial parts; the student defect detection model is a transformer network model; the student defect detection model optimizes model parameters using the feature difference between a first feature information and a second feature information; the first feature information is the image features of the defect images extracted by the teacher defect detection model, and the second feature information is the image features of the defect images extracted by the student defect detection model.
[0017] In one possible implementation, the training module is specifically used to determine the distillation loss between the first feature information and the second feature information based on the feature difference; the training module is specifically used to input the defect image into the student defect detection model to obtain the predicted defect information of the defect image; the training module is specifically used to determine the prediction loss based on the deviation between the predicted defect information and the actual defect information; the training module is specifically used to optimize the student defect detection model based on the distillation loss and the prediction loss.
[0018] In one possible implementation, the first feature extraction layer of the convolutional neural network (CNN) model includes multiple first feature extraction sub-layers, and the second feature extraction layer of the transformer network model includes multiple second feature extraction sub-layers; the first feature information is the feature information corresponding to a portion of the multiple first feature extraction sub-layers; and the second feature information is the feature information corresponding to a portion of the multiple second feature extraction sub-layers.
[0019] In one possible implementation, the first feature extraction layer of the convolutional neural network (CNN) model includes multiple first feature extraction sub-layers, and the second feature extraction layer of the transformer network model includes multiple second feature extraction sub-layers. The first feature information includes first global feature information corresponding to the first feature extraction layer and first local feature information corresponding to a portion of the multiple first feature extraction sub-layers. The second feature information includes second global feature information corresponding to the second feature extraction layer and second local feature information corresponding to a portion of the multiple second feature extraction sub-layers. A training module is specifically used to determine a global feature loss based on the feature difference between the first and second global feature information. A training module is specifically used to determine a local feature loss based on the feature difference between the first and second local feature information. A training module is specifically used to determine a distillation loss based on the global feature loss and the local feature loss.
[0020] In one possible implementation, the predicted defect information includes: a predicted probability that the defect type is a preset defect type, a confidence level of the predicted probability, and a predicted defect location; the actual defect information includes: the actual defect type and the actual defect location; the training module is specifically used to determine a classification prediction loss based on the predicted probability and the actual defect type; the training module is specifically used to determine a location prediction loss based on the predicted defect location and the actual defect location; the training module is specifically used to determine a confidence loss based on the confidence level of the predicted probability; and the training module is specifically used to determine a prediction loss based on the classification prediction loss, the location prediction loss, and the confidence loss.
[0021] According to the fourth aspect provided in this application, a lightweight transformer defect detection device based on knowledge distillation is provided. The lightweight transformer defect detection device based on knowledge distillation includes: an acquisition module and a detection module; the acquisition module is used to acquire an image of a defect to be detected, wherein the image of the defect to be detected is an image of an industrial part to be detected; the detection module is used to input the image of the defect to be detected into a student defect detection model to detect the defect type and defect location of the industrial part to be detected.
[0022] According to a fifth aspect provided in this application, an electronic device is provided, comprising: a processor and a memory; wherein the memory is used to store one or more programs, the one or more programs including computer-executable instructions, and when the electronic device is running, the processor executes the computer-executable instructions stored in the memory, and the electronic device executes the method of the first aspect described above and any possible implementation thereof, and / or the method of the implementation of the second aspect described above.
[0023] According to the sixth aspect provided in this application, a computer-readable storage medium is provided, wherein when computer instructions stored in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device performs the method of the first aspect and any possible implementation thereof, and / or the method of the implementation of the second aspect.
[0024] According to the seventh aspect provided in this application, a computer program product is provided, the computer program product including computer instructions, which, when executed on an electronic device, cause the electronic device to perform the method of the first aspect described above and any possible implementation thereof, and / or the method of the implementation of the second aspect described above.
[0025] According to the eighth aspect provided in this application, a vehicle is provided, the vehicle including industrial parts, the industrial parts being defect-detected using the method of the second aspect described above.
[0026] It should be noted that the technical effects of any of the implementation methods in aspects two through eight can be found in the technical effects of the corresponding implementation methods in aspect one, and will not be repeated here.
[0027] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0028] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application, and do not constitute an undue limitation of this application.
[0029] Figure 1 This is a schematic diagram of the structure of a defect detection system according to an exemplary embodiment;
[0030] Figure 2 This is a flowchart illustrating a lightweight transformer defect detection method based on knowledge distillation, according to an exemplary embodiment. Figure 1 ;
[0031] Figure 3 This is a schematic diagram illustrating the framework structure of a teacher defect detection model according to an exemplary embodiment;
[0032] Figure 4 This is a schematic diagram illustrating the framework structure of a teacher-student visual detection network model according to an exemplary embodiment;
[0033] Figure 5 This is a flowchart illustrating a lightweight transformer defect detection method based on knowledge distillation, according to an exemplary embodiment. Figure 2 ;
[0034] Figure 6 This is a flowchart illustrating a lightweight transformer defect detection method based on knowledge distillation, according to an exemplary embodiment. Figure 3 ;
[0035] Figure 7 This is a flowchart illustrating a lightweight transformer defect detection method based on knowledge distillation, according to an exemplary embodiment. Figure 4 ;
[0036] Figure 8 This is a flowchart illustrating a lightweight transformer defect detection method based on knowledge distillation, according to an exemplary embodiment.
[0037] Figure 9 This is a schematic diagram illustrating the framework structure of a student defect detection model according to an exemplary embodiment;
[0038] Figure 10 This is a block diagram illustrating a lightweight transformer defect detection device based on knowledge distillation, according to an exemplary embodiment.
[0039] Figure 11 This is a block diagram illustrating a lightweight transformer defect detection device based on knowledge distillation, according to an exemplary embodiment.
[0040] Figure 12 This is a block diagram illustrating an electronic device according to an exemplary embodiment;
[0041] Figure 13 This is a schematic diagram of the structure of a computer system according to an exemplary embodiment. Detailed Implementation
[0042] To enable those skilled in the art to better understand the technical solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0043] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0044] The lightweight transformer defect detection method based on knowledge distillation provided in this application can be applied to defect detection systems. Figure 1 A schematic diagram of a defect detection system is shown. Figure 1 As shown, the defect detection system includes an electronic device 11 and an industrial part 12. The electronic device 11 is used to perform defect detection on the industrial part 12 to determine the defect type and location of the industrial part 12. The industrial part 12 includes the industrial part corresponding to the defect image used to construct the defect detection dataset, and the industrial part to be detected.
[0045] Electronic device 11 is used to construct a defect detection dataset and, based on the defect detection dataset and the teacher defect detection model, to train a student defect detection model.
[0046] The electronic device 11 is also used to input the image of the defect to be detected into the trained student defect detection model to determine the defect type and defect location of the industrial part to be detected.
[0047] In one possible implementation, the defect detection system may specifically include: a dataset construction module, a model training module, and a defect detection module.
[0048] The dataset construction module is used to build a defect detection dataset, divide the defect detection dataset into training and testing sets, and annotate the defect images of industrial parts included in the training and testing sets with actual defect information.
[0049] The dataset building module is also used to preprocess defect images of industrial parts included in the training and test sets.
[0050] The model training module is used to train the teacher defect detection model and the student defect detection model using the pre-processed training set. During the training process, the pre-trained teacher defect detection model is used to guide the training of the student defect detection model.
[0051] The defect detection module is used to test the trained student defect detection model using a preprocessed test set.
[0052] For ease of understanding, the lightweight transformer defect detection method based on knowledge distillation provided in this application will be described in detail below with reference to the accompanying drawings.
[0053] Figure 2 This is a flowchart illustrating a lightweight transformer defect detection method based on knowledge distillation, according to an exemplary embodiment, applied to electronic devices, such as... Figure 2 As shown, this lightweight transformer defect detection method based on knowledge distillation includes the following steps S201-S202:
[0054] S201. Construct a defect detection dataset.
[0055] The defect detection dataset includes defect images of industrial parts; the defect images are labeled with actual defect information.
[0056] In one possible implementation, defect images of multiple industrial parts can be acquired, and actual defect information can be labeled in each defect image. This actual defect information may include: defect type, defect location, and defect size, etc.
[0057] In one possible implementation, the defect detection dataset can be divided according to a preset ratio (e.g., 4:1, 3:2) to obtain a training set and a validation set, thereby training and validating the teacher defect detection model and the student defect detection model based on the training set and the validation set.
[0058] In one possible implementation, the training and validation sets can be preprocessed, specifically including resizing and data augmentation. Resizing transforms the input defect image to a preset size, while data augmentation enhances the data information in the defect image using methods such as random horizontal flipping, random vertical flipping, and random rotation.
[0059] In one possible implementation, a large convolutional neural network (CNN) model with pre-trained weights can be trained based on a training set and a validation set, thereby causing the CNN model to converge to obtain a pre-trained CNN model.
[0060] In one possible implementation, a teacher-student visual detection network model can be further constructed by using a converged convolutional neural network (CNN) model and a lightweight transformer network model. The CNN model is used as the teacher network (teacher model, i.e., teacher defect detection model), and the transformer network model is used as the student network (student model, i.e., student defect detection model).
[0061] In some embodiments, the convolutional neural network (CNN) model (teacher defect detection model) also needs to be trained based on the defect detection dataset (i.e., training set and validation set) so that the CNN model converges.
[0062] Specifically, the training and validation sets are input into the teacher defect detection model for feature extraction. The features extracted in the final stage of the model are then input into the detector to obtain the predicted defect information for defective images (i.e., defective images included in the training set). The predicted defect information is then compared with the actual defect information labeled in the defective images to determine the prediction loss (also known as the label loss). The teacher defect detection model is then updated through backpropagation based on the prediction loss to improve its performance until it converges.
[0063] In this embodiment, the teacher defect detection model can be constructed using a Residual Network (ResNet-101) model, which consists of five stages. The first stage includes a 7x7 convolutional layer and a 3x3 max-pooling layer, with an output feature size of 56×56×64. The second stage contains three residual blocks, with an output feature size of 56×56×256. The third stage contains four residual blocks, with an output feature size of 14×14×512. The fourth stage contains 23 residual blocks, with an output feature size of 14×14×1024. The fifth stage contains five residual blocks, with an output feature size of 7×7×2048. The student model can be a Transformer network model, which has four stages, each consisting of multiple stacked Transformer encoders. The four stages generate four one-dimensional feature maps of different scales, with feature sizes of 3136×64, 784×128, 196×256 and 49×512, respectively.
[0064] For example, such as Figure 3The diagram shows the framework of the teacher defect detection model. By sequentially inputting the defect detection dataset (i.e., the training set and the validation set) into the five stages (i.e., the first stage, the second stage, the third stage, the fourth stage, and the fifth stage) of the teacher defect detection model, feature extraction is performed to extract feature information (i.e., the first feature information). Then, the extracted feature information is input into the detector to obtain the predicted defect information of the defect image. Based on the predicted defect information and the actual defect information, the prediction loss (also known as the label loss) can be determined.
[0065] S202. Based on the defect detection dataset and the teacher defect detection model, train the student defect detection model.
[0066] Among them, the teacher defect detection model is a convolutional neural network (CNN) model used to detect defects in industrial parts; the student defect detection model is a transformer network model; the student defect detection model optimizes the model parameters by utilizing the feature differences between the first feature information and the second feature information; the first feature information is the image features of the defect image extracted by the teacher defect detection model, and the second feature information is the image features of the defect image extracted by the student defect detection model.
[0067] In one possible implementation, after training the convolutional neural network (CNN) model (teacher defect detection model) to obtain a converged teacher defect detection model, it is also necessary to input the defect detection dataset (i.e., the training set and validation set) into the teacher-student visual detection network model for training, so as to train the student defect detection model and make the student defect detection model converge.
[0068] Specifically, the training and test sets can be simultaneously input into both the teacher defect detection model and the student defect detection model for feature extraction. This allows for the extraction of accurate features (i.e., first feature information) based on the teacher defect detection model. Furthermore, based on the four stages of the student defect detection model, features at four different scales (i.e., second feature information) are extracted from shallow to deep levels.
[0069] In this embodiment of the application, since the feature information extracted by the convolutional neural network (CNN) model is two-dimensional feature information and the feature information extracted by the transformer network model is one-dimensional feature information, in order to perform subsequent calculations between the feature information extracted by the transformer network model and the CNN model, a feature conversion module is needed to convert the one-dimensional feature information (i.e., the first feature information) extracted by the transformer network model into two-dimensional feature information of the same size as that extracted by the CNN model.
[0070] The conversion formula used in the feature conversion module is Formula 1.
[0071] S i =1×1Conv(Reshape(s) i Formula 1: i∈(1,2,3,4)
[0072] Among them, s i S is the one-dimensional feature information extracted from the four stages of the transformer network model. i Two-dimensional feature information is obtained by transforming one-dimensional feature information.
[0073] In one possible implementation, the feature information extracted from the last of the four stages of the transformer network model needs to be input into the detector to obtain the predicted defect information of the defect image by the transformer network model. Then, the feature difference between the predicted defect information and the actual defect information of the defect image is calculated to obtain the prediction loss (also known as the label loss).
[0074] In one possible implementation, the feature information extracted from the second to fourth stages of the five stages of the Convolutional Neural Network (CNN) model, which has strong feature extraction capabilities, can be used as labels to calculate the feature information extracted from the first three stages of the four stages of the Transformer network model, so as to constrain the feature information extracted from the first three stages of the Transformer network model.
[0075] In one possible implementation, the final feature information extracted by the convolutional neural network (CNN) model can be used as a label to calculate the distillation loss, thereby constraining the transformer network model and guiding it to extract more accurate feature information.
[0076] Thus, based on the determined prediction loss and distillation loss (including feature loss), the total loss of the teacher-student visual detection network model can be calculated, and then the teacher-student visual detection network model is updated through backpropagation until the student model converges.
[0077] For example, such as Figure 4The diagram illustrates the framework of a teacher-student visual detection network model. By inputting the defect detection datasets (i.e., training and validation sets) into the five stages (stage 1, stage 2, stage 3, stage 4, and stage 5) of the teacher defect detection model, feature information T1 (generated by stage 2), feature information T2 (generated by stage 3), feature information T3 (generated by stage 4), and feature information T4 (generated by stage 5) are obtained from the latter four stages of the teacher defect detection model. Similarly, by inputting the defect detection datasets (i.e., training and validation sets) into the four stages (i.e., stage 1, stage 2, stage 3, and stage 4) of the student defect detection model, feature information s1 (generated by stage 1), feature information s2 (generated by stage 2), feature information s3 (generated by stage 3), and feature information s4 (generated by stage 4) are obtained from the four stages of the student defect detection model. Furthermore, the one-dimensional feature information s1, s2, s3, and s4 need to be converted into two-dimensional feature information S1, S2, S3, and S4 through a feature transformation module. Then, feature difference 1 can be obtained based on feature information T1 and S1, feature difference 2 can be obtained based on feature information T2 and S2, feature difference 3 can be obtained based on feature information T3 and S3, and feature difference 4 (i.e., distillation loss) can be obtained based on feature information T4 and S4. The feature information s1, s2, s3, and s4 extracted by the student defect detection model are then input into the detector to obtain the predicted defect information for the defect image. Based on the predicted defect information and the actual defect information, the prediction loss (also known as label loss) can be determined.
[0078] In this embodiment, defect images of industrial parts labeled with actual defect information are used as the defect detection dataset. A convolutional neural network (CNN) model is then used as the teacher defect detection model to accurately extract image features from the defect images in the dataset. Based on the image features extracted by the teacher defect detection model, the image features extracted by a lightweight transformer network model (serving as the student defect detection model) from the defect images in the dataset are evaluated. Furthermore, based on the feature differences between the two models, the model parameters of the student defect detection model are optimized, resulting in an optimized student defect detection model. Thus, by using a knowledge distillation compression model and incorporating the advantages of CNN models in extracting low-level features such as local details, edges, and textures into the transformer network model, the lightweight transformer network model retains the advantage of extracting global features while also possessing the advantages of CNN models in this area. This allows the lightweight transformer network model to achieve high detection performance based on sample data. Therefore, using the optimized student defect detection model to detect defects in industrial parts can efficiently and accurately detect defects in industrial parts.
[0079] In some embodiments, such as Figure 5 As shown in the embodiment of this application, in a lightweight transformer defect detection method based on knowledge distillation, the above step S202 may specifically include S501-S504.
[0080] S501. Determine the distillation loss between the first feature information and the second feature information based on feature differences.
[0081] In some embodiments, the first feature extraction layer of the convolutional neural network (CNN) model includes multiple first feature extraction sub-layers, and the second feature extraction layer of the transformer network model includes multiple second feature extraction sub-layers; the first feature information is the feature information corresponding to a portion of the multiple first feature extraction sub-layers; and the second feature information is the feature information corresponding to a portion of the multiple second feature extraction sub-layers.
[0082] In this embodiment, the feature extraction layer of a Convolutional Neural Network (CNN) model can use the feature information corresponding to a portion of its multiple feature extraction sub-layers as the first feature information, and the feature extraction layer of a Transformer network model can use a portion of its multiple feature extraction sub-layers as the second feature information. Thus, the student defect detection model can be optimized and trained by combining the feature information extracted from the partial feature extraction sub-layers of the CNN and Transformer networks, without needing to comprehensively consider the feature information extracted from all feature extraction sub-layers of both models. The optimized student defect detection model can be obtained by training with only partial feature information.
[0083] In some embodiments, the first feature extraction layer of the convolutional neural network (CNN) model includes multiple first feature extraction sub-layers, and the second feature extraction layer of the transformer network model includes multiple second feature extraction sub-layers; the first feature information includes first global feature information corresponding to the first feature extraction layer and first local feature information corresponding to some of the multiple first feature extraction sub-layers; the second feature information includes second global feature information corresponding to the second feature extraction layer and second local feature information corresponding to some of the multiple second feature extraction sub-layers.
[0084] In some embodiments, such as Figure 6 As shown in the embodiment of this application, in a lightweight transformer defect detection method based on knowledge distillation, the above step S501 may specifically include S601-S603.
[0085] S601. Determine the global feature loss based on the feature difference between the first global feature information and the second global feature information.
[0086] In one possible implementation, to capture complex distribution information and improve the robustness and generalization ability of the transformer network model, a global feature loss is used. dis The maximum mean difference function can be used for calculation, as shown in Formula 2.
[0087]
[0088] Among them, F s ' represents the flattened matrix of the feature information (i.e., the second feature information) extracted by the transformer network model, F t' represents the flattened matrix of the feature information (i.e., the first feature information) extracted by the Convolutional Neural Network (CNN) model. The dimensions of the flattened matrix are (w×h)×c, where w represents the width, h represents the height, and c represents the number of channels.
[0089] Specifically, F s It contains m = n = w × h eigenvectors, each with a size of c. Similarly, F t It also contains m = n = w × h eigenvectors, each of which has a size of c. This represents the i-th feature vector after the feature information extracted by the transformer network model has been flattened. This represents the i-th feature vector after the feature information extracted by the Convolutional Neural Network (CNN) model has been flattened. The kernel function (e.g., the Gaussian kernel function, also known as the radial basis function (RBF)) is used to calculate the similarity between two feature vectors, σ. 2 This represents the bandwidth of the Gaussian kernel.
[0090] S602. Determine the local feature loss based on the feature difference between the first local feature information and the second local feature information.
[0091] In one possible implementation, to improve the learning efficiency of the transformer network model by enabling it to quickly learn defect features from a small number of samples (i.e., defect images), feature information extracted by a convolutional neural network (CNN) model can be used to supervise the training of the transformer network model.
[0092] Specifically, the feature information extracted from the second to fourth stages of the Convolutional Neural Network (CNN) model can be used as labels to constrain the feature extraction layers of the first three stages of the Transformer network model, and feature loss (i.e., local feature loss) can be calculated. To avoid noise in the feature information extracted by the CNN model misleading the Transformer network model about learning defective features, the local feature loss is calculated accordingly. f The smooth L1 loss can be used for calculation, and its calculation formula is shown in Formula 3.
[0093]
[0094] in, This represents the difference between predicted defect information and actual defect information. This represents the element in the i-th row and j-th column of the feature information extracted by the Convolutional Neural Network (CNN) model. This represents the element in the i-th row and j-th column of the feature information extracted by the transformer network model. w and h represent the width and height of the feature information, respectively.
[0095] In one possible implementation, to reduce the interference of noise contained in the low-level feature information extracted by the Convolutional Neural Network (CNN) model on the learning of the Transformer network model, and to enhance the supervision of the Transformer network model by the high-level feature information extracted by the CNN model, the calculated feature losses for the three stages (i.e., the second to fourth stages of the CNN model and the first three stages of the Transformer network model) can be multiplied by different weight coefficients to determine the local feature loss. f As shown in Formula 4.
[0096]
[0097] in, The feature loss is calculated using the feature information extracted in the second stage of the Convolutional Neural Network (CNN) model and the feature information extracted in the first stage of the Transformer network model. The feature loss is calculated using the feature information extracted in the third stage of the Convolutional Neural Network (CNN) model and the feature information extracted in the second stage of the Transformer Network model. This represents the feature loss calculated using the feature information extracted in the fourth stage of the Convolutional Neural Network (CNN) model and the feature information extracted in the third stage of the Transformer network model. α, β, and γ represent hyperparameters (i.e., variable coefficients). For example, to reduce the impact of noise in shallow features on the Transformer network model, the initial values of the hyperparameters can be set to 0.4, 0.6, and 0.8.
[0098] S603. Determine the distillation loss based on global feature loss and local feature loss.
[0099] In this embodiment, the global feature loss is determined by combining the global feature information extracted by all feature extraction sub-layers of the Convolutional Neural Network (CNN) model with the global feature information extracted by all feature extraction sub-layers of the Transformer network model. Furthermore, the local feature loss is determined by combining the local feature information extracted by some feature extraction sub-layers of the CNN model with the local feature information extracted by some feature extraction sub-layers of the Transformer network model. Therefore, based on the global and local feature losses, the distillation loss can be accurately determined. Furthermore, by combining the distillation loss with the prediction loss, the model's loss can be considered more comprehensively, allowing for efficient and accurate model training to obtain an optimized model.
[0100] S502. Input the defect image into the student defect detection model to obtain the predicted defect information of the defect image.
[0101] S503. Determine the predicted loss based on the deviation between the predicted defect information and the actual defect information.
[0102] S504. Optimize the student defect detection model based on distillation loss and prediction loss.
[0103] In this embodiment, the student defect detection model predicts defect information from a defect image. Then, the predicted defect information is combined with the actual defect information, and the deviation between the two is determined as the prediction loss. Additionally, a distillation loss is calculated between the determined first and second feature information. Based on the distillation loss and the prediction loss, the student defect detection model is optimized, resulting in an optimized model. Thus, by combining the prediction loss and the distillation loss, this application can comprehensively consider multiple losses in the model during training, thereby efficiently and accurately optimizing the model.
[0104] In some embodiments, the predicted defect information includes: the predicted probability that the defect type is a preset defect type, the confidence level of the predicted probability, and the predicted defect location; the actual defect information includes: the actual defect type and the actual defect location.
[0105] In some embodiments, such as Figure 7 As shown in the embodiment of this application, in a lightweight transformer defect detection method based on knowledge distillation, the above step S503 may specifically include S701-S704.
[0106] S701. Determine the classification prediction loss based on the predicted probability and the actual defect type.
[0107] In one possible implementation, to improve the feature extraction capabilities of Convolutional Neural Network (CNN) and Transformer network models, prediction loss can be used to supervise the predictions of these models. label It consists of three parts: classification prediction loss, location prediction loss, and confidence loss, as shown in Formula 5.
[0108] loss label =loss cls +loss conf +loss reg Formula 5
[0109] Where, loss cls This represents the classification prediction loss. conf This represents the confidence loss. reg This indicates the location prediction loss.
[0110] Specifically, the classification prediction loss can be calculated using the multi-class cross-entropy loss function, as shown in Formula 6.
[0111]
[0112] Among them, y ia p represents the true probability that the i-th predicted bounding box (i.e., the i-th defect in the predicted defect image) belongs to the a-th defect category. ia This represents the prediction probability of the i-th prediction box for the a-th defect category. N is the total number of prediction boxes in the defect image (i.e., the number of defects predicted for the defect image), and a is the total number of defect categories.
[0113] That is, actual defects are represented by ground truth bounding boxes in the defect image, and predicted defects are represented by predicted bounding boxes in the defect image.
[0114] S702. Determine the location prediction loss based on the predicted defect location and the actual defect location.
[0115] Specifically, the location prediction loss (also known as regression loss) can be calculated using the GIoU loss function, as shown in Formula 7.
[0116]
[0117] Where I represents the area where the ground truth bounding box (i.e., the real defect) and the predicted bounding box (i.e., the predicted defect) overlap in the defect image, and A p A represents the area of the prediction box (i.e., the predicted defect). g Represents the area of the true bounding box (i.e., the actual defect). A cThis represents the area of the smallest bounding rectangle between the ground truth bounding box and the predicted bounding box.
[0118] S703. Determine the confidence loss based on the confidence level of the predicted probability.
[0119] Specifically, the confidence loss can be calculated using the binary cross-entropy loss function, as shown in Formula 8.
[0120]
[0121] Among them, y i p represents the actual confidence level of the i-th prediction box. i This represents the prediction confidence of the i-th prediction box.
[0122] S704. Determine the prediction loss based on classification prediction loss, location prediction loss, and confidence loss.
[0123] Thus, based on Formula 5 above, the prediction loss can be determined based on the classification prediction loss, confidence loss, and location prediction loss.
[0124] In this embodiment, the application comprehensively considers the predicted probability of a defect type being a preset defect type, the confidence level of the predicted probability, and the predicted defect location, combined with the actual defect type and actual defect location, to further determine the classification prediction loss, location prediction loss, and confidence loss. This is further achieved by summing the classification prediction loss, location prediction loss, and confidence loss to accurately determine the prediction loss. Furthermore, by combining the prediction loss with the distillation loss, the model's loss can be considered more comprehensively, allowing for efficient and accurate model training to obtain an optimized model.
[0125] In one possible implementation, the total loss can be determined based on the prediction loss and the distillation loss (including global feature loss and local feature loss). total As shown in Formula Nine.
[0126] loss total =ε×loss label +θ×loss dis +μ×loss f Formula Nine
[0127] Where ε, θ, and μ are variable hyperparameters (i.e., variable coefficients, also called weight coefficients), their initial values can be set to 0.6, 1, and 0.2, respectively. To enable the Transformer network model to learn the advantages of Convolutional Neural Networks (CNNs), such as inductive bias and translation invariance, while retaining its powerful representational capabilities, the weight coefficients of the distillation loss (including global and local feature losses) can be gradually decreased with each training iteration using cosine annealing. This gradually reduces the constraints on the CNN model, while the weight coefficients of the prediction loss gradually increase.
[0128] Figure 8 This is a flowchart illustrating a lightweight transformer defect detection method based on knowledge distillation, according to an exemplary embodiment, applied to electronic devices, such as... Figure 8 As shown, this lightweight transformer defect detection method based on knowledge distillation includes the following S801-S802:
[0129] S801. Obtain the image of the defect to be detected.
[0130] Among them, the defect image to be detected is an image of the industrial part to be inspected.
[0131] S802. Input the image of the defect to be detected into the student defect detection model to determine the defect type and location of the industrial part to be detected.
[0132] In one possible implementation, the defect image of the industrial part to be detected can be input into a trained student defect detection model (transformer network model) to obtain the detection results (i.e., defect type and defect location).
[0133] For example, such as Figure 9 As shown, the image of the defect to be detected can be input into the trained student defect detection model. The student defect detection model then detects the image of the defect through four stages (i.e., the first stage, the second stage, the third stage, and the fourth stage), extracts the feature information of the image of the defect to be detected, and inputs the extracted feature information into the detector to obtain the detection result (defect type and defect location) of the industrial part to be detected corresponding to the image of the defect to be detected.
[0134] In this embodiment, a large convolutional neural network (CNN) model with powerful feature extraction capabilities is used as the teacher model (i.e., the teacher defect detection model), and a lightweight Transformer network model with powerful global context modeling capabilities is used as the student model (i.e., the student defect detection model). Features extracted by the teacher model at multiple stages are then used to provide more supervision information to the student model. This allows the student model to not only capture high-level features of defects and solve the information loss problem using global context modeling capabilities, but also to gain the advantages of the CNN model in extracting low-level features such as local details, edges, and textures of defects. This enables the student model to quickly learn defect features from a small amount of data, reducing dependence on large amounts of defect data and improving the detection performance of the lightweight Transformer network model for surface defects on industrial parts.
[0135] Furthermore, this application utilizes the knowledge learned by a large-scale network to guide the training of a lightweight network. This allows the lightweight network to achieve performance comparable to that of a large-scale network while significantly reducing the number of parameters and computational costs. This enables model compression and acceleration, reducing unnecessary computational and storage requirements during visual inspection, greatly improving the efficiency of industrial visual inspection and reducing project development costs. Consequently, it can be used for defect detection on the surface of industrial parts, such as dents, cracks, and scratches in stamped parts. This application effectively utilizes a small amount of defect data to improve the performance of a lightweight transformer network model in detecting surface defects on industrial parts.
[0136] The above primarily describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, the lightweight transformer defect detection device or electronic device based on knowledge distillation includes corresponding hardware structures and / or software modules for performing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is implemented in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0137] This application embodiment can, according to the above method, exemplarily divide the lightweight transformer defect detection device or electronic device based on knowledge distillation into functional modules. For example, the lightweight transformer defect detection device or electronic device based on knowledge distillation may include functional modules corresponding to each functional division, or two or more functions may be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0138] Figure 10 This is a block diagram illustrating a lightweight transformer defect detection device based on knowledge distillation, according to an exemplary embodiment. (Refer to...) Figure 10 The lightweight transformer-based defect detection device 1000 based on knowledge distillation includes: a construction module 1001 and a training module 1002; the construction module 1001 is used to construct a defect detection dataset, wherein the defect detection dataset includes defect images of industrial parts; the defect images are labeled with actual defect information; the training module 1002 is used to train a student defect detection model based on the defect detection dataset and a teacher defect detection model; wherein the teacher defect detection model is a convolutional neural network (CNN) model for detecting defects in industrial parts; the student defect detection model is a transformer network model; the student defect detection model optimizes the model parameters by utilizing the feature difference between a first feature information and a second feature information; the first feature information is the image features of the defect images extracted by the teacher defect detection model, and the second feature information is the image features of the defect images extracted by the student defect detection model.
[0139] In one possible implementation, the training module 1002 is specifically used to determine the distillation loss between the first feature information and the second feature information based on the feature difference; the training module 1002 is specifically used to input the defect image into the student defect detection model to obtain the predicted defect information of the defect image; the training module 1002 is specifically used to determine the prediction loss based on the deviation between the predicted defect information and the actual defect information; the training module 1002 is specifically used to optimize the student defect detection model based on the distillation loss and the prediction loss.
[0140] In one possible implementation, the first feature extraction layer of the convolutional neural network (CNN) model includes multiple first feature extraction sub-layers, and the second feature extraction layer of the transformer network model includes multiple second feature extraction sub-layers; the first feature information is the feature information corresponding to a portion of the multiple first feature extraction sub-layers; and the second feature information is the feature information corresponding to a portion of the multiple second feature extraction sub-layers.
[0141] In one possible implementation, the first feature extraction layer of the convolutional neural network (CNN) model includes multiple first feature extraction sub-layers, and the second feature extraction layer of the transformer network model includes multiple second feature extraction sub-layers. The first feature information includes first global feature information corresponding to the first feature extraction layer and first local feature information corresponding to a portion of the multiple first feature extraction sub-layers. The second feature information includes second global feature information corresponding to the second feature extraction layer and second local feature information corresponding to a portion of the multiple second feature extraction sub-layers. The training module 1002 is specifically used to determine a global feature loss based on the feature difference between the first global feature information and the second global feature information; the training module 1002 is specifically used to determine a local feature loss based on the feature difference between the first local feature information and the second local feature information; and the training module 1002 is specifically used to determine a distillation loss based on the global feature loss and the local feature loss.
[0142] In one possible implementation, the predicted defect information includes: a predicted probability that the defect type is a preset defect type, a confidence level of the predicted probability, and a predicted defect location; the actual defect information includes: the actual defect type and the actual defect location; the training module 1002 is specifically used to determine a classification prediction loss based on the predicted probability and the actual defect type; the training module 1002 is specifically used to determine a location prediction loss based on the predicted defect location and the actual defect location; the training module 1002 is specifically used to determine a confidence loss based on the confidence level of the predicted probability; and the training module 1002 is specifically used to determine a prediction loss based on the classification prediction loss, the location prediction loss, and the confidence loss.
[0143] Figure 11 This is a block diagram illustrating a lightweight transformer defect detection device based on knowledge distillation, according to an exemplary embodiment. (Refer to...) Figure 11The lightweight transformer defect detection device 1100 based on knowledge distillation includes: an acquisition module 1101 and a detection module 1102; the acquisition module 1101 is used to acquire an image of a defect to be detected, which is an image of an industrial part to be detected; the detection module 1102 is used to input the image of the defect to be detected into a student defect detection model to detect the defect type and defect location of the industrial part to be detected.
[0144] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0145] Figure 12 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Figure 12 As shown, the electronic device 1200 includes, but is not limited to, a processor 1201 and a memory 1202.
[0146] The memory 1202 described above is used to store the executable instructions of the processor 1201. It is understood that the processor 1201 is configured to execute instructions to implement the lightweight transformer defect detection method based on knowledge distillation in the above embodiments.
[0147] It should be noted that those skilled in the art will understand that Figure 12 The electronic device structure shown does not constitute a limitation on the electronic device; the electronic device may include, but is not limited to, other electronic devices. Figure 12 This may indicate more or fewer components, or a combination of certain components, or a different arrangement of components.
[0148] Processor 1201 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in memory 1202, and by calling data stored in memory 1202, it performs various functions and processes data, thereby providing overall monitoring of the electronic device. Processor 1201 may include one or more processing units. Optionally, processor 1201 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into processor 1201.
[0149] The memory 1202 can be used to store software programs and various data. The memory 1202 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, application programs required by at least one functional module (such as a processing module, a storage module, etc.), etc. Furthermore, the memory 1202 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0150] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 1202 including instructions, which can be executed by a processor 1201 of an electronic device 1200 to implement the lightweight transformer defect detection method based on knowledge distillation in the above embodiments.
[0151] In actual implementation, Figure 10 The functions of the building module 1001 and the training module 1002, and Figure 11 The functions of the acquisition module 1101 and the detection module 1102 can both be provided by Figure 12 The processor 1201 calls the computer program stored in the memory 1202 to implement the process. The specific execution process can be found in the description of the lightweight transformer defect detection method based on knowledge distillation in the previous embodiment, and will not be repeated here.
[0152] Optionally, the computer-readable storage medium may be a non-transitory computer-readable storage medium, such as a read-only memory (ROM), random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.
[0153] In an exemplary embodiment, this application also provides a computer program product including one or more instructions, which can be executed by the processor 1201 of the electronic device 1200 to complete the lightweight transformer defect detection method based on knowledge distillation in the above embodiments.
[0154] It should be noted that when one or more instructions in the computer-readable storage medium or computer program product are executed by the processor of an electronic device, they implement the various processes of the above-described lightweight transformer defect detection method based on knowledge distillation, and can achieve the same technical effect as the above-described lightweight transformer defect detection method based on knowledge distillation. To avoid repetition, they will not be described again here.
[0155] Figure 13 A schematic diagram of the structure of a computer system for an electronic device is shown. It should be noted that... Figure 13 The computer system of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0156] like Figure 13 As shown, the computer system includes a Central Processing Unit (CPU), which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) or loaded from storage into Random Access Memory (RAM), such as executing the methods described in the above embodiments. The RAM also stores various programs and data required for system operation. The CPU, ROM, and RAM are interconnected via a bus. Input / Output (I / O) interfaces are also connected to the bus. The I / O interfaces are used to implement functions such as data input, output, communication, and storage; the storage function can be specifically implemented through removable media.
[0157] The following components are connected to the I / O interface: input sections including keyboards, mice, etc.; output sections including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage sections including hard drives; and communication sections including network interface cards such as LAN (Local Area Network) cards and modems. The communication sections perform communication processing via networks such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical discs, magneto-optical discs, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as required.
[0158] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), it performs various functions defined in the system of this application.
[0159] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0160] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0161] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0162] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the classified units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0163] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0164] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, essentially, or the part that contributes to the prior art, or a complete or partial classification of the technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0165] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A lightweight transformer defect detection method based on knowledge distillation, characterized in that, The method includes: Construct a defect detection dataset, wherein the defect detection dataset includes defect images of industrial parts; the defect images are labeled with actual defect information; Based on the aforementioned defect detection dataset and teacher defect detection model, train the student defect detection model; The teacher defect detection model is a convolutional neural network (CNN) model used to detect defects in industrial parts. The student defect detection model is a transformer network model; The student defect detection model optimizes model parameters by utilizing the feature differences between the first feature information and the second feature information; The first feature information is the image feature of the defect image extracted by the teacher defect detection model, and the second feature information is the image feature of the defect image extracted by the student defect detection model.
2. The lightweight transformer defect detection method based on knowledge distillation according to claim 1, characterized in that, The step of training a student defect detection model based on the defect detection dataset and the teacher defect detection model includes: The distillation loss between the first feature information and the second feature information is determined based on the feature differences; The defect image is input into the student defect detection model to obtain the predicted defect information of the defect image; The prediction loss is determined based on the deviation between the predicted defect information and the actual defect information; The student defect detection model is optimized based on the distillation loss and the prediction loss.
3. The lightweight transformer defect detection method based on knowledge distillation according to claim 2, characterized in that, The first feature extraction layer of the convolutional neural network (CNN) model includes multiple first feature extraction sub-layers, and the second feature extraction layer of the transformer network model includes multiple second feature extraction sub-layers. The first feature information is the feature information corresponding to a portion of the first feature extraction sub-layers among multiple first feature extraction sub-layers; The second feature information is the feature information corresponding to a portion of the second feature extraction sub-layers among the multiple second feature extraction sub-layers.
4. The lightweight transformer defect detection method based on knowledge distillation according to claim 2, characterized in that, The first feature extraction layer of the convolutional neural network (CNN) model includes multiple first feature extraction sub-layers, and the second feature extraction layer of the transformer network model includes multiple second feature extraction sub-layers. The first feature information includes the first global feature information corresponding to the first feature extraction layer and the first local feature information corresponding to a portion of the first feature extraction sub-layers in the plurality of first feature extraction sub-layers; The second feature information includes the second global feature information corresponding to the second feature extraction layer and the second local feature information corresponding to some of the second feature extraction sub-layers in the multiple second feature extraction sub-layers; Determining the distillation loss between the first feature information and the second feature information based on the feature differences includes: Based on the feature differences between the first global feature information and the second global feature information, the global feature loss is determined; Based on the feature differences between the first local feature information and the second local feature information, the local feature loss is determined; The distillation loss is determined based on the global feature loss and the local feature loss.
5. The lightweight transformer defect detection method based on knowledge distillation according to any one of claims 1-4, characterized in that, The predicted defect information includes: the predicted probability that the defect type is a preset defect type, the confidence level of the predicted probability, and the predicted defect location; the actual defect information includes: the actual defect type and the actual defect location. The step of determining the prediction loss based on the deviation between the predicted defect information and the actual defect information includes: Based on the predicted probability and the actual defect type, determine the classification prediction loss; Based on the predicted defect location and the actual defect location, determine the location prediction loss; Based on the confidence level of the predicted probability, determine the confidence loss; The prediction loss is determined based on the classification prediction loss, the location prediction loss, and the confidence loss.
6. A lightweight transformer defect detection method based on knowledge distillation, characterized in that, The method includes: Acquire an image of the defect to be detected, wherein the image of the defect to be detected is an image of the industrial part to be detected; The image of the defect to be detected is input into the student defect detection model according to any one of claims 1-5 to determine the defect type and defect location of the industrial part to be detected.
7. A lightweight transformer defect detection device based on knowledge distillation, characterized in that, The lightweight transformer defect detection device based on knowledge distillation includes: a construction module and a training module; The construction module is used to construct a defect detection dataset, wherein the defect detection dataset includes defect images of industrial parts; the defect images are labeled with actual defect information; The training module is used to train the student defect detection model based on the defect detection dataset and the teacher defect detection model. The teacher defect detection model is a convolutional neural network (CNN) model used to detect defects in industrial parts. The student defect detection model is a transformer network model; The student defect detection model optimizes model parameters by utilizing the feature differences between the first feature information and the second feature information; The first feature information is the image feature of the defect image extracted by the teacher defect detection model, and the second feature information is the image feature of the defect image extracted by the student defect detection model.
8. A lightweight transformer defect detection device based on knowledge distillation, characterized in that, The lightweight transformer defect detection device based on knowledge distillation includes: an acquisition module and a detection module; The acquisition module is used to acquire an image of a defect to be detected, wherein the image of the defect to be detected is an image of an industrial part to be detected; The detection module is used to input the image of the defect to be detected into the student defect detection model according to any one of claims 1-5, and to detect the defect type and defect location of the industrial part to be detected.
9. A vehicle, characterized in that, The vehicle includes industrial parts, which are defect-detected using the lightweight transformer defect detection method based on knowledge distillation as described in claim 6.
10. An electronic device, characterized in that, include: A processor and a memory; wherein the memory is used to store one or more programs, the one or more programs including computer-executable instructions, wherein when the electronic device is running, the processor executes the computer-executable instructions stored in the memory, and the electronic device executes any one of claims 1-5, and / or the lightweight transformer defect detection method based on knowledge distillation as described in claim 6.