Industrial detection classification model, model training method and model training device
By adopting the method of unsupervised pre-training model and frozen parameters in industrial detection, the problem of low detection accuracy of traditional algorithms in complex environments is solved, and efficient industrial parts classification is achieved.
Patent Information
- Application Number
- CN202311626986.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional algorithms are difficult to effectively extract complex and diverse defect characteristics of industrial parts in industrial inspection, resulting in low detection accuracy.
The big data unsupervised pre-trained model is used as the backbone model, combined with the classifier and the output prediction probability module, by freezing the backbone model parameters, only the classifier parameters are updated, and the unsupervised training method is used to improve feature extraction capabilities and classification accuracy.
It improves the training efficiency of industrial detection classification models, reduces computer computing power and video memory requirements, improves classification accuracy, and is suitable for complex and diverse industrial detection environments.
Smart Images

Figure CN120298735A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of industrial inspection, and particularly relates to an industrial inspection classification model, a model training method, a model training device, and a computer-readable storage medium. Background Art
[0002] In recent years, deep learning technology has been continuously developing, and deep learning models have also begun to be applied in the field of industrial inspection. When using a classification model to detect the qualified and unqualified categories of industrial parts and other products, due to the complex and diverse defect categories and forms of industrial products, it is difficult for traditional algorithms to extract effective features, resulting in poor detection effects and low classification accuracy. Summary of the Invention
[0003] Embodiments of this application provide an industrial inspection classification model, a model training method, a model training device, and a computer-readable storage medium to solve at least one of the above-mentioned technical problems.
[0004] The industrial inspection classification model according to the embodiments of this application includes:
[0005] A backbone model for extracting image features of an input image;
[0006] A classifier for mapping a multi-dimensional feature space to a linear feature space;
[0007] An output prediction probability module for performing feature normalization processing to output the prediction probability of the category to which the industrial part corresponding to the input image belongs;
[0008] Wherein, the backbone model adopts a big data unsupervised pre-training model.
[0009] The model training method according to the embodiments of this application is applied to an industrial inspection classification model, and the industrial inspection classification model includes a backbone model, a classifier, and an output prediction probability module. The model training method includes:
[0010] Inputting a training image into the backbone model;
[0011] Extracting the image features of the training image through the backbone model;
[0012] Mapping the multi-dimensional feature space to a linear feature space through the classifier;
[0013] Performing feature normalization processing through the output prediction probability module to output the prediction probability of the category to which the industrial part corresponding to the training image belongs;
[0014] Updating the model parameters according to the classification error of the prediction probability;
[0015] Among them, the backbone model adopts a big data unsupervised pre-training model.
[0016] In some embodiments, the model training method includes:
[0017] During the process of updating the model parameters, freeze the parameters of the backbone model;
[0018] Updating the model parameters according to the classification error of the prediction probability includes:
[0019] Update the parameters of the classifier according to the classification error of the prediction probability.
[0020] In some embodiments, the function of the classifier is:
[0021] y = xA T + b
[0022] Where x is the multi-dimensional feature vector extracted by the backbone model, A is the parameter matrix, b is the linear space bias parameter, and y is the output value of the classifier;
[0023] Updating the parameters of the classifier according to the classification error of the prediction probability includes:
[0024] Fit the parameter matrix and the linear space bias parameter according to the classification error of the prediction probability.
[0025] In some embodiments, the parameters of the classifier are fewer than those of the backbone model.
[0026] In some embodiments, the model training method further includes:
[0027] Train the big data unsupervised pre-training model using an unsupervised training method in a contrastive manner.
[0028] In some embodiments, the model training method further includes:
[0029] Train the big data unsupervised pre-training model using an unsupervised training method in a reconstruction manner.
[0030] The model training device according to the embodiments of the present application is applied to an industrial detection classification model, and the industrial detection classification model includes a backbone model, a classifier, and an output prediction probability module. The model training device includes an input module and an update module;
[0031] The input module is used to input training images into the backbone model;
[0032] The backbone model is used to extract the image features of the training images;
[0033] The classifier is used to map a multi-dimensional feature space to a linear feature space;
[0034] The output prediction probability module is used to perform feature normalization processing to output the prediction probability of the category to which the industrial part corresponding to the training image belongs;
[0035] The update module is used to update the model parameters according to the classification error of the prediction probability;
[0036] Wherein, the backbone model adopts a big data unsupervised pre-training model.
[0037] The model training device according to the embodiment of the present application, the model training device includes one or more processors and a memory, the memory stores a computer program, and when the computer program is executed by the processor, the model training method according to any one of the above embodiments is realized.
[0038] The computer-readable storage medium according to the embodiment of the present application, on which a computer program is stored, and when the program is executed by a processor, the model training method according to any one of the above embodiments is realized.
[0039] In the industrial detection classification model, model training method, model training device and computer-readable storage medium according to the embodiments of the present application, the backbone model adopts a big data unsupervised pre-training model. In this way, the training efficiency of the industrial detection classification model is effectively improved, the requirements for the computing power and video memory of the computer are reduced, thereby reducing the use threshold of the industrial detection classification model in the industrial detection field, and the industrial detection classification model can have a high classification accuracy.
[0040] The additional aspects and advantages of the embodiments of the present application will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, wherein:
[0042] Figure 1 is a schematic diagram of the modules of the industrial detection classification model according to some embodiments of the present application;
[0043] Figure 2 is a schematic diagram of the structure of the backbone model according to some embodiments of the present application;
[0044] Figure 3 is a schematic flowchart of the model training method according to some embodiments of the present application;
[0045] Figure 4It is a schematic diagram of the working process of the model training method according to some embodiments of the present application;
[0046] Figure 5 It is a schematic flowchart of the model training method according to some embodiments of the present application;
[0047] Figure 6 It is a schematic flowchart of the model training method according to some embodiments of the present application;
[0048] Figure 7 It is a schematic flowchart of the model training method according to some embodiments of the present application;
[0049] Figure 8 It is a schematic flowchart of the model training method according to some embodiments of the present application;
[0050] Figure 9 It is a schematic diagram of the modules of the model training device according to some embodiments of the present application;
[0051] Figure 10 It is a schematic diagram of the modules of the model training device according to some embodiments of the present application;
[0052] Figure 11 It is a schematic diagram of the connection state between the computer-readable storage medium and the processor according to some embodiments of the present application. Specific embodiments
[0053] The following further describes the embodiments of the present application with reference to the accompanying drawings. The same or similar reference numerals in the drawings represent the same or similar elements or elements having the same or similar functions throughout. In addition, the embodiments of the present application described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application, and should not be construed as a limitation of the present application.
[0054] Please refer to Figure 1 , the embodiments of the present application provide an industrial detection classification model 100. The industrial detection classification model 100 includes a backbone model 10, a classifier 20, and an output prediction probability module 30. The backbone model 10 is used to extract the image features of the input image. The classifier 20 is used to map the multi-dimensional feature space to a linear feature space. The output prediction probability module 30 is used to perform feature normalization processing to output the prediction probability of the category to which the industrial part corresponding to the input image belongs. Among them, the backbone model 10 adopts a big data unsupervised pre-training model.
[0055] In the industrial detection classification model 100 of the embodiment of the present application, the backbone model 10 adopts a big data unsupervised pre-training model. In this way, the training efficiency of the industrial detection classification model 100 is effectively improved, the requirements for the computing power and video memory of the computer are reduced, thereby reducing the usage threshold of the industrial detection classification model 100 in the field of industrial detection, and the industrial detection classification model 100 can have a high classification accuracy.
[0056] Among them, the input image is an image of an industrial part. The input image can be a single-channel image or a multi-channel image. For example, the input image can be a single-channel image, a three-channel image, or a four-channel image. The input image can be the training image in the following text or the detection image in the industrial detection process. When the input image is a training image, there are usually multiple images. For example, 100, 200, or 500 training images can be input for training. When the input image is a detection image, there can be any number of images. For example, 1 detection image can be input for individual detection, or 100, 200, or 500 detection images can be input for batch detection. The image features include: the number of images m, the number of image feature channels c, the height h of the image features, and the width w of the image features. The number of image feature channels c corresponds to the number of channels of the input image. The above image features together constitute a multi-dimensional image feature vector in the multi-dimensional image feature space. The number of image feature channels c, the height h of the image features, and the width w of the image features can be combined to obtain the deformed image feature chw. The multi-dimensional image feature vector contains the deformed image feature chw of m images, and the size of the multi-dimensional image feature vector is m×1. Feature normalization processing refers to normalizing the features output by the classifier 20 using the normalization exponential function (softmax). The normalization processing result represents the prediction probability of the category to which the industrial part corresponding to the input image belongs.
[0057] The big data unsupervised pre-training model refers to a machine learning model that is pre-trained in an unsupervised learning manner using a large number of unlabeled samples. Please refer to Figure 2, a big data unsupervised pre-training model is used as the backbone model 10. At this time, the backbone model 10 may include a convolutional layer 11, a normalization layer 12, an activation layer 13, and a pooling layer 14. The convolutional layer 11 is used to extract information in the input image, and this information is called image features. The normalization layer 12 is used to normalize each extracted image feature to the same range, reduce the interference caused by the range difference of each image feature, and accelerate the model convergence. The activation layer 13 is used to perform a non-linear mapping on the image features and map the image features to a non-linear interval. The pooling layer 14 is used to perform downsampling to compress the image features and reduce the dimension of the image features. It can be understood that the backbone model 10 may include n convolutional layers 11, normalization layers 12, and activation layers 13, as well as a corresponding pooling layer 14. The backbone model 10 may also include n pooling layers, and each pooling layer corresponds to n convolutional layers 11, normalization layers 12, and activation layers 13. The numbers of the convolutional layer 11, the normalization layer 12, the activation layer 13, and the pooling layer 14 can be adjusted according to the actual application scenario of industrial inspection.
[0058] In the implementation manner of this application, a big data unsupervised pre-training model is used as the backbone model 10 to model a multi-dimensional image feature space by extracting the number of images m, the number of image feature channels c, the height h of the image features, and the width w of the image features of the input image. The classifier 20 maps the multi-dimensional feature space to a linear feature space. The output prediction probability module 30 normalizes the linear features output by the classifier 20 into the prediction probability of the industrial part category corresponding to the input image and outputs it. According to the size of the prediction probability, it can be determined which category (such as a qualified category or an unqualified category) the industrial part corresponding to the input image belongs to, and the classification is completed.
[0059] It should be noted that traditional supervised training models rely on a large number of labeled samples, and it takes a lot of time and labor costs to label the samples, which is not suitable for complex and diverse industrial inspection environments. Unsupervised training models do not require sample labeling, have higher model training efficiency, and unsupervised training models have better image feature extraction capabilities and are suitable for complex and diverse industrial inspection environments. The big data unsupervised pre-training model has good feature extraction generalization and the ability to represent visual features in an open world, and can extract any image features, including the image features of industrial inspection images. Using the big data unsupervised pre-training model as the backbone model 10 to extract the image features of the input image can make the industrial inspection classification model 100 have a high classification accuracy.
[0060] Please refer to Figure 3 and Figure 4 , the implementation manner of this application provides a model training method, which is applied to the industrial inspection classification model 100. The industrial inspection classification model 100 includes a backbone model 10, a classifier 20, and an output prediction probability module 30. The model training method includes:
[0061] 010: Input the training image into the backbone model 10;
[0062] 020: Extract the image features of the training image through the backbone model 10;
[0063] 030: Map the multi-dimensional feature space to a linear feature space through the classifier 20;
[0064] 040: Perform feature normalization processing through the output prediction probability module 30 to output the prediction probability of the category to which the industrial part corresponding to the training image belongs;
[0065] 050: Update the model parameters according to the classification error of the prediction probability;
[0066] Among them, the backbone model 10 adopts a big data unsupervised pre-training model.
[0067] Specifically, the training image is input as an input image into the big data unsupervised pre-training model. The big data unsupervised pre-training model extracts the number of images m, the number of image feature channels c, the height h of the image features, and the width w of the image features of the training image to model a multi-dimensional feature space. The classifier 20 maps the multi-dimensional feature space to a linear feature space. The output prediction probability module 30 normalizes the features output by the classifier 20 through the softmax function to the prediction probability of the category to which the industrial part corresponding to the input image belongs and outputs it. Update the parameters of the industrial detection classification model 100 according to the classification error of the prediction probability of the category to which the industrial part belongs, so as to optimize the parameters of the industrial detection classification model 100.
[0068] Please refer to Figure 4 and Figure 5 , in some embodiments, the model training method includes:
[0069] 060: During the process of updating the model parameters, freeze the parameters of the backbone model 10;
[0070] Updating the model parameters according to the classification error of the prediction probability (i.e., 050) includes:
[0071] 051: Update the parameters of the classifier 20 according to the classification error of the prediction probability.
[0072] It can be understood that the image feature extraction ability of the backbone model 10 for the input image will affect the classification accuracy of the industrial detection classification model 100. In order to better extract the image features of the input image and model a multi-dimensional image feature space, it is usually necessary to update the parameters of the backbone model 10. However, the backbone model 10 usually contains multiple convolutional and non-linear operations, which involve a large amount of computation and require a high computer video memory. Updating the parameters of the backbone model 10 takes a lot of time and has a low training fitting efficiency. The model training method of the embodiment of the present application uses a big data unsupervised pre-training model as the backbone model 10. By utilizing the characteristics of the big data unsupervised pre-training model, which has good feature extraction generalization and the ability to represent visual features in the open world, it can well extract the image features of the input image without updating its model parameters. Therefore, during the process of updating the model parameters, the parameters of the backbone model 10 can be frozen. In this way, the training efficiency of the industrial detection classification model 100 is effectively improved, the requirements for the computing power and video memory of the computer are reduced, and thus the usage threshold of the industrial detection classification model 100 in the industrial detection field is lowered. At the same time, the industrial detection classification model 100 still maintains a high classification accuracy.
[0073] In the embodiment of the present application, during the process of updating the parameters of the industrial detection classification model 100, the parameters of the backbone model 10 can be frozen, and only the parameters of the classifier 20 need to be updated by backpropagation according to the classification error of the prediction probability to optimize the parameters of the classifier 20, so as to complete the training of the industrial detection classification model 100.
[0074] Please refer to Figure 6 , in some embodiments, the function of the classifier 20 is:
[0075] y = xA T + b
[0076] where x is the multi-dimensional feature vector extracted by the backbone model 10, A is the parameter matrix, b is the linear space bias parameter, and y is the output value of the classifier 20;
[0077] Updating the parameters of the classifier 20 (i.e., 051) according to the classification error of the prediction probability includes:
[0078] 0511: Fitting the parameter matrix and the linear space bias parameter according to the classification error of the prediction probability.
[0079] Among them, the multi-dimensional feature vector x is the multi-dimensional feature vector composed of the number of images m, the number of image feature channels c, the height h of the image features, and the width w of the image features of the above training images. The number of image feature channels c, the height h of the image features, and the width w of the image features are combined to obtain the feature chw. The multi-dimensional image feature vector contains the features chw of m images, and the size of the multi-dimensional image feature vector is m×1. A is a parameter matrix composed of the number of image feature channels c, the height h of the image features, and the width w of the image features of n predicted categories. The number of image feature channels c, the height h of the image features, and the width w of the image features are combined to obtain the feature chw, and the size of the parameter matrix A is n×1. Matrix A T with the superscript T indicating the transpose of the parameter matrix A, that is, matrix A T is the transposed matrix of the parameter matrix A. It can be understood that the parameter matrix A needs to be transposed to matrix A T to perform operations with the multi-dimensional feature vector x. The linear space bias parameter b is a matrix with a size of m×n. The classifier 20 maps the multi-dimensional feature space to the linear feature space through the function y = xA T +b, linearly transforming the multi-dimensional feature vector x of the training image into the image linear feature y. The size of the image linear feature y is m×n, and it contains the linear features of m training images under n predicted categories.
[0080] In the implementation manner of this application, the classifier 20 outputs the image linear feature y, which is transformed into the prediction probability that the industrial part corresponding to the training image belongs to each category after normalization processing and then output. According to the classification error of the prediction probability of the category to which the industrial part belongs, backpropagation is used to update the parameter matrix A and the linear space bias parameter b of the classifier 20, and the parameter matrix A and the linear space bias parameter b are fitted to optimize the parameter matrix A and the linear space bias parameter b. Taking the prediction probability of a training image as an example, for this training image for three predicted categories, the prediction probabilities for category 1, category 2, and category 3 are 80%, 60%, and 30% respectively. It is predicted that the industrial part corresponding to this training image belongs to category 1. At this time, the classification error between the prediction probability and the actual category is 20%. According to this classification error, backpropagation is used to update the parameter matrix A and the linear space bias parameter b of the classifier 20, and the parameter matrix A and the linear space bias parameter b are fitted to make the classification error between the prediction probability and the actual category as small as possible (for example, approaching 0), that is, to optimize the parameter matrix A and the linear space bias parameter b, and thus the training of the industrial detection classification model 100 can be completed.
[0081] In some implementation manners, the parameters of the classifier 20 are fewer than those of the backbone model 10.
[0082] Specifically, the backbone model 10 includes multiple convolutional layers 11, each convolutional layer 11 containing a large number of parameters for extracting image features, while the parameter matrix A and the linear space bias parameter b included in the classifier 20 are much fewer than the parameters of the entire backbone model 10. When training the industrial detection classification model 100, only a small number of training images need to be input to fine-tune the parameter matrix A and the linear space bias parameter b of the classifier 20, and thus the fitting of the parameter matrix A and the linear space bias parameter b of the classifier 20 can be completed. In this way, the training of the industrial detection classification model 100 can be completed by inputting a small number of training images, effectively improving the training efficiency of the industrial detection classification model 100, reducing the requirements for the computing power and video memory of the computer, and thus lowering the usage threshold of the industrial detection classification model 100 in the field of industrial detection.
[0083] Please refer to Figure 7 , in some embodiments, the model training method further includes:
[0084] 070: Training the big data unsupervised pre-training model using an unsupervised training method in a contrastive manner.
[0085] Specifically, the unsupervised training method in a contrastive manner refers to performing data augmentation on a training image to obtain multiple homogeneous training images, and learning the common features among the homogeneous training images through training to distinguish the differences between different types of training images. The goal of the unsupervised training method in a contrastive manner is to make the big data unsupervised pre-training model encode the image features of homogeneous training images similarly, and make the encoding results of the image features of different types of training images as different as possible. For example, performing data augmentation on the training image of an industrial part, that is, performing operations such as cropping, resizing, color distortion, or grayscaling on the training image. One data augmentation operation can be performed on the training image, or any combination of multiple data augmentation operations can be performed on the training image. For example, only cropping can be performed on the training image, or the training image can be resized after being grayscaled. Multiple homogeneous training images can be obtained from one training image, and all training images and the training images after data augmentation are used to train and optimize the big data unsupervised pre-training model, so that the finally obtained big data unsupervised pre-training model encodes the image features of homogeneous input images similarly, and makes the encoding results of the image features of different types of input images as different as possible, thereby enabling the big data unsupervised pre-training model to learn more discriminative representations.
[0086] It should be noted that the big data unsupervised pre-training model can be trained using any unsupervised training method based on contrast. For example, a series of unsupervised training methods based on contrast can be used, such as Momentum Contrast (MoCo), Simple Framework for Contrastive Learning of Visual Representations (SimCLR), Swapping Assignments between multiple Views of the same images (SwAV), and a form of knowledge distillation with no labels (dino), etc., to train the big data unsupervised pre-training model.
[0087] Please refer to Figure 8 , in some embodiments, the model training method further includes:
[0088] 080: Training the big data unsupervised pre-training model using an unsupervised training method based on reconstruction.
[0089] Specifically, the unsupervised training method based on reconstruction means randomly covering a large number of pixel blocks in the training images, extracting the image features of the uncovered training images, reconstructing the covered pixel blocks and outputting them to restore the training images. By randomly covering a large number of pixel blocks in the images, the big data unsupervised pre-training model is forced to learn better representations and reconstruct the missing pixels. The finally obtained big data unsupervised pre-training model has strong image feature extraction capabilities. The big data unsupervised pre-training model can adopt any unsupervised training method based on reconstruction. For example, masked autoencoders (MAE) can be used to train the big data unsupervised pre-training model.
[0090] It should be noted that, according to different industrial detection scenarios, an unsupervised training method with different comparison methods or an unsupervised training method with different reconstruction methods can be selected. The big data unsupervised pre-training model can also select any classification model. For example, classification models of the Deep Residual Network (ResNet) series and the Vision Transformer (ViT) series can be adopted. Different models have different scales and sizes, and different classification models will have different efficiencies and accuracies. The big data unsupervised pre-training model pre-trained by using a complex unsupervised training method will have better image feature space representation ability, and the extracted image features will be better and richer, and the classification accuracy of the industrial detection classification model 100 will be higher. The classification accuracy of a classification model with a general size scale is relatively weak, but the efficiency will be higher. In practical applications, different classification models and unsupervised training methods can be selected according to different industrial detection scenario requirements, so that the industrial detection classification model 100 can adapt to the requirements of complex and diverse industrial detection scenarios and improve the robustness of the industrial detection classification model 100.
[0091] Please refer to Figure 9 , an embodiment of the present application provides a model training device 200, which is applied to the industrial detection classification model 100. The industrial detection classification model 100 includes a backbone model 10, a classifier 20, and an output prediction probability module 30. The model training device 200 includes an input module 210 and an update module 220. The input module 210 is used to input training images into the backbone model 10. The backbone model 10 is used to extract image features of the training images. The classifier 20 is used to map the multi-dimensional feature space to a linear feature space. The output prediction probability module 30 is used to perform feature normalization processing to output the prediction probability of the category to which the industrial part corresponding to the training image belongs. The update module 220 is used to update the model parameters according to the classification error of the prediction probability. Among them, the backbone model 10 adopts a big data unsupervised pre-training model.
[0092] In some embodiments, the model training device 200 includes a freezing module. The freezing module is used to freeze the parameters of the backbone model 10 during the process of updating the model parameters. The update module 220 is specifically used to update the parameters of the classifier 20 according to the classification error of the prediction probability.
[0093] In some embodiments, the function of the classifier 20 is:
[0094] y = xA T + b
[0095] where x is the multi-dimensional feature vector extracted by the backbone model 10, A is the parameter matrix, b is the linear space bias parameter, and y is the output value of the classifier 20;
[0096] The update module 220 is specifically configured to fit the parameter matrix and the linear space bias parameter according to the classification error of the prediction probability.
[0097] In some embodiments, the classifier 20 has fewer parameters than the backbone model 10.
[0098] In some embodiments, the model training device 200 further includes a training module. The training module is used to train the big data unsupervised pre-training model by using an unsupervised training method in a contrastive manner.
[0099] In some embodiments, the model training device 200 further includes a training module. The training module is used to train the big data unsupervised pre-training model by using an unsupervised training method in a reconstruction manner.
[0100] It should be noted that the explanations of the model training method in the foregoing embodiments are equally applicable to the model training device 200 in the embodiments of the present application, and will not be elaborated herein.
[0101] Please refer to Figure 10 , the embodiments of the present application further provide a model training device 300. The model training device 300 includes one or more processors 310 and a memory 320. The memory 320 stores a computer program, and when the computer program is executed by the processor 310, the model training method of any of the foregoing embodiments is implemented.
[0102] For example, when the computer program is executed by the processor 310, the following model training method is implemented:
[0103] 010: Input the training image into the backbone model 10;
[0104] 020: Extract the image features of the training image through the backbone model 10;
[0105] 030: Map the multi-dimensional feature space to a linear feature space through the classifier 20;
[0106] 040: Perform feature normalization processing through the output prediction probability module 30 to output the prediction probability of the category to which the industrial part corresponding to the training image belongs;
[0107] 050: Update the model parameters according to the classification error of the prediction probability;
[0108] Wherein, the backbone model 10 adopts a big data unsupervised pre-training model.
[0109] Again, for example, when the computer program is executed by the processor 310, the following model training method is implemented:
[0110] 060: During the process of updating the model parameters, freeze the parameters of the backbone model 10;
[0111] Update the model parameters (i.e., 050) according to the classification error of the prediction probability, including:
[0112] 051: Update the parameters of the classifier 20 according to the classification error of the prediction probability.
[0113] It should be noted that the explanations of the model training method in the foregoing embodiments are equally applicable to the model training device 300 of the embodiments of the present application, and will not be elaborated here.
[0114] Please refer to Figure 11 , the embodiments of the present application also provide a computer-readable storage medium 400, on which a computer program 410 is stored. When the program is executed by a processor 420, the model training method of any of the foregoing embodiments is implemented.
[0115] For example, when the program is executed by the processor 420, the following model training method is implemented:
[0116] 010: Input the training image into the backbone model 10;
[0117] 020: Extract the image features of the training image through the backbone model 10;
[0118] 030: Map the multi-dimensional feature space to a linear feature space through the classifier 20;
[0119] 040: Perform feature normalization processing through the output prediction probability module 30 to output the prediction probability of the category to which the industrial part corresponding to the training image belongs;
[0120] 050: Update the model parameters according to the classification error of the prediction probability;
[0121] Among them, the backbone model 10 adopts a big data unsupervised pre-training model.
[0122] Again, for example, when the program is executed by the processor 420, the following model training method is implemented:
[0123] 060: During the process of updating the model parameters, freeze the parameters of the backbone model 10;
[0124] Update the model parameters (i.e., 050) according to the classification error of the prediction probability, including:
[0125] 051: Update the parameters of the classifier 20 according to the classification error of the prediction probability.
[0126] It should be noted that the explanations of the model training method and the model training device 200 in the foregoing embodiments are equally applicable to the computer-readable storage medium 400 of the embodiments of the present application, and will not be elaborated herein.
[0127] In summary, for the industrial detection classification model 100, the model training method, the model training device 200, and the computer-readable storage medium 400 of the embodiments of the present application, the industrial detection classification model 100 includes a backbone model 10, a classifier 20, and an output prediction probability module 30. The backbone model 10 is used to extract image features of the input image. The classifier 20 is used to map the multi-dimensional feature space to a linear feature space. The output prediction probability module 30 is used to perform feature normalization processing to output the prediction probability of the category to which the industrial part corresponding to the input image belongs. Among them, the backbone model 10 adopts a big data unsupervised pre-training model. In this way, the training efficiency of the industrial detection classification model 100 is effectively improved, the requirements for the computing power and video memory of the computer are reduced, thereby reducing the usage threshold of the industrial detection classification model 100 in the field of industrial detection, and the industrial detection classification model 100 can have a high classification accuracy.
[0128] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0129] Any process or method description in the flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art of the embodiments of the present application.
[0130] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definable sequence of executable instructions for implementing a logical function, and can be embodied specifically in any computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a computer-readable storage medium can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable storage medium include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable storage medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0131] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0132] Those of ordinary skill in the art can understand that all or part of the steps carried out in the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments. In addition, in each of the embodiments of the present application, the functional units can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. If the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage media mentioned above can be read-only memories, magnetic disks or optical discs, etc.
[0133] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application. The scope of the present application is defined by the claims and their equivalents.
Claims
1. An industrial inspection and classification model, characterized in that, It includes: A backbone model for extracting image features of the input image; A classifier for mapping a multi-dimensional feature space to a linear feature space; An output prediction probability module for performing feature normalization processing to output the prediction probability of the category to which the industrial part corresponding to the input image belongs; Wherein, the backbone model adopts a big data unsupervised pre-training model.
2. A model training method, characterized in that, Applied to an industrial detection classification model, the industrial detection classification model includes a backbone model, a classifier and an output prediction probability module, and the model training method includes: Inputting a training image into the backbone model; Extracting the image features of the training image through the backbone model; Mapping the multi-dimensional feature space to a linear feature space through the classifier; Performing feature normalization processing through the output prediction probability module to output the prediction probability of the category to which the industrial part corresponding to the training image belongs; Updating the model parameters according to the classification error of the prediction probability; Wherein, the backbone model adopts a big data unsupervised pre-training model.
3. The model training method according to claim 2, wherein The model training method includes: During the process of updating the model parameters, freezing the parameters of the backbone model; The updating the model parameters according to the classification error of the prediction probability includes: Updating the parameters of the classifier according to the classification error of the prediction probability.
4. The model training method according to claim 3, wherein The function of the classifier is: y = xAT + b Wherein, x is the multi-dimensional feature vector extracted by the backbone model, A is the parameter matrix, b is the linear space bias parameter, and y is the output value of the classifier; The updating the parameters of the classifier according to the classification error of the prediction probability includes: Fitting the parameter matrix and the linear space bias parameter according to the classification error of the prediction probability.
5. The model training method according to claim 3, wherein The parameters of the classifier are fewer than those of the backbone model.
6. The model training method according to claim 2, wherein The model training method further includes: Training the big data unsupervised pre-training model by using an unsupervised training method in a contrast manner.
7. The model training method according to claim 2, wherein The model training method further includes: Training the big data unsupervised pre-training model by using an unsupervised training method in a reconstruction manner.
8. A model training device, characterized in that, Applied to an industrial detection classification model, the industrial detection classification model includes a backbone model, a classifier and an output prediction probability module, and the model training device includes an input module and an update module; The input module is used for inputting a training image into the backbone model; The backbone model is used for extracting the image features of the training image; The classifier is used for mapping the multi-dimensional feature space to a linear feature space; The output prediction probability module is used for performing feature normalization processing to output the prediction probability of the category to which the industrial part corresponding to the training image belongs; The update module is used for updating the model parameters according to the classification error of the prediction probability; Wherein, the backbone model adopts a big data unsupervised pre-training model.
9. A model training device, characterized in that The model training device includes one or more processors and a memory, and the memory stores a computer program, and when the computer program is executed by the processor, the model training method according to any one of claims 2-7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the model training method according to any one of claims 2-7 is implemented.