Defect identification model training method, defect identification method, device and equipment
By using a pre-trained model based on space and channel attention mechanism in the distribution cabinet defect recognition and training on the small sample data set, the problem of performance degradation under the existing technology of small sample multi-defect recognition tasks is solved, and the accuracy and generalization ability of defect recognition are improved.
Patent Information
- Application Number
- CN202510245658.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-23
AI Technical Summary
The existing intelligent identification algorithm has deteriorated performance under the small sample multiple defect recognition task, poor generalization capabilities, and it is difficult to accurately identify defects in the distribution cabinet.
A pre-trained defect recognition model is adopted based on spatial attention mechanism, channel attention mechanism and basic model, and a small sample data set is used to train it to obtain the target defect recognition model.
It improves the accuracy of the identification of power distribution cabinet defects, solves the problem of performance degradation of identification tasks under small sample data, and enhances the generalization ability of the model on small sample data.
Smart Images

Figure CN120032183A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of defect detection technology, and in particular to a defect recognition model training method, a defect recognition method, a device and equipment. Background Art
[0002] Due to the long-term bumps and vibrations during train operation, the electrical equipment in the distribution cabinet may have loose terminals, loose wire sheaths, hot melting of cables, and other abnormal conditions, causing the distribution cabinet to malfunction. If detection, positioning, and maintenance are not timely, it will cause significant economic losses, lead to railway safety accidents, and even cause catastrophic consequences. However, due to insufficient data volume, the existing intelligent recognition algorithms inevitably suffer from performance degradation in small sample multi-defect recognition tasks, poor generalization ability, and difficulty in accurate defect recognition.
[0003] Therefore, how to improve the accuracy of defect identification is a technical problem that technical personnel in this field need to solve urgently. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a defect recognition model training method, a defect recognition method, a device and equipment, which solve the technical problem of low accuracy of defect recognition of distribution cabinets in the prior art.
[0005] In order to solve the above technical problems, the present invention provides a defect recognition model training method, comprising:
[0006] Annotate the defects in each power distribution cabinet defect image in the power distribution cabinet defect image dataset to obtain a small sample dataset; wherein the small sample dataset is a dataset whose data volume is less than a first data volume threshold;
[0007] Obtain a pre-trained defect recognition model constructed based on a spatial attention mechanism, a channel attention mechanism, and a basic model; wherein the basic model is a model trained based on a large sample data set; the large sample data set is a data set whose data volume is greater than a second data volume threshold;
[0008] The pre-trained defect recognition model is trained using the small sample data set to obtain a target defect recognition model.
[0009] Optionally, the defects in each power distribution cabinet defect image in the power distribution cabinet defect image dataset are annotated to obtain a small sample dataset, including:
[0010] Constructing the distribution cabinet defect image dataset based on the distribution cabinet defect image; wherein the distribution cabinet defect image is a high-speed train distribution cabinet defect image;
[0011] Cropping an area of interest in each power distribution cabinet defect image in the power distribution cabinet defect image dataset to remove information irrelevant to defect detection, thereby obtaining a cropped power distribution cabinet defect image dataset;
[0012] A normalization operation is performed on the images in the cropped distribution cabinet defect image dataset to obtain the small sample dataset.
[0013] Optionally, before obtaining the pre-trained defect recognition model based on the spatial attention mechanism, the channel attention mechanism, and the basic model, the following is also included:
[0014] Adding a spatial attention module based on the spatial attention mechanism and a channel attention module based on the channel attention mechanism to the end of the basic model to obtain an initial defect recognition model;
[0015] The initial defect recognition model is trained based on a large sample defect data set to obtain the pre-trained defect recognition model; wherein the large sample defect data set includes a distribution cabinet defect data set.
[0016] Optionally, the pre-trained defect recognition model is trained using the small sample data set to obtain a target defect recognition model, including:
[0017] The target defect recognition model is obtained by freezing some convolutional layers of the pre-trained defect recognition model and using the small sample data set to train the initialized defect recognition model; wherein the some convolutional layers are convolutional layers whose parameters do not need to be adjusted.
[0018] Optionally, after the pre-trained defect recognition model is trained using the small sample data set to obtain a target defect recognition model, the method further includes:
[0019] The target defect recognition model is trimmed based on the entropy value to obtain a lightweight target defect recognition model; wherein the entropy value trimming method is a method of determining the importance of the filter based on the entropy value and trimming the filter based on the importance.
[0020] Optionally, the entropy-based clipping method clips the target defect recognition model to obtain a lightweight target defect recognition model, including:
[0021] Divide each filter in the target defect recognition model into multiple buckets, determine the probability of each bucket, and determine the entropy value corresponding to the current filter through the probability of each bucket;
[0022] Determine whether the entropy value is greater than a set minimum entropy value threshold;
[0023] When the entropy value of the filter is greater than the minimum entropy value threshold, determining not to trim the current filter;
[0024] When the entropy value of the filter is not greater than the minimum entropy value threshold, it is determined to trim the current filter.
[0025] Optionally, the target defect recognition model is trimmed based on an entropy value to obtain a lightweight target defect recognition model, including:
[0026] The target defect recognition model is trimmed based on the trimming method of the entropy value to obtain an initial lightweight target defect recognition model;
[0027] The initial lightweight target defect recognition model is fine-tuned to obtain the lightweight target defect recognition model, so that the output of the lightweight target defect recognition model is consistent with the output of the target defect recognition model.
[0028] The present invention also provides a defect identification method, comprising:
[0029] Acquire the image of the power distribution cabinet of the high-speed train to be inspected;
[0030] The target defect recognition model is used to perform defect recognition on the image of the high-speed train distribution cabinet to be detected to determine the fault type of the distribution cabinet; wherein the target defect recognition model is a model obtained based on the above-mentioned defect recognition model training method.
[0031] Optionally, the using a target defect recognition model to perform defect recognition on the image of the high-speed train power distribution cabinet to be detected to determine the fault type of the power distribution cabinet includes:
[0032] The target defect recognition model is used to perform defect recognition on the image of the high-speed train power distribution cabinet to be detected, and the fault type and fault location of the power distribution cabinet are determined.
[0033] The present invention also provides a defect recognition model training device, comprising:
[0034] A small sample data set determination module is used to mark the defects in each power distribution cabinet defect image in the power distribution cabinet defect image data set to obtain a small sample data set; wherein the small sample data set is a data set whose data volume is less than a first data volume threshold;
[0035] A pre-trained defect recognition model determination module is used to obtain a pre-trained defect recognition model constructed based on a spatial attention mechanism, a channel attention mechanism and a basic model; wherein the basic model is a model obtained by training based on a large sample data set; the large sample data set is a data set whose data volume is greater than a second data volume threshold;
[0036] The target defect model determination module is used to train the pre-trained defect recognition model using the small sample data set to obtain a target defect recognition model.
[0037] The present invention also provides a defect identification device, comprising:
[0038] A data acquisition module, used to acquire an image of a power distribution cabinet of a high-speed train to be inspected;
[0039] The defect recognition module is used to use the target defect recognition model to perform defect recognition on the image of the high-speed train distribution cabinet to be detected and determine the fault type of the distribution cabinet; wherein the target defect recognition model is a model obtained based on the above-mentioned defect recognition model training method.
[0040] The present invention also provides an electronic device, comprising:
[0041] Memory for storing computer programs;
[0042] A processor is used to execute the computer program to implement the steps of the above-mentioned defect recognition model training method and the steps of the above-mentioned defect recognition method.
[0043] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned defect recognition model training method and the steps of the above-mentioned defect recognition method are implemented.
[0044] The present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned defect recognition model training method and the steps of the above-mentioned defect recognition method.
[0045] It can be seen that the present invention obtains a small sample data set by marking the defects in each distribution cabinet defect image in the distribution cabinet defect image data set; wherein the small sample data set is a data set with a data volume less than the first data volume threshold; obtains a pre-trained defect recognition model constructed based on a spatial attention mechanism, a channel attention mechanism and a basic model; wherein the basic model is a model obtained by training based on a large sample data set; the large sample data set is a data set with a data volume greater than the second data volume threshold; the pre-trained defect recognition model is trained using a small sample data set to obtain a target defect recognition model. Compared with the current complex background and small target, and the insufficient sample size of the existing intelligent recognition algorithm during training, resulting in a relatively low diagnostic accuracy, the present application uses a small sample data set to train a pre-trained defect recognition model based on a spatial attention mechanism, a channel attention mechanism and a basic model, and the basic model is a model obtained by training using a large sample data set, which makes up for the defect of insufficient data set, and based on the trained spatial attention mechanism and channel attention mechanism, respectively simulates the semantic interdependence in the spatial and channel dimensions, which can solve the problem that the existing intelligent recognition algorithm is difficult to capture the full picture of data distribution under the small sample multi-defect recognition task, thereby improving the accuracy of distribution cabinet defect recognition.
[0046] In addition, the present invention also provides a defect recognition model training device, equipment and computer-readable storage medium, as well as a defect recognition method, device, equipment and computer-readable storage medium, which also have the above-mentioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0048] Figure 1 A flowchart of a defect recognition model training method provided by an embodiment of the present invention;
[0049] Figure 2 A schematic diagram of a defect image of a power distribution cabinet provided by an embodiment of the present invention;
[0050] Figure 3 A dual attention mechanism based on a spatial attention mechanism and a channel attention mechanism is provided in an embodiment of the present invention;
[0051] Figure 4 An example diagram of a process for fine-tuning a basic model provided in an embodiment of the present invention;
[0052] Figure 5 A flowchart of a network pruning method provided by an embodiment of the present invention;
[0053] Figure 6 An example flowchart of a defect identification method provided by an embodiment of the present invention;
[0054] Figure 7 An example diagram of a defect recognition model training and use process provided by an embodiment of the present invention;
[0055] Figure 8 A schematic diagram of a framework for training and using a defect recognition model provided by an embodiment of the present invention;
[0056] Fig. 9 A schematic diagram of a defect recognition framework provided by an embodiment of the present invention;
[0057] Fig.10 A structural schematic diagram of a defect recognition model training device provided by an embodiment of the present invention;
[0058] Fig.11 A schematic diagram of the structure of a defect identification device provided by an embodiment of the present invention;
[0059] Fig.12 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0060] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0061] Please refer to Figure 1 , Figure 1 A flowchart of a defect recognition model training method provided by an embodiment of the present invention. The method may include:
[0062] S101, marking defects in each distribution cabinet defect image in the distribution cabinet defect image dataset to obtain a small sample dataset; wherein the small sample dataset is a dataset whose data volume is less than a first data volume threshold.
[0063] The execution subject of this embodiment is an electronic device. The electronic device in this embodiment can be a distribution cabinet device; or the electronic device in this embodiment can be a computer. The distribution cabinet defect image data set in this embodiment is a data set in which the amount of all defect image data is lower than the first data amount threshold. This embodiment does not limit the specific defect types included in the distribution cabinet defect image data set. For example, it can be a high-speed train distribution cabinet (a high-speed train distribution cabinet is an electrical device used on a high-speed train, responsible for distributing electrical energy from the main power supply of the train to various electrical devices). Generally, the more types of distribution cabinet defect images collected, the more accurate the training of the model. Therefore, the types of distribution cabinet defect images in the distribution cabinet defect image data set can be all types in the prior art, such as loose terminals, burnt and discolored terminals, loose wire sheaths, broken terminals, abnormal motor starter position, and missing resistor disks. Or the distribution cabinet in this embodiment can also be a high-voltage distribution cabinet (with relatively few defect images). In this embodiment, the defects in each distribution cabinet defect image can be annotated to obtain an annotated image; or this embodiment can also annotate the defects and defect locations so that the defects and defect locations can be identified later. This embodiment does not limit the first data volume threshold; for example, the first data volume threshold in this embodiment may be 80; or the first data volume threshold in this embodiment may be 90.
[0064] It should be further explained that, based on any of the above embodiments, in order to improve the accuracy of constructing a small sample data set, the above-mentioned labeling of defects in each distribution cabinet defect image in the distribution cabinet defect image data set to obtain a small sample data set may include:
[0065] S1011, constructing a distribution cabinet defect image dataset based on the distribution cabinet defect image; wherein the distribution cabinet defect image is a high-speed train distribution cabinet defect image;
[0066] S1012, cropping the region of interest in each power distribution cabinet defect image in the power distribution cabinet defect image dataset to remove information irrelevant to defect detection, thereby obtaining a cropped power distribution cabinet defect image dataset;
[0067] S1013, performing a normalization operation on the cropped images in the distribution cabinet defect image dataset to obtain a small sample dataset.
[0068] The defect images of high-speed train power distribution cabinets in the power distribution cabinet defect image dataset of this embodiment include loose terminals, burnt and discolored terminals, loose wire sheaths, broken terminals, abnormal motor starter positions, and missing resistor disks. Figure 2 , Figure 2A schematic diagram of a defective image of a power distribution cabinet provided in an embodiment of the present invention. After collecting data, this embodiment preprocesses the image to meet the input requirements of the model and further improve the generalization. The specified coordinates of the image are cropped to remove the interference of irrelevant information of the image on the classification of image quality defects, thereby weakening data noise and increasing model stability, and also avoiding image deformation caused by changes in aspect ratio when resizing the image. Secondly, the image is normalized to reduce the value of the input layer, facilitate the selection of learning rate, and increase the training speed. Then, in order to enrich the data and enhance the generalization ability of the model, this paper can perform data enhancement technology on the above-mentioned processed data to expand the number of samples and form a small sample data set. Commonly used data enhancement techniques are shown in Table 1, which is an example table of a data enhancement method provided in an embodiment of the present invention.
[0069] Table 1 Example table of a data augmentation method
[0070]
[0071] S102, obtaining a pre-trained defect recognition model constructed based on a spatial attention mechanism, a channel attention mechanism and a basic model; wherein the basic model is a model trained based on a large sample data set; and the large sample data set is a data set whose data volume is greater than a second data volume threshold.
[0072] This embodiment does not limit the second data volume threshold. The second data volume threshold in this embodiment may be the same as the first data volume threshold or may be different. For example, the second data volume threshold in this embodiment may be 1000; or the second data volume threshold in this embodiment may also be 2000. The spatial attention mechanism in this embodiment is a CA attention mechanism. CA attention is an efficient attention mechanism that selectively aggregates the features of each position by weighting the features at all positions, alleviating the defect of position information loss caused by global pooling. CA attention decomposes channel attention into two parallel 1D (identification) feature encoding processes, effectively integrating spatial coordinate information into the generated attention map, and realizing the model's precise pointing to the target object of interest. For the input X, the pooling kernels (H, 1) and (1, W) are used to encode the horizontal and vertical features, that is, the output of the c-th dimension feature is: ; ;in, Indicates that in the vertical direction, all pixel values of the hth row of the cth dimension feature map are averaged to obtain the pooled output in the horizontal direction. Indicates that the c-th dimension feature map is at position The pixel value of represents the width of the feature map (the size of the horizontal dimension), c represents the channel index of the feature map, that is, the c-th dimension feature, and h represents the index in the horizontal direction (width). Represents the output value after pooling the c-th dimension feature map in the vertical direction, where is the vertical index; H: represents the height of the feature map (the size of the vertical dimension); w represents the vertical index (height), Represents the pixel value of the c-th dimension feature map at position (j, w). Use Concat to combine the feature maps of two directions, and use convolution, BN and non-linear activation for feature transformation. ,in, is the intermediate feature containing horizontal and vertical spatial information, r is the reduction factor, Usually represents a nonlinear activation function, such as ReLU, etc. Refers to a network layer or operation that contains 1x1 convolution, batch normalization (BN), and nonlinear activation. Indicates that the horizontal feature and vertical features To splice, represents a set of real numbers, and C represents the total number of features or channels. Divide f into two independent features and , using the other two Convolution, Sigmoid function and reverse feature transformation make its dimension consistent with input X: ; , , Represents the output features after feature conversion, corresponding to the height and width directions respectively, Represents the Sigmoid function, which is a commonly used nonlinear activation function that compresses the input value to between 0 and 1. , Represents two 1x1 convolution operations, which are used to transform features. 1x1 convolution is usually used to adjust the number of channels of the feature map without changing its spatial dimension. , It represents the two parts after the input feature f is split, corresponding to the features in the height and width directions respectively. and Merged into a weight matrix for calculating the output of the CA module: ,in, Indicates the pixel value of the feature map output by the CA module at position (i, j), belonging to the cth channel, Represents the pixel value of the input feature map at position (i, j), belonging to the cth channel, Represents the weight value obtained after the height direction feature conversion, which is used to adjust the feature of the cth channel at position i. Represents the weight value obtained after width-wise feature conversion, which is used to adjust the features of the c-th channel at position j.
[0073] The channel attention mechanism is the SE attention mechanism, and the SE attention module can be determined based on the SER attention mechanism. The SE attention module integrates the relevant features between all feature map channel mappings through global pooling and fully connected layers to selectively emphasize the interdependent channel mappings and complete the recalibration of the feature map channel adaptive relationship. The main calculation method of SE attention is as follows: First, use global average pooling to extract the global information of each channel: ;in, represents the global information of the cth channel after global average pooling, represents the height of the feature map, represents the width of the feature map, U represents the feature map tensor, and c represents the channel index. Represents the pixel value of the cth channel at position (m,n) in the feature map tensor U. Then, a fully connected layer and activation function are used to generate the channel weight vector S: ;in, and They are the dimension reduction weights and raw dimension weights of the fully connected layer respectively; Represents the Sigmoid function; represents the linear rectification function, Represents global information, Represents a function that receives global information Z and weight parameters As input, the output is the result of processing by the fully connected layer and the activation function. Finally, the channel weight vector S is multiplied by U to obtain the recalibrated feature map; CA attention learns the spatial interdependence of features, SE attention simulates the interdependence of channels, and dual attention fuses the outputs of CA attention and SE attention to obtain a more accurate segmentation result by modeling rich contextual dependencies on local features and increasing the feature strength of each dimension while keeping the feature dimension unchanged. Figure 3 As shown, Figure 3 A dual attention mechanism based on a spatial attention mechanism and a channel attention mechanism is provided in an embodiment of the present invention. Figure 3The Input in it represents input, Residual represents residual, Avg Pool represents average pooling, FC represents fully connected layer, ReLU represents linear rectification function (RectifiedLinear Unit), Sigmoid represents activation function, Scale represents scaling, X Avg Pool represents average pooling in the X direction, Y Avg Pool represents average pooling in the Y direction, Concat+Conv2D represents concatenation + two-dimensional convolution, BatchNorm+Non-linear represents batch normalization + non-linear activation, Conv2D represents two-dimensional convolution, Re-weight represents re-weighting, and Sum Fusion represents sum fusion.
[0074] This embodiment does not limit the specific basic model. The basic model is a model pre-trained on a large sample data set. These models have learned rich feature representations and can capture common patterns in the data. The basic model can be a convolutional neural network (CNN), a recurrent neural network (RNN), YOLOV8 (a new model in the field of target detection), etc. It should be noted that since YOLOV8 itself belongs to the target detection model, the use of the YOLOV8 model can reduce the size of the small sample data set used for adjustment later, so that only fine-tuning can be used to obtain a distribution cabinet defect recognition model. The large sample data set in this embodiment generally refers to a data set containing a large number of samples (data points). The number of samples can range from thousands to millions or even more. The number of samples contained in the small sample data set is relatively small, and may only be tens to hundreds of samples. This embodiment does not limit the specific process of obtaining a pre-trained defect recognition model based on the spatial attention mechanism, the channel attention mechanism and the basic model. For example, this embodiment can first train the basic model based on a large sample data set of defects similar to the distribution cabinet (for example, a large sample data set of defects in substation control cabinets, a large sample data set of defects in controllers, etc.) to obtain a pre-trained basic model, and then add a spatial attention mechanism and a channel attention mechanism to the end of the pre-trained basic model to obtain an initial defect recognition model, and then fine-tune the initial defect recognition model to obtain a pre-trained defect recognition model, that is, migrate the model trained with a large sample data set to a small sample data set, and initialize it based on the weight parameters of the pre-trained model. Compared to training from random initialization, transfer learning utilizes the knowledge that has been learned by the pre-trained model, which can significantly improve the generalization ability of the model on small sample data and make the model converge faster and more efficiently. The parameters of the pre-trained basic model are migrated layer by layer, the model parameters and weights of the pre-trained basic model are loaded, and a complete model including a fully connected layer is retrained. Load the trained fully connected layer weights, and fine-tune certain key layers of the convolutional network or the entire model through methods such as layer freezing, such as Figure 4 As shown, Figure 4An example diagram of a process for fine-tuning a basic model provided in an embodiment of the present invention shows the training process of the YOLOv8 network model (basic model), which is divided into three stages: model pre-training, model fine-tuning, and model retraining. Each stage involves the processing of different network layers and weights. The following is an explanation of the process and letters in the figure: In the first stage, feature extraction training is performed using a large sample data set. After the training is completed, the weights of the feature extraction part are saved, but the fully connected layer is not included. In the second stage, some convolutional layers are frozen and only the remaining layers are fine-tuned. During the fine-tuning process, the weights of the feature extraction layer saved in the first stage are loaded to save the weights of the fully connected layer. Freeze some convolutional layers, train the pre-trained basic model using a small sample data set, add the spatial attention mechanism and the channel attention mechanism to the end of the pre-trained basic model, and obtain a pre-trained defect recognition model. Finally, a target defect detection model migrated from a large-scale data set and adapted to small sample data can be obtained. Alternatively, the implementation may directly add a spatial attention mechanism and a channel attention mechanism at the end of the basic model to obtain an initial defect recognition model, and then use a large sample data set to train the initial defect recognition model to obtain a pre-trained defect recognition model.
[0075] It should be further explained that in order to improve the efficiency of pre-trained defect recognition model training, before obtaining the pre-trained defect recognition model constructed based on the spatial attention mechanism, the channel attention mechanism and the basic model, it can also include: attaching the spatial attention module based on the spatial attention mechanism and the channel attention module based on the channel attention mechanism to the end of the basic model to obtain an initial defect recognition model; training the initial defect recognition model based on a large sample defect data set to obtain a pre-trained defect recognition model; wherein the large sample defect data set includes a distribution cabinet defect data set. This embodiment only needs to use a large sample defect data set to train the initial defect recognition model, without first training the basic model, and then re-fine-tuning the basic model after training with the addition of the spatial attention mechanism and the channel attention mechanism, so the efficiency of pre-trained defect recognition model training can be improved.
[0076] S103, using a small sample data set to train the pre-trained defect recognition model to obtain a target defect recognition model.
[0077] This embodiment uses a small sample data set to train the pre-trained defect recognition model, and fine-tunes the model parameters to obtain a target defect recognition model. It can be understood that the pre-trained defect recognition model is a model that uses a large sample data set (a data set of distribution cabinet defects in other fields) to train a basic model, and adds a spatial attention mechanism and a channel attention mechanism, or the pre-trained defect recognition model is a model obtained by training a basic model with a large sample data set (a data set of distribution cabinet defects in other fields) with an additional spatial attention mechanism and a channel attention mechanism. After training the model with a small sample data set, this embodiment can test the model with a test set (composed of a distribution cabinet defect image data set). When the test results show that the target defect recognition model meets the defect recognition capability requirements, it is determined to use the model for defect recognition.
[0078] It should be further explained that, in order to improve the efficiency of model training, the above-mentioned use of a small sample data set to train the pre-trained defect recognition model to obtain a target defect recognition model may include: freezing some convolutional layers of the pre-trained defect recognition model, and using a small sample data set to train the initialized defect recognition model to obtain a target defect recognition model; wherein some convolutional layers are convolutional layers whose parameters do not need to be adjusted. It is understandable that in the pre-trained defect recognition model, the parameters of some convolutional layers are frozen, that is, the parameters of these layers are no longer updated during the training process. These convolutional layers are usually located in the lower layers of the model and are responsible for extracting common features such as edges, textures, etc. Since some convolutional layers are frozen, the model can converge quickly on a small sample data set, thereby improving the efficiency of model training.
[0079] It should be further explained that in order to improve the model reasoning speed, after the pre-trained defect recognition model is trained with a small sample data set to obtain the target defect recognition model, it can also include: clipping the target defect recognition model based on the entropy value to obtain a lightweight target defect recognition model; wherein the entropy value clipping method is a method of determining the importance of the filter based on the entropy value, and clipping the filter based on the importance. This embodiment can determine the entropy value based on the probability of the pixel value in the feature map; or this embodiment can also determine the entropy value based on the probability of the bucket. This embodiment can clip filters with entropy values lower than the set value to obtain a lightweight target defect recognition model, so that the distribution cabinet defect recognition can be performed directly based on the lightweight target defect recognition model. Lightweight models usually consume fewer computing resources, thereby reducing energy consumption and improving model reasoning speed.
[0080] It should be further explained that, in order to improve the accuracy of clipping, the above entropy-based clipping method clips the target defect recognition model to obtain a lightweight target defect recognition model, which may include:
[0081] S1, divide each filter in the target defect recognition model into multiple buckets, determine the probability of each bucket, and determine the entropy value corresponding to the current filter through the probability of each bucket;
[0082] S2, determining whether the entropy value is greater than a set minimum entropy value threshold;
[0083] S3, when the entropy value of the filter is greater than the minimum entropy value threshold, determining not to trim the current filter;
[0084] S4: When the entropy value of the filter is not greater than the minimum entropy value threshold, determine to trim the current filter.
[0085] In the high-speed train distribution cabinet scenario, real-time and accurate target detection is crucial to ensure the safe operation of the distribution cabinet. This step uses the network pruning method to compress the model, such as Figure 5 As shown, Figure 5 An example flow chart of a network pruning method provided for an embodiment of the present invention includes S1-S4. Network pruning is usually performed after model training is completed. First, an evaluation criterion needs to be defined to measure the importance of each connection. Commonly used criteria include the size of the connection weight, etc. Then, according to the evaluation criteria, neurons or filters that are considered to be "unimportant" are deleted. However, it is difficult to judge the importance of the filter by the size of the weight value, and some useful filters may be pruned. Therefore, this step adopts an entropy-based pruning method, which uses entropy to determine the importance of the filter. Entropy is a common measure of disorder or uncertainty in information theory. A larger entropy value means that the system contains more information. In the filter pruning scenario, if the channel of the activation tensor contains less information, the corresponding filter is less important and can be discarded. The output of each layer is converted into a vector of length c (the number of filters) through a global average pooling. For n images, a For each filter, divide it into m bins, count the probability of each bin, and then calculate its entropy value. Use the entropy value to determine the importance of the filter, and then cut the unimportant filters. The entropy value of the jth filter is calculated as follows: ;in, yes The probability of is the entropy of channel j. Usually, if some layers are weak enough, for example, most of their activations are 0, then their entropy is relatively small. Therefore, our entropy-based method can be used to evaluate the importance of each channel. A smaller score of means that channel j is not that important in this layer and can therefore be removed. These weak filters are then removed from the original model altogether, resulting in a more compact network architecture. The corresponding channels of the filters in the next layer are also removed, achieving the purpose of compressing the model. Finally, the network needs to be fine-tuned to restore performance. Pruning and fine-tuning can be repeated until the best balance between model size and performance is achieved.
[0086] It should be further explained that in order to improve the performance of the lightweight target defect recognition model, the target defect recognition model is trimmed based on the entropy value to obtain a lightweight target defect recognition model, which may include: trimming the target defect recognition model based on the entropy value to obtain an initial lightweight target defect recognition model; fine-tuning the initial lightweight target defect recognition model to obtain a lightweight target defect recognition model, so that the output of the lightweight target defect recognition model is consistent with the output of the target defect recognition model. When fine-tuning this embodiment, a small sample data set determined based on the distribution cabinet defect image data set can be used for fine-tuning. In this embodiment, after the model is trimmed, the model will be fine-tuned so that the performance of the model can be maintained.
[0087] The defect recognition model training method provided by an embodiment of the present invention may include: S101, marking the defects in each distribution cabinet defect image in a distribution cabinet defect image data set to obtain a small sample data set; wherein the small sample data set is a data set with a data volume less than a first data volume threshold; S102, obtaining a pre-trained defect recognition model constructed based on a spatial attention mechanism, a channel attention mechanism and a basic model; wherein the basic model is a model obtained by training based on a large sample data set; the large sample data set is a data set with a data volume greater than a second data volume threshold; S103, using the small sample data set to train the pre-trained defect recognition model to obtain a target defect recognition model. Compared with the current situation where the background is relatively complex and the target is relatively small, and the existing intelligent recognition algorithms have insufficient sample size during training, resulting in relatively low diagnostic accuracy, the present application uses a small sample data set to train a pre-trained defect recognition model based on the spatial attention mechanism, the channel attention mechanism and the basic model, so that the target defect recognition model can simulate the semantic interdependence in the spatial and channel dimensions based on the trained spatial attention mechanism and the channel attention mechanism, respectively, to solve the problem that the existing intelligent recognition algorithms are difficult to capture the full picture of data distribution in small sample multi-defect recognition tasks, thereby improving the accuracy of distribution cabinet defect recognition.
[0088] The complex scene environment and massive equipment types of the distribution cabinet make it difficult to directly train the monitoring model, such as insufficient sample size and difficulty in migration, which will seriously restrict the development of intelligent monitoring technology for electrical equipment. At present, the mainstream methods for small sample target detection include meta-learning, data enhancement, and self-supervised learning, but the existing intelligent recognition algorithms will inevitably suffer from performance degradation in small sample multi-defect recognition tasks, poor generalization ability, and difficulty in capturing the full picture of data distribution. They are often accompanied by huge parameters to adapt to different tasks, with high computational costs. In resource-constrained scenarios such as high-speed train distribution cabinets, it is difficult to achieve online real-time detection.
[0089] In order to make the present invention easier to understand, please refer to Figure 6 , Figure 6 A flowchart of a defect identification method provided by an embodiment of the present invention may specifically include:
[0090] S601, obtaining an image of a power distribution cabinet of a high-speed train to be inspected.
[0091] The distribution cabinet image in this embodiment is an image of a high-speed train distribution cabinet. A distribution cabinet is a device widely used in power systems, mainly used to manage and allocate electric energy to ensure normal, stable and safe power supply.
[0092] S602, using a target defect recognition model to perform defect recognition on the image of the high-speed train distribution cabinet to be inspected, and determine the fault type of the distribution cabinet; wherein the target defect recognition model is a model obtained based on the above-mentioned defect recognition model training method.
[0093] In addition to using the target defect recognition model to determine the type of distribution cabinet fault, this embodiment can also determine the fault location by simply marking the location of each fault during the training process.
[0094] The target defect recognition model in the embodiment of the present invention is a model obtained based on the above-mentioned defect recognition model training method. Since a model with a dual attention module based on CA attention and SE attention is designed, the model can simulate semantic interdependence in spatial and channel dimensions, and improve the characterization effect of target features in application scenarios. Moreover, a large sample data set will be used for training first, so the generalization ability of the model on small sample data can be improved. Therefore, the defect recognition method provided by the present invention is more accurate.
[0095] In order to make the present invention easier to understand, please refer to Figure 7 , Figure 7 An example flow chart of defect recognition model training and use provided in an embodiment of the present invention may specifically include:
[0096] S701. Collect and generate target defect images of the high-speed train power distribution cabinet during operation. Among them, the target defect images include six types: terminal looseness, terminal burning and discoloration, wire protection sleeve loosening, terminal fracture, abnormal position of the motor starter, and missing resistor plate.
[0097] For easy understanding, please refer to Figure 8 , Figure 8 which is a framework schematic diagram for training and using a defect recognition model provided by an embodiment of the present invention. It can be seen from this figure that as a whole, it mainly includes processing the data set to obtain a training set and a test set (data collection, generating a sample set using data augmentation techniques; labeling tags, dividing the training set and the test set), training the BAYoLo model based on the training set, using the test set to determine the performance of the trained model, determining whether the classification ability is optimal. If not, fine-tuning is performed. If it is already optimal, model pruning is performed and then fine-tuning is carried out to obtain the final model (target defect recognition model), so as to perform online detection on the power distribution cabinet image to be detected based on the target defect recognition model.
[0098] S702. Label the target defects in the target defect images to obtain a small sample data set.
[0099] S703. Preprocess the images in the small sample data set to meet the input requirements of the model, and crop the regions of interest in the images to remove irrelevant information in the pictures, obtaining a preprocessed small sample data set.
[0100] S704. Perform a normalization operation on the images in the preprocessed small sample data set to obtain a normalized small sample data set.
[0101] S705. Perform data augmentation on the normalized small sample data set to obtain an expanded small sample data set. Among them, the data augmentation methods include at least one of rotation, scaling, translation, color transformation, adding noise, cropping, and flipping.
[0102] S706. Attach SE attention and CA attention to the Yolov8 model to obtain an initial defect recognition model.
[0103] The Yolov8 model in this embodiment corresponds to the basic model above. The initial defect recognition model in this embodiment is the BAYoLo model.
[0104] S707. Train the initial defect recognition model using a large sample data set to obtain a pre-trained defect recognition model.
[0105] S708. Train the pre-trained defect recognition model using the small sample data set to obtain a target defect recognition model.
[0106] S709, online acquisition of the image of the power distribution cabinet of the high-speed train to be detected, and processing of the image of the power distribution cabinet of the high-speed train to be detected to obtain the target image to be detected.
[0107] S710, using the target defect recognition model to identify defects in the target image to be detected, and obtain defect categories.
[0108] It should be further explained that this embodiment can develop an intelligent warning system for power distribution cabinets based on infrared thermal imaging, deploy image processing algorithms and defect recognition algorithms to mobile terminals, and conduct secondary development of visible light cameras and infrared cameras. The system modules include an infrared image real-time preview module, a function selection module, and a high-speed rail power distribution cabinet typical abnormality diagnosis module, providing a simple interactive environment. The algorithm development and model training are completed on the computer side, and the trained BAYolo model is deployed to the mobile terminal using TensorFlowLite technology, and experimental verification is carried out on the high-speed train power distribution cabinet. For ease of understanding, please refer to Fig. 9 , Fig. 9 A schematic diagram of a defect recognition framework provided by an embodiment of the present invention, from Fig. 9 It can be seen that the defect recognition framework mainly includes data set construction, model construction, defect recognition and alarm system, among which data set construction, model construction and defect recognition correspond to the defect recognition model training method and defect recognition method mentioned above.
[0109] The beneficial effects of the present invention are: (1) Existing intelligent recognition algorithms will inevitably experience performance degradation problems under small sample multi-defect recognition tasks, with poor generalization ability and difficulty in capturing the full picture of data distribution. The present invention pre-trains and fine-tunes the Yolov8 model on a large-scale data set to improve the generalization ability of the model on small sample data; designs a dual attention module to simulate semantic interdependence in spatial and channel dimensions, and improves the characterization effect of target features in application scenarios; (2) The pre-trained model based on transfer learning has a large number of parameters and cannot be deployed on the resource-constrained edge, making it difficult to achieve online real-time detection. Using the idea of model compression, an entropy-based evaluation criterion is designed to measure the importance of each connection, and neurons or filters that are considered "unimportant" are deleted, so as to compress the size of the defect recognition model while maintaining its detection accuracy and improving the model reasoning speed.
[0110] The defect recognition model training device provided in an embodiment of the present invention is introduced below. The defect recognition model training device described below and the defect recognition model training method described above can be referenced to each other.
[0111] Please refer to Fig.10 , Fig.10 A structural diagram of a defect recognition model training device provided by an embodiment of the present invention may include:
[0112] A small sample data set determination module 100 is used to mark defects in each power distribution cabinet defect image in the power distribution cabinet defect image data set to obtain a small sample data set; wherein the small sample data set is a data set whose data volume is less than a first data volume threshold;
[0113] A pre-trained defect recognition model determination module 200 is used to obtain a pre-trained defect recognition model constructed based on a spatial attention mechanism, a channel attention mechanism and a basic model; wherein the basic model is a model obtained by training based on a large sample data set; and the large sample data set is a data set whose data volume is greater than a second data volume threshold;
[0114] The target defect model determination module 300 is used to train the pre-trained defect recognition model using the small sample data set to obtain a target defect recognition model.
[0115] Further, based on the above embodiment, the small sample data set determination module 100 may include:
[0116] A data set construction unit, used to construct the distribution cabinet defect image data set based on the distribution cabinet defect image; wherein the distribution cabinet defect image is a high-speed train distribution cabinet defect image;
[0117] A cropping unit, used to crop an area of interest in each power distribution cabinet defect image in the power distribution cabinet defect image dataset, remove information irrelevant to defect detection, and obtain a cropped power distribution cabinet defect image dataset;
[0118] A normalization unit is used to perform a normalization operation on the images in the cropped distribution cabinet defect image data set to obtain the small sample data set.
[0119] Further, based on any of the above embodiments, the above defect recognition model training device may further include:
[0120] An initial defect recognition model determination module, used for attaching a spatial attention module based on the spatial attention mechanism and a channel attention module based on the channel attention mechanism to the end of the basic model to obtain an initial defect recognition model;
[0121] A pre-trained defect recognition model determination module is used to train the initial defect recognition model based on a large sample defect data set to obtain the pre-trained defect recognition model; wherein the large sample defect data set includes a distribution cabinet defect data set.
[0122] Further, based on any of the above embodiments, the target defect model determination module 300 may include:
[0123] A fine-tuning unit is used to freeze part of the convolutional layers of the pre-trained defect recognition model and train the initialized defect recognition model using the small sample data set to obtain the target defect recognition model; wherein the part of the convolutional layers is a convolutional layer whose parameters do not need to be adjusted.
[0124] Further, based on any of the above embodiments, the above defect recognition model training device may further include:
[0125] A lightweight processing module is used to trim the target defect recognition model based on an entropy value trimming method to obtain a lightweight target defect recognition model; wherein the entropy value trimming method is a method of determining the importance of the filter based on the entropy value and trimming the filter based on the importance.
[0126] Further, based on the above embodiment, the above lightweight processing module may include:
[0127] An entropy value determination unit, used to divide each filter in the target defect recognition model into a plurality of buckets, determine the probability of each bucket, and determine the entropy value corresponding to the current filter through the probability of each bucket;
[0128] A judgment unit, used to determine whether the entropy value is greater than a set minimum entropy value threshold;
[0129] A clipping determination unit, configured to determine not to clip the current filter when the entropy value of the filter is greater than the minimum entropy value threshold;
[0130] The non-pruning determination unit is used to determine to prune the current filter when the entropy value of the filter is not greater than the minimum entropy value threshold.
[0131] Further, based on the above embodiment, the above lightweight processing module may include:
[0132] An initial lightweight unit, used for cutting the target defect recognition model based on the cutting method of the entropy value to obtain an initial lightweight target defect recognition model;
[0133] The lightweight model fine-tuning unit is used to fine-tune the initial lightweight target defect recognition model to obtain the lightweight target defect recognition model, so that the output of the lightweight target defect recognition model is consistent with the output of the target defect recognition model.
[0134] It should be noted that the order of the modules and units in the above-mentioned defect recognition model training device can be changed without affecting the logic.
[0135] The defect recognition model training device provided by an embodiment of the present invention may include: a small sample data set determination module 100, used to mark the defects in each distribution cabinet defect image in the distribution cabinet defect image data set to obtain a small sample data set; wherein the small sample data set is a data set with a data volume less than a first data volume threshold; a pre-trained defect recognition model determination module 200, used to obtain a pre-trained defect recognition model constructed based on a spatial attention mechanism, a channel attention mechanism and a basic model; wherein the basic model is a model obtained by training based on a large sample data set; the large sample data set is a data set with a data volume greater than a second data volume threshold; a target defect model determination module 300, used to train the pre-trained defect recognition model using the small sample data set to obtain a target defect recognition model. Compared with the current situation where the background is relatively complex and the target is relatively small, and the existing intelligent recognition algorithms have insufficient sample size during training, resulting in relatively low diagnostic accuracy, the present application uses a small sample data set to train a pre-trained defect recognition model based on the spatial attention mechanism, the channel attention mechanism and the basic model, so that the target defect recognition model can simulate the semantic interdependence in the spatial and channel dimensions based on the trained spatial attention mechanism and the channel attention mechanism, respectively, to solve the problem that the existing intelligent recognition algorithms are difficult to capture the full picture of data distribution in small sample multi-defect recognition tasks, thereby improving the accuracy of distribution cabinet defect recognition.
[0136] The defect recognition device provided by an embodiment of the present invention is introduced below. The defect recognition device described below and the defect recognition method described above can be referenced to each other.
[0137] Please refer to Fig.11 , Fig.11 A structural diagram of a defect identification device provided by an embodiment of the present invention may include:
[0138] The data acquisition module 400 is used to acquire the image of the power distribution cabinet of the high-speed train to be detected;
[0139] The defect recognition module 500 is used to use a target defect recognition model to perform defect recognition on the image of the high-speed train distribution cabinet to be detected and determine the fault type of the distribution cabinet; wherein the target defect recognition model is a model obtained based on the above-mentioned defect recognition model training method.
[0140] Further, based on any of the above embodiments, the defect identification module 500 may include:
[0141] The defect recognition unit is used to use the target defect recognition model to perform defect recognition on the image of the high-speed train distribution cabinet to be detected, and determine the fault type and fault location of the distribution cabinet.
[0142] It should be noted that the order of the modules and units in the above-mentioned defect identification device can be changed without affecting the logic.
[0143] The defect recognition device provided in the embodiment of the present invention may include: a data acquisition module 400, which is used to acquire an image of a high-speed train distribution cabinet to be detected; a defect recognition module 500, which is used to use a target defect recognition model to perform defect recognition on the image of the high-speed train distribution cabinet to be detected, and determine the type of fault in the distribution cabinet; wherein the target defect recognition model is a model obtained based on the above-mentioned defect recognition model training method. The target defect recognition model in the embodiment of the present invention is a model obtained based on the above-mentioned defect recognition model training method. Since a model with dual attention modules designed based on CA attention and SE attention is designed, the model can simulate semantic interdependence in spatial and channel dimensions, improve the characterization effect of target features in application scenarios, and will first use a large sample data set for training, so the generalization ability of the model on small sample data can be improved. Therefore, the defect recognition method provided by the present invention has higher accuracy.
[0144] An electronic device provided by an embodiment of the present invention is introduced below. The electronic device described below and the steps of the defect recognition model training method and the defect recognition method described above can be referenced to each other.
[0145] Please refer to Fig.12 , Fig.12 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention may include:
[0146] A memory 10, used for storing computer programs;
[0147] The processor 20 is used to execute a computer program to implement the above-mentioned defect recognition model training method and defect recognition method.
[0148] The memory 10 , the processor 20 , and the communication interface 30 all communicate with each other via a communication bus 40 .
[0149] In an embodiment of the present invention, the memory 10 is used to store one or more programs, the program may include program code, the program code includes computer operation instructions, and in an embodiment of the present invention, the memory 10 may store steps for implementing a defect recognition model training method and a defect recognition method.
[0150] In a possible implementation, the memory 10 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function, etc.; the data storage area may store data created during use.
[0151] In addition, the memory 10 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include an NVRAM. The memory stores an operating system and operating instructions, executable modules or data structures, or a subset thereof, or an extended set thereof, wherein the operating instructions may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic tasks and processing hardware-based tasks.
[0152] The processor 20 may be a central processing unit (CPU), an application specific integrated circuit, a digital signal processor, a field programmable gate array or other programmable logic device, a microprocessor or any conventional processor, etc. The processor 20 may call a program stored in the memory 10 .
[0153] The communication interface 30 may be an interface of a communication module, and is used to connect to other devices or systems.
[0154] Of course, it should be noted that Fig.12 The structure shown does not constitute a limitation on the electronic device in the embodiment of the present invention. In actual applications, the electronic device may include Fig.12 More or fewer components than shown, or combinations of certain components.
[0155] The computer-readable storage medium provided in an embodiment of the present invention is introduced below. The computer-readable storage medium described below and the defect recognition model training method and defect recognition method described above can be referenced to each other.
[0156] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned defect recognition model training method and defect recognition method steps are implemented.
[0157] The computer-readable storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0158] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0159] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0160] Finally, it should be noted that, in this article, relationships such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variations are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0161] The above is a detailed introduction to a defect recognition model training method, defect recognition method, device and equipment provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A defect recognition model training method, characterized in that: include: Annotate the defects in each power distribution cabinet defect image in the power distribution cabinet defect image dataset to obtain a small sample dataset; wherein the small sample dataset is a dataset whose data volume is less than a first data volume threshold; Obtain a pre-trained defect recognition model constructed based on a spatial attention mechanism, a channel attention mechanism, and a basic model; wherein the basic model is a model trained based on a large sample data set; the large sample data set is a data set whose data volume is greater than a second data volume threshold; The pre-trained defect recognition model is trained using the small sample data set to obtain a target defect recognition model.
2. The defect recognition model training method according to claim 1, characterized in that: The defects in each distribution cabinet defect image in the distribution cabinet defect image dataset are annotated to obtain a small sample dataset, including: Constructing the distribution cabinet defect image dataset based on the distribution cabinet defect image; wherein the distribution cabinet defect image is a high-speed train distribution cabinet defect image; Cropping an area of interest in each power distribution cabinet defect image in the power distribution cabinet defect image dataset to remove information irrelevant to defect detection, thereby obtaining a cropped power distribution cabinet defect image dataset; A normalization operation is performed on the images in the cropped distribution cabinet defect image dataset to obtain the small sample dataset.
3. The defect recognition model training method according to claim 1, characterized in that: Before obtaining the pre-trained defect recognition model based on the spatial attention mechanism, channel attention mechanism and basic model, it also includes: Adding a spatial attention module based on the spatial attention mechanism and a channel attention module based on the channel attention mechanism to the end of the basic model to obtain an initial defect recognition model; The initial defect recognition model is trained based on a large sample defect data set to obtain the pre-trained defect recognition model; wherein the large sample defect data set includes a distribution cabinet defect data set.
4. The defect recognition model training method according to any one of claims 1 to 3, characterized in that: The pre-trained defect recognition model is trained using the small sample data set to obtain a target defect recognition model, including: The target defect recognition model is obtained by freezing some convolutional layers of the pre-trained defect recognition model and using the small sample data set to train the initialized defect recognition model; wherein the some convolutional layers are convolutional layers whose parameters do not need to be adjusted.
5. The defect recognition model training method according to any one of claims 1 to 3, characterized in that: After the pre-trained defect recognition model is trained using the small sample data set to obtain a target defect recognition model, the method further includes: The target defect recognition model is trimmed based on the entropy value to obtain a lightweight target defect recognition model; wherein the entropy value trimming method is a method of determining the importance of the filter based on the entropy value and trimming the filter based on the importance.
6. The defect recognition model training method according to claim 5, characterized in that: The entropy-based clipping method clips the target defect recognition model to obtain a lightweight target defect recognition model, including: Divide each filter in the target defect recognition model into multiple buckets, determine the probability of each bucket, and determine the entropy value corresponding to the current filter through the probability of each bucket; Determine whether the entropy value is greater than a set minimum entropy value threshold; When the entropy value of the filter is greater than the minimum entropy value threshold, determining not to trim the current filter; When the entropy value of the filter is not greater than the minimum entropy value threshold, it is determined to trim the current filter.
7. The defect recognition model training method according to claim 5, characterized in that: The target defect recognition model is trimmed based on the entropy value to obtain a lightweight target defect recognition model, including: The target defect recognition model is trimmed based on the trimming method of the entropy value to obtain an initial lightweight target defect recognition model; The initial lightweight target defect recognition model is fine-tuned to obtain the lightweight target defect recognition model, so that the output of the lightweight target defect recognition model is consistent with the output of the target defect recognition model.
8. A defect identification method, characterized in that: include: Acquire the image of the power distribution cabinet of the high-speed train to be inspected; The target defect recognition model is used to perform defect recognition on the image of the high-speed train distribution cabinet to be detected to determine the fault type of the distribution cabinet; wherein the target defect recognition model is a model obtained based on the defect recognition model training method described in any one of claims 1 to 7.
9. The defect identification method according to claim 8, characterized in that: The method of using the target defect recognition model to perform defect recognition on the image of the high-speed train power distribution cabinet to be detected and determining the fault type of the power distribution cabinet includes: The target defect recognition model is used to perform defect recognition on the image of the high-speed train power distribution cabinet to be detected, and the fault type and fault location of the power distribution cabinet are determined.
10. A defect recognition model training device, characterized in that: include: A small sample data set determination module is used to mark the defects in each power distribution cabinet defect image in the power distribution cabinet defect image data set to obtain a small sample data set; wherein the small sample data set is a data set whose data volume is less than a first data volume threshold; A pre-trained defect recognition model determination module is used to obtain a pre-trained defect recognition model constructed based on a spatial attention mechanism, a channel attention mechanism and a basic model; wherein the basic model is a model obtained by training based on a large sample data set; the large sample data set is a data set whose data volume is greater than a second data volume threshold; The target defect model determination module is used to train the pre-trained defect recognition model using the small sample data set to obtain a target defect recognition model.
11. A defect identification device, characterized in that: include: A data acquisition module, used to acquire an image of a power distribution cabinet of a high-speed train to be inspected; A defect recognition module is used to use a target defect recognition model to perform defect recognition on the image of the high-speed train distribution cabinet to be detected and determine the fault type of the distribution cabinet; wherein the target defect recognition model is a model obtained based on the defect recognition model training method described in any one of claims 1 to 7.
12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, used to execute the computer program to implement the steps of the defect recognition model training method according to any one of claims 1 to 7, and the steps of the defect recognition method according to claim 8 or 9.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the defect recognition model training method according to any one of claims 1 to 7, and the steps of the defect recognition method according to claim 8 or 9.