A Method for Detecting Damage of Steel Floors of Railway Freight Cars Based on Knowledge Distillation

By applying knowledge distillation from a ResNet101 teacher model to a ResNet18 student model, the method addresses the resource constraints of Transformer-based target detection, achieving efficient and accurate steel deck damage detection on resource-limited devices.

CN116703819BActive Publication Date: 2025-07-15SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310399454.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-14
Publication Date
2025-07-15
Estimated Expiration
2043-04-14

AI Technical Summary

Technical Problem

The object detection method model based on Transformer network is relatively large and cannot be deployed on terminal devices with limited storage space and computing resources.

Method used

The knowledge distillation technology is introduced to pass the knowledge of the teacher model to the student model layer by layer. Through progressive multi-level knowledge distillation and teacher feature distillation, a railway truck steel floor damage detection method is constructed based on Transformer, and ResNet101 is used as the teacher model and ResNet18 is used as the student model for training and knowledge transmission.

Benefits of technology

It realizes efficient deployment of railway truck steel floor damage detection model on resource-constrained terminal equipment, improving detection accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116703819B_ABST
    Figure CN116703819B_ABST
Patent Text Reader

Abstract

The present invention provides a method for detecting damage to the steel floor of railway freight cars based on knowledge distillation, which includes the following steps: obtaining images of the steel floor area of railway freight cars to construct a training set; building a teacher network and a student network for steel floor damage detection and training them, using the teacher network to distill the student network, and obtaining the final fault detection model by adjusting parameters; obtaining the image to be detected, processing it and inputting it into the fault detection model to obtain the detection result of steel floor damage. Based on deep convolutional neural network and knowledge distillation, it adopts the structure of an encoder and a decoder, and establishes a prediction match between queries through progressive multi-level knowledge distillation to gradually transfer useful knowledge to the student model. The present invention provides an automated detection method with high accuracy and precision, which solves the problem of false detection and missed detection caused by visual fatigue due to the fact that at the current stage, faults can only be identified by dynamic car inspectors through visual inspection of images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of object detection in computer vision, and particularly relates to a method for detecting damage to the steel floor of railway wagons based on knowledge distillation. Background Art

[0002] In recent years, the continuous complexity of deep convolutional neural network models, with the continuous expansion of parameters and computational volume, has brought huge challenges to computing resources and storage resources. To solve the problem of difficult deployment on resource-constrained devices such as embedded devices, some neural network model compression technologies are adopted to reduce the volume and computational volume of the model, so as to achieve efficient deployment of deep learning models in resource-limited environments.

[0003] Knowledge distillation is a commonly used model compression method and is applied to the field of image classification. Generally, a complex model with excellent performance is trained as a teacher model, and then the knowledge learned by the teacher model is used to guide the training of a simpler student model. Eventually, the performance of the student model is comparable to that of the teacher model, but the number of network parameters and complexity are greatly reduced, thus achieving model compression and acceleration.

[0004] Although knowledge distillation has achieved good results in traditional object detection methods based on convolutional neural networks, due to the different architectures of convolutional neural networks and Transomers, traditional algorithms are difficult to be directly applied to object detection methods based on Transformer. This is because in object detection methods based on convolutional neural networks, object information is carried by image feature maps, while in object detection methods based on Transformer, object information is mainly encoded in query vectors. This huge difference leads to a significant difference in the feature distribution of object information in these two detection methods. Summary of the Invention

[0005] Object of the Invention: To solve the problem that the object detection method model based on the Transformer network is relatively large and cannot be deployed to terminals with limited storage space and computing resources, the present invention introduces knowledge distillation into the object detection method based on Transformer, and proposes a method for detecting damage to the steel floor of railway wagons based on knowledge distillation, including the following steps:

[0006] Step 1, obtain images of multiple angles at the bottom of the train;

[0007] Step 2, select the pictures containing damage to the steel floor and the pictures without damage, label the damaged parts, and distinguish the samples without faults from a small number of samples with faults;

[0008] Step 3, construct a fault detection model, including a teacher model and a student model;

[0009] Step 4: Use the train images of the train bottom obtained in Step 1 to train the teacher model and the student model in parallel. Transfer the knowledge of the teacher model to the student model using progressive multi-level knowledge distillation and teacher feature distillation. The teacher model transfers the dark knowledge of the decoder layer in the teacher model to the student model layer by layer. According to the knowledge distillation loss function, train the student model to achieve knowledge distillation, and finally obtain a trained fault detection model.

[0010] Step 5: Verify the model detection result. Obtain the image to be detected, input it into the fault detection model, calculate the anomaly score, and obtain the detection result of the steel floor damage.

[0011] Furthermore, the teacher model and the student model have the same structure, both including a backbone network module, an encoder module, a decoder module, and a prediction output module. However, the teacher model and the student model use backbone network modules of different sizes. After the image passes through the backbone network module, high-dimensional vector information is extracted and sent to the encoder module. The encoder module performs semantic encoding on the features and then sends them to the decoder module. The decoder performs cross-attention on the key values of the feature map and the corresponding regional features, and finally outputs the final detection result through the prediction module.

[0012] The backbone network module includes an input layer, a first group of convolutional layers, a max pooling layer, a second group of convolutional layers, a third group of convolutional layers, a fourth group of convolutional layers, and a fifth group of convolutional layers connected in sequence. Among them, the input size of the input layer is image data of 513x513. The first group of convolutional layers includes two parts: a 7x7 convolutional operation and a non-linear activation function operation in sequence. The second group of convolutional layers includes 9 convolutional layers, a non-linear activation layer, and an average pooling layer in sequence. The third group of convolutional layers includes 12 convolutional layers, a non-linear activation layer, and an average pooling layer in sequence. The fourth group of convolutional layers includes 69 convolutional layers, a non-linear activation layer, and an average pooling layer in sequence. The fifth group of convolutional layers includes 9 convolutional layers, a non-linear activation layer, and an average pooling layer in sequence. Among them, each convolutional layer in the second group to the fifth group of convolutional layers passes through a 1x1 convolution, a 3x3 convolution, and a 1x1 convolution operation in sequence.

[0013] The encoder module includes six encoders. The encoder is a Transformer encoder. The encoder module is stacked by 6 identical encoders. Each encoder has two sub-layers. The first sub-layer is a multi-head self-attention pooling layer. The second sub-layer is a position-based feed-forward neural network layer. Each sub-layer uses a residual connection. The encoder adds the serialized feature map and the position encoding to obtain the query Q and the key value K. After passing through a multi-head self-attention layer, it is added to the feature map and normalized, and then passed through a feed-forward network to obtain the output of a single encoder, which is used as the input of the next encoder. After passing through 6 identical encoder structures, the output of the encoder part is obtained.

[0014] Decoder module: The input of the decoder consists of three parts: the output of the encoder, the positional encoding, and the query. Among them, the dimension of the query is (300, 4), where the first dimension is the number of predefined target queries, and the second dimension is the number of hidden layers. The decoder calculates the first half of the multi-head self-attention in the same way as the encoder, and then calculates the cross-attention with the output of the encoder. After passing through 6 identical decoder structures, the output of the decoder part is obtained.

[0015] Prediction output module: It includes a feed-forward neural network and a fully connected layer. The feed-forward neural network layer is mainly divided into two parts. One part predicts the category, and the other part predicts the position. The branch for predicting the category consists of a linear layer with a hidden layer dimension of 512. Since there is a background class (empty class), the output dimension is the number of classes plus 1. The other branch of the feed-forward neural network for predicting the position mainly consists of 3 linear layers with a hidden layer dimension of 512. Both of these branches will pass through a sigmoid activation function.

[0016] Furthermore, the backbone network of the teacher model is ResNet101, and the backbone network of the student model is ResNet18.

[0017] Furthermore, the loss function of the student model in step 4 is a composite loss function composed of the teacher soft label and the student hard label:

[0018]

[0019] where α and β are hyperparameters, is the teacher soft label loss, is the student hard label loss.

[0020] Furthermore, through positive and negative sample anchor box query sampling, during the distillation training, the selected distillation anchor boxes do not participate in the backpropagation of the student model. The loss functions of the positive and negative sample anchor boxes and the randomly sampled anchor boxes are

[0021] where, is the positive sample anchor box distillation loss, is the negative sample anchor box distillation loss, is the random anchor box distillation loss.

[0022] Beneficial effects: The method for detecting the damage of the steel floor of railway wagons based on knowledge distillation proposed by the present invention mainly consists of two parts: a teacher network model and a student network model. The backbone network of the teacher network model is ResNet101, and the backbone network of the student network model is ResNet18. Progressive multi-level knowledge distillation is used to establish predictive matches between queries to gradually transfer useful knowledge to the student model; at the same time, the method of teacher feature distillation is adopted to make full use of the intermediate features of the teacher to provide additional information for the one-to-one assignment strategy group in the student model. Description of the Drawings

[0023] Figure 1 It is a schematic flow chart of the method for detecting the damage of the steel floor of railway wagons based on knowledge distillation of the present invention.

[0024] Figure 2 It is the network structure diagram and detailed module diagram of the present invention.

[0025] Figure 3 It is the structure diagram of teacher model feature distillation.

[0026] Figure 4 It is the comparison chart of the map indexes of the improved algorithm and the student model algorithm of the present invention. Detailed Embodiment

[0027] As Figure 1 shown, a method for detecting the damage of the steel floor of railway wagons based on knowledge distillation of the present invention includes the following steps:

[0028] Step 1: Obtain images at multiple angles of the bottom of the train.

[0029] First, obtain the whole vehicle image during the operation of the railway wagon through a high-speed camera, including the side frame, the middle part, and the coupler buffer part, select the image of the bottom containing the steel floor from them, and detect the obtained image of the bottom.

[0030] Step 2: Screen the obtained pictures, and retain the pictures containing the target parts; screen the pictures, select the pictures with damaged steel floors and the pictures without damage in a quantity ratio of 1:1, mark the damaged parts, and distinguish the samples without faults from a small number of samples with faults;

[0031] Step 3: Construct a teacher model and a student model;

[0032] The encoder module structures of the teacher model and the student model are the same, and both are encoder modules composed of 6-layer Transformer encoders.

[0033] As Figure 2As shown, both the teacher model and the student model include a backbone network module, an encoder module, a decoder module, and a prediction output module. After the image passes through the backbone network module, high-dimensional vector information is extracted and sent to the encoder module. The encoder module performs semantic encoding on the features and then sends them to the decoder module. The decoder performs cross-attention on the key-value of the feature map and the corresponding regional features, and finally outputs the final detection result through the prediction module.

[0034] The backbone network module includes an input layer, a first group of convolutional layers, a max pooling layer, a second group of convolutional layers, a third group of convolutional layers, a fourth group of convolutional layers, and a fifth group of convolutional layers connected in sequence.

[0035] Among them, the input size of the input layer is image data of 513x513; the first group of convolutional layers includes 1 7x7 convolutional operation and 1 non-linear activation function operation in sequence; the second group of convolutional layers includes 9 convolutional layers, a non-linear activation layer, and an average pooling layer in sequence; the third group of convolutional layers includes 12 convolutional layers, a non-linear activation layer, and an average pooling layer in sequence; the fourth group of convolutional layers includes 69 convolutional layers, a non-linear activation layer, and an average pooling layer in sequence; the fifth group of convolutional layers includes 9 convolutional layers, a non-linear activation layer, and an average pooling layer in sequence. Among them, each convolutional layer in the second group to the fifth group of convolutional layers undergoes 1 1x1 convolution, 1 3x3 convolution, and 1 1x1 convolution operation in sequence.

[0036] Encoder module: The encoder module includes six encoders, and the encoder is a Transformer encoder. The encoder module is composed of 6 identical encoders stacked together. Each encoder has two sub-layers. The first sub-layer is a multi-head self-attention pooling layer; the second sub-layer is a position-based feed-forward neural network layer, and each sub-layer uses a residual link. The encoder adds the serialized feature map and the position encoding to obtain the query Q and the key-value K, adds them to the feature map after passing through a multi-head self-attention layer and normalizes them, and then obtains the output of a single encoder through a feed-forward network and uses it as the input of the next encoder. After passing through 6 identical encoder structures, the output of the encoder part is obtained.

[0037] Decoder module: The input of the decoder consists of three parts: the output of the encoder, the position encoding, and the query. Among them, the dimension of the query is (300,4), the first dimension is the number of predefined target queries, and the second dimension is the number of hidden layers. That is, according to the features encoded by the encoder, the decoder converts 300 queries into 300 targets. The first half of the decoder calculates the multi-head self-attention in the same way as the encoder operation, and then calculates the cross-attention with the output of the encoder. After passing through 6 identical decoder structures, the output of the decoder part is obtained.

[0038] Prediction Output Module: It continues to calculate based on the features output by the decoder and is mainly composed of a feed-forward neural network and a fully connected layer. The feed-forward neural network layer is mainly divided into two parts. One part predicts the category, and the other part predicts the position. The branch for predicting the category consists of a linear layer with a hidden layer dimension of 512. Since there is a background class (empty class), the output dimension is the number of classes plus 1. The other branch of the feed-forward neural network for predicting the position mainly consists of three linear layers with a hidden layer dimension of 512. Both of these branches will go through a sigmoid activation function.

[0039] Step 3, as Figure 3 shown, use the train bottom images obtained in Step 1 to train the teacher model and the student model in parallel. Transfer the knowledge of the teacher model to the student model using progressive multi-level knowledge distillation and teacher feature distillation. Use the Hungarian algorithm to perform one-to-one matching on the forward inference results of the teacher model and the student model, and layer by layer transfer the dark knowledge of the decoder layer in the teacher model to the student model. At the same time, adopt the method of teacher feature distillation to provide the probability information before model normalization for the student model.

[0040] Input the input data into the teacher model, layer by layer transfer the dark knowledge of the decoder layer in the teacher model to the student model, and train the student model according to the knowledge distillation loss function to achieve knowledge distillation. The implementation process is as follows: The backbone network of the teacher model is ResNet101, and the backbone network of the student model is ResNet18. Both are composed of five layers of convolutional layer residuals connected in series. The difference is that there is only downsampling operation in the convolutional layer of ResNet18, and the size of its convolutional kernel is 3x3. The output feature depth of the fifth convolutional layer is 512; the convolutional layer of ResNet101 has both upsampling and downsampling operations, and the output feature depth of the fifth convolutional layer is 2048. Both the teacher model and the student model will output the feature maps from the third layer to the fifth layer, and downsample to 256 when inputting the encoder part as the hidden layer dimension. The encoder structures of the teacher model and the student model are the same, both are Transformer encoder structures.

[0041] The query distillation anchor box in the decoder part is a combination of the image and the query. Since the query has the function of detecting and aggregating some instance features, its distribution in different backbone models may be inconsistent. Therefore, select the same number of similar positive sample (foreground) anchor boxes and negative sample (background) anchor boxes as the distillation anchor boxes. At the same time, since the decoder also has a multi-level structure, on this basis, use progressive multi-level knowledge distillation to better obtain the dark knowledge of the teacher model. In the part of calculating cross attention in each layer of the decoder, use the attention weight matrix of the teacher model to guide the student model, and weighted fuse the attention weights corresponding to the distillation points, so that the student model can obtain target features with richer semantic information.

[0042] Train a teacher network model with ResNet101 as the backbone network. Based on the trained teacher network model, generate soft labels using high temperature. At this time, the loss function of the student model is no longer the loss function of hard labels, but a composite loss function composed of teacher soft labels and student hard labels. In the loss function, the teacher soft labels make the class probability distribution of the student model as close as possible to that of the teacher model, and make the feature responses of the student model and the teacher model as close as possible under the squared error loss; the student hard labels in the loss function are the prediction results of the student model. The composite loss function is obtained by weighting the distillation loss (teacher soft label part) and the student model loss (student hard label), and the loss function is

[0043] Through positive and negative sample anchor box query sampling, the student model can focus on the areas that the teacher pays more attention to, while random sampling provides the teacher model's view of the features. During distillation training, these selected distillation anchor boxes do not participate in the backpropagation of the student model. The loss functions of positive and negative sample anchor boxes and random sampling anchor boxes are

[0044] Step 4: Obtain the image to be detected, process it and input it into the fault detection model, calculate the anomaly score, and obtain the steel floor breakage detection result.

[0045] The test results of the present invention are as Figure 4 shown.

[0046] The above is only the specific implementation manner of the present invention. Any feature disclosed in this specification, unless specifically described, can be replaced by other equivalent or alternative features with similar purposes; all the features disclosed, or all the steps in any method or process, except for mutually exclusive features or steps, can be combined in any way.

Claims

1. A method for detecting damage to the steel floor of a railway freight car based on knowledge distillation, characterized in that, It includes the following steps: Step 1: Obtain images of the bottom of the train from multiple angles; Step 2: Select the pictures with damaged steel floors and the pictures without damage, mark the damaged parts, and separate the samples without failures from a small number of samples with failures; Step 3: Construct a fault detection model, including a teacher model and a student model; Step 4: Use the images of the bottom of the train obtained in Step 1 to train the teacher model and the student model in parallel. Transfer the knowledge of the teacher model to the student model by using progressive multi-level knowledge distillation and teacher feature distillation. At the same time, the teacher model layer by layer transfers the dark knowledge of the decoder layer in the teacher model to the student model. According to the knowledge distillation loss function, train the student model to achieve knowledge distillation, and finally obtain a trained fault detection model; Step 5: Verify the model detection result, obtain the image to be detected, input it into the fault detection model, calculate the anomaly score, and obtain the detection result of the damaged steel floor; Through positive and negative sample anchor box query sampling, during the distillation training, the selected distillation anchor boxes do not participate in the backpropagation of the student model. The loss functions of the positive and negative sample anchor boxes and the randomly sampled anchor boxes are Among them, is the positive sample anchor box distillation loss, is the negative sample anchor box distillation loss, is the random anchor box distillation loss.

2. The method for detecting the damage of the steel floor of a railway freight car based on knowledge distillation according to claim 1, wherein, The teacher model and the student model have the same structure, both including a backbone network module, an encoder module, a decoder module, and a prediction output module, except that the teacher model and the student model use backbone network modules of different sizes; After the image passes through the backbone network module, high-dimensional vector information is extracted and sent to the encoder module. The encoder module performs semantic encoding on the features and then sends them to the decoder module. The decoder performs cross-attention on the key values of the feature map and the corresponding region features, and finally outputs the final detection result through the prediction module.

3. The method for detecting damage of the steel floor of a railway freight car based on knowledge distillation according to claim 2, wherein, The backbone network module includes an input layer, a first group of convolutional layers, a max pooling layer, a second group of convolutional layers, a third group of convolutional layers, a fourth group of convolutional layers, and a fifth group of convolutional layers connected in sequence; Among them, the input size of the input layer is image data of 513x513; the first group of convolutional layers includes 1 7x7 convolutional operation and 1 non-linear activation function operation in sequence. The second group of convolutional layers includes 9 convolutional layers, a non-linear activation layer, and an average pooling layer in sequence. The third group of convolutional layers includes 12 convolutional layers, a non-linear activation layer, and an average pooling layer in sequence. The fourth group of convolutional layers includes 69 convolutional layers, a non-linear activation layer, and an average pooling layer in sequence. The fifth group of convolutional layers includes 9 convolutional layers, a non-linear activation layer, and an average pooling layer in sequence. Among them, each convolutional layer in the second group to the fifth group of convolutional layers passes through 1 1x1 convolution, 1 3x3 convolution, and 1 1x1 convolution operation in sequence.

4. The method for detecting the damage of the steel floor of a railway freight car based on knowledge distillation according to claim 2, characterized in that The encoder module includes six encoders, which are Transformer encoders. The encoder module is formed by stacking 6 identical encoders. Each encoder has two sub-layers. The first sub-layer is a multi-head self-attention pooling layer; the second sub-layer is a position-based feed-forward neural network layer; each sub-layer adopts a residual connection; the encoder adds the serialized feature map and the position encoding to obtain the query Q and the key-value K, adds them to the feature map and normalizes them after passing through a multi-head self-attention layer, and then passes through a feed-forward network to obtain the output of a single encoder, which is used as the input of the next encoder. After passing through 6 identical encoder structures, the output of the encoder part is obtained; Decoder module: The input of the decoder consists of three parts: the output of the encoder, the position encoding, and the query; among them, the dimension of the query is (300, 4), the first dimension is the number of predefined target queries, and the second dimension is the number of hidden layers; The first half of the decoder's calculation of the multi-head self-attention is the same as that of the encoder operation, and then it calculates the cross-attention with the output of the encoder. After passing through 6 identical decoder structures, the output of the decoder part is obtained; Prediction output module: It includes a feed-forward neural network and a fully connected layer; the feed-forward neural network layer is mainly divided into two parts, one part predicts the category, and the other part predicts the position; the branch for predicting the category consists of a linear layer with a hidden layer dimension of 512; since there is a background class (empty class), the output dimension is the number of classes plus 1; the other branch of the feed-forward neural network for predicting the position mainly consists of 3 linear layers with a hidden layer dimension of 512, and both of these branches will pass through a sigmoid activation function.

5. The method for detecting damage of the steel floor of a railway freight car based on knowledge distillation according to claim 2, characterized in that, The backbone network of the teacher model is ResNet101, and the backbone network of the student model is ResNet18.

6. The method for detecting damage of the steel floor of a railway wagon based on knowledge distillation according to claim 1, wherein The loss function of the student model in step 4 is a composite loss function composed of the teacher soft label and the student hard label: where α and β are hyperparameters, is the teacher soft label loss, is the student hard label loss.

Citation Information

Patent Citations

  • Railway wagon bottom floor damage fault detection method

    CN111652227A

  • Pre-trained language model compression method and platform based on Knowledge distillation

    CN111767711A