A Fault Detection Method for Railway Freight Car Parts Based on Deep Anomaly Detection
Through the method based on deep anomaly detection, normal samples are learned and generative adversarial network is constructed, which solves the problems of poor fault detection robustness and inability to identify unknown faults in the prior art, and efficient identification and accurate detection of various faults are achieved.
Patent Information
- Application Number
- CN202210229806.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-10
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-03-10
AI Technical Summary
The existing railway truck fault detection method is based on template matching, with poor robustness and low detection accuracy, and it cannot identify unknown faults, and can only detect known faults based on target detection method.
Using a method based on deep anomaly detection, by learning a large number of normal samples, the infrequent representations are regarded as abnormal, and a generative adversarial network is constructed to generate high-quality unfailed sample pictures, and the pictures are encoded through the encoded network to achieve fault detection.
It realizes the identification of various faults, reduces the intensity of manual labor, improves the robustness and accuracy of detection, and the AUC reaches 98.07%, which can effectively assist in manual fault detection.
Smart Images

Figure CN114663370B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of fault detection of railway freight car parts, and particularly to a fault detection method for railway freight car parts based on deep anomaly detection. Background Art
[0002] At present, the fault detection of railway freight cars mainly relies on high-speed cameras to capture passing trains, and then dynamic inspectors judge faults with the naked eye. Traditional railway freight car fault detection algorithms rely on template matching methods. By establishing a fault template library, the collected pictures are matched, and if a match is found, it is considered that a fault has occurred. However, this method has many limitations. One is that it is easily affected by environmental conditions such as light and shooting angle. The other is that due to the various manifestations of faults, it is difficult to establish a complete fault template library. At present, the object detection method based on deep learning can well solve the first problem, but the second problem still exists. The object detection method can only detect faults with known manifestations. Therefore, an algorithm that can identify all faults is very necessary. Summary of the Invention
[0003] The object of the present invention is to address the problems of poor robustness and low detection accuracy of the current railway freight car fault detection method based on template matching, and the inability of the object detection method to identify unknown faults. A fault detection method for railway freight car parts based on deep anomaly detection is proposed. By learning a large number of normal samples, the infrequently occurring manifestations are regarded as anomalies and reported as faults, which can cope with various manifestations of faults and reduce the labor intensity of manual work.
[0004] The present invention adopts the following technical solutions to achieve the above object: A fault detection method for railway freight car parts based on deep anomaly detection, specifically including:
[0005] Step S1: Obtain line-array images;
[0006] Step S2: Screen the obtained pictures and retain the images containing the target parts; there is no need to perform additional annotation on the images, and only a large number of samples without faults need to be distinguished from a small number of samples with faults.
[0007] Preferably, according to the images obtained in step S1, and according to the position information of the photos, screen the pictures containing the position of the cock handle; at the same time, judge whether the target part has a fault, and distinguish the picture samples without faults from the picture samples with faults, without performing additional annotation on the images.
[0008] Step S3: Use deep learning technology to construct a generative adversarial network for generating samples of normal targets.
[0009] Preferably, in order to make the generated images contain more details for convenient fault location, a generative adversarial network is constructed using a proGAN-style structure, which can generate high-quality and high-resolution sample images; the network is divided into a generator Generator and a discriminator Discriminator.
[0010] The generator Generator is mainly composed of three modules: GenInitBlock, GenConvBlock, and Convert2RGB. GenInitBlock is used to convert the input random noise into an initial feature map. GenConvBlock is used to increase the size of the feature map and enrich the quantity and details of features through upsampling and convolution operations. Convert2RGB is used to convert the multi-channel feature map into an RGB three-channel image.
[0011] Preferably, GenInitBlock first performs a PixelwiseNorm operation on the input random noise. Here, the dimension of the input noise is 512, and the PixelwiseNorm formula is: mean represents taking the mean along the feature channel dimension. Then, through a transposed convolution operation, the noise is converted into a 4x4 feature map. After activation by the LeakyRelu activation function, it passes through a 3x3 convolution and LeakyRelu, and finally passes through another PixelwiseNorm operation for output.
[0012] GenConvBlock consists of an upsampling function and two convolutional layers. The upsampling function is used to enlarge the size of the feature map, and the two convolutional layers are used to supplement features so that the upsampled feature map has richer detailed information.
[0013] The Convert2RGB module is implemented using a convolutional layer Conv(n,3), where n represents the number of feature map channels and 3 represents the RGB three channels.
[0014] Preferably, the discriminator Discriminator is mainly composed of three modules: FromRGB, DisConvBlock, and DisFinalBlock. The FromRGB module is used to convert the RGB three-channel image into a multi-channel feature map. The DisConvBlock is used to reduce the dimension of the feature map to reduce the number of parameters and extract higher-order features. The DisFinalBlock is used to score the image based on the features extracted by the previous modules to determine whether the image is real or fake.
[0015] Preferably, the FromRGB module is implemented using a convolutional layer Conv(3,n), where 3 represents the three RGB channels and n represents the number of feature map channels.
[0016] The DisConvBlock consists of two convolutional layers and a downsampling function. The convolutional layers are used to extract higher-order features from the previous-level features, and the downsampling is used to reduce the size of the feature map and reduce the computational amount.
[0017] The DisFinalBlock consists of a MinibatchStdDev function and three convolutional layers. The MinibatchStdDev is used to replace Batchnorm to improve the diversity of the data generated by the generator model. The formula is: That is, the root mean square of the feature map is calculated in the Batch dimension, and then this value is extended to the size of the feature map and concatenated with the input x. Here, mean represents the mean in the Batch dimension. The three convolutional layers are used to convert the obtained feature map into a score. The higher the score, the more the discriminator believes that the input image is a real image.
[0018] Step S4: Use the non-faulty samples described in Step S2 to train the network described in Step S3 so that it can generate high-quality non-faulty sample pictures.
[0019] Preferably, the training is carried out in an adversarial manner, and the generator and the discriminator are updated alternately, so that finally the generator can generate pictures that the discriminator cannot distinguish between true and false. At the same time, the network is trained step by step. Starting from generating pictures with a size of 4x4, and then gradually training to generate 8x8, 16x16... Finally, clear pictures with a size of 256x256 can be generated.
[0020] Step S5: Use deep learning technology to construct an encoding network for encoding pictures. The purpose of encoding is to enable the generative adversarial network described in Step S3 to regenerate pictures according to the encoding.
[0021] Preferably, the encoder network Encoder is used. The encoder consists of an EncoderInitBlock and several EncoderConvBlocks, and is used to encode pictures and map a picture to a 512-dimensional vector.
[0022] Preferably, the EncoderInitBlock consists of three convolutional layers and an average pooling AvgPool. The three convolutional layers are used to extract different features from the input picture to form a multi-channel feature map, and the average pooling layer is used to reduce the size of the feature map and reduce the computational amount.
[0023] The EncoderConvBlock consists of two convolutional layers and an average pooling AvgPool. The convolutional layers are used to extract higher-order features from the features of the previous level, and the average pooling AvgPool is used to reduce the size of the feature map and reduce the computational amount.
[0024] Step S6: Use the samples without failures described in Step S2 to train the encoding network described in Step S5, so that the network described in Step S3 can reconstruct the picture well according to the encoding.
[0025] Preferably, initialize the parameters of the encoder network described in Step S5, input the picture x in the dataset into the encoding network for encoding to obtain the encoding z; then input z into the generator trained in Step S4 to generate the picture x'; use x and x' to calculate the error for backpropagation. The error function is L =
[0026] L rec +L dis where L rec represents the reconstruction error, calculated using MSE (root mean square error), and L dis represents the discriminator error, that is, use the discriminator to discriminate x and x', and the judgment results of the two should be similar. L sis = MSE(D(x), D(x')), where D represents the discriminator. During the training process, fix the parameters of the generator and the discriminator. After training the entire dataset for 100 epochs, obtain the trained encoder network.
[0027] Step S7: Use the failure samples described in Step S2 to test the entire network. When a picture without failure is input into the network, the network encodes and then reconstructs the picture, and the error between the reconstructed picture and the input picture is small; while when a picture with failure is input into the network, the network encodes and then reconstructs the picture, and the error between the reconstructed picture and the input picture is large. Pixel-level failure localization can be performed according to the difference between the input and output pictures, and at the same time, an anomaly score is calculated through the input picture and the reconstructed picture to determine whether the picture contains a failure. According to the test results, the AUC of this method can reach 98.07%, and the failure can be well identified.
[0028] Preferably, the anomaly score calculation method is F-Score = MSE(x, x') + MSE(D(x), D(x')), where MSE represents the root mean square error function, x and x' represent the input picture and the reconstructed picture respectively, and D represents the discriminator.
[0029] Use AUC to measure the performance of the method. AUC represents the area enclosed by the ROC curve and the coordinate axes. The ROC curve is composed of the false positive rate and the true positive rate The closer the AUC is to 1, the better it is at distinguishing normal samples from abnormal samples.
[0030] Beneficial effects: Compared with the traditional railway freight car fault detection algorithm based on template matching, this method is not restricted by environmental conditions such as lighting and shooting angle, and has better robustness; compared with the method based on target detection, this method summarizes the normal representation of parts by learning a large number of normal samples, and regards other representations as abnormalities. It can detect faults well without exhaustively enumerating possible representation forms of faults, and achieved an AUC of 98.07% on the test set. It can well assist manual fault detection and greatly reduce the labor intensity of inspectors. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is the algorithm flow chart of the present invention
[0032] Figure 2 This is the structure and training diagram of the generative adversarial network in the present invention.
[0033] Figure 3 This is a schematic diagram of the structure of the generator part of the generative adversarial network in the present invention.
[0034] Figure 4 This is a schematic diagram of the structure of the discriminator part of the generative adversarial network in the present invention.
[0035] Figure 5 It is a schematic diagram of the structure of the encoder part in the present invention
[0036] Figure 6 Schematic diagram of the encoder training method in the present invention DETAILED DESCRIPTION
[0037] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation methods.
[0038] like Figure 1 As shown, the present invention aims at the shortcomings of poor robustness of traditional template matching methods and inability of target detection methods to exhaustively enumerate fault types. A railway freight car fault detection method based on deep anomaly detection is proposed, which can greatly improve the generalization and accuracy of the detection method and reduce the labor intensity of manual work. Specifically, it includes the following steps:
[0039] Step S1: Collect the train image to be detected. High-speed cameras are arranged around the rails to capture images of various parts of the train when it passes by.
[0040] Step S2: Screen the acquired images and keep the images containing the target parts. No additional annotation is required for the images, and only a large number of samples without faults need to be distinguished from a small number of samples with faults.
[0041] Step S21: Here, take the handle part of the railway wagon cock as an example. According to the image obtained in step S1 and the position information of the photo, filter the picture containing the handle part of the cock. At the same time, judge whether the target part has a fault, distinguish the picture samples without faults from the picture samples with faults, and there is no need to perform additional annotation on the images.
[0042] Step S3: As Figure 2 shown, use deep learning technology to construct a generative adversarial network for generating samples with normal targets. In order to make the generated pictures contain more details for convenient fault location, use the proGAN-style structure to construct the generative adversarial network. The generative adversarial network can be divided into two parts, the generator Generator and the discriminator Discriminator. The former is used to generate forged pictures (fake images) with the same distribution as the training samples, and the discriminator needs to judge whether the input picture is a forged picture generated by the generator or a real picture from nature. The two learn against each other, and finally the generator can generate forged pictures that are the same as the real pictures.
[0043] Step S31: As Figure 2 shown, the generator Generator is mainly composed of three modules: GenInitBlock, GenConvBlock, and Convert2RGB. GenInitBlock is used to convert the input random noise into an initialized feature map. GenConvBlock is used to increase the size of the feature map through upsampling and convolution operations, enrich the quantity and details of the features. Convert2RGB is used to convert the multi-channel feature map into an RGB three-channel image.
[0044] Step S311: As Figure 3 (a) shown, GenInitBlock first performs a PixelwiseNorm operation on the input random noise. Here, the dimension of the input noise is 512, and the PixelwiseNorm formula is: mean represents taking the mean in the feature channel dimension. Then, through the transposed convolution operation, the noise is converted into a 4x4 feature map. After activation by the LeakyRelu activation function, it passes through a 3x3 convolution and LeakyRelu, and finally passes through another PixelwiseNorm operation for output.
[0045] Step S312: As Figure 3 (b) shown, GenConvBlock consists of an upsampling function and two convolutional layers. The upsampling function is used to enlarge the size of the feature map, and the two convolutional layers are used to supplement features, so that the upsampled feature map has richer detailed information.
[0046] Step S313: The role of the Convert2RGB module is to convert the multi-channel feature map into an RGB three-channel color image. Therefore, only one convolutional layer Conv(n,3) can achieve this, where n represents the number of channels of the feature map and 3 represents the RGB three channels.
[0047] Step S32: As Figure 2 shown, the discriminator is mainly composed of three modules: FromRGB, DisConvBlock, and DisFinalBlock. The FromRGB module is used to convert the RGB three-channel image into a multi-channel feature map. The DisConvBlock is used to reduce the dimensionality of the feature map to reduce the number of parameters and extract higher-order features. The DisFinalBlock is used to score the image based on the features extracted by the previous modules to determine whether the image is real or forged.
[0048] Step S321: The FromRGB module is mainly used to convert the RGB three-channel image into a multi-channel feature map for matching subsequent modules. Therefore, only one convolutional layer Conv(3,n) can achieve this, where 3 represents the RGB three channels and n represents the number of channels of the feature map.
[0049] Step S322: As Figure 4 (a) shown, the DisConvBlock consists of two convolutional layers and a downsampling function. The convolutional layers are used to extract higher-order features from the previous-level features, and the downsampling is used to reduce the size of the feature map and reduce the computational amount.
[0050] Step S323: As Figure 4 (b) shown, the DisFinalBlock consists of a MinibatchStdDev function and three convolutional layers. The MinibatchStdDev is used to replace Batchnorm to improve the diversity of the data generated by the generator model. The formula is: That is, calculate the root mean square of the feature map in the Batch dimension, and then expand this value to the size of the feature map and concatenate it with the input x. Here, mean represents the mean in the Batch dimension. The three convolutional layers are used to convert the obtained feature map into a score. The higher the score, the more the discriminator believes that the input image is a real image.
[0051] Step S4: As Figure 2As shown, use the non-faulty samples described in step S2 to train the network described in step S3 so that it can generate high-quality non-faulty sample images. The training is carried out in an adversarial manner, with the generator and discriminator updated alternately, so that finally the generator can generate images that the discriminator cannot distinguish between true and false. At the same time, the network is trained step by step, starting from generating images of size 4x4, and then gradually training to generate 8x8, 16x16... Finally, clear images of size 256x256 can be generated.
[0052] Step S41: As Figure 2 shown, the initial training starts from 4x4 images. Downsample the real images in the dataset to 4x4 as the input to the discriminator. At the same time, sample a 512-dimensional noise from a normal distribution as the input to the generator. The generator generates a forged 4x4 image based on the input noise. Update the parameters of the generator and discriminator using the adversarial loss. Train the entire dataset for 120 epochs.
[0053] Step S411: The loss function formula is where D and G represent the discriminator and generator respectively, x ∼ p data denotes selecting an image x from the dataset, and z ∼ p z (z) represents randomly sampling a noise z from a normal distribution.
[0054] Step S42: As Figure 2 shown, after the previous step of training is completed, insert a GenConvBlock between the GenInitBlock and Convert2RGB of the generator so that the generator can generate images twice the size of the previous stage; at the same time, insert a DisConvBlock between the FromRGB and DisFinalBlock of the discriminator so that the discriminator can process images twice the size of the previous stage. At the same time, downsample the real images in the dataset, and the size is also twice that of the previous stage. Update the parameters of the generator and discriminator using the adversarial loss. Train the entire dataset for 120 epochs.
[0055] Step S43: As Figure 2 shown, repeat step S42 until the generator can generate images of size 128.
[0056] Step S5: As Figure 5As shown in (a), using deep learning technology, an encoding network is constructed to encode pictures. The purpose of encoding is to enable the generative adversarial network described in step S3 to regenerate pictures based on the encoding. The encoder consists of an EncoderInitBlock and several EncoderConvBlocks, which are used to encode pictures and map a picture to a 512-dimensional vector.
[0057] Step S51: As Figure 5 shown in (b), the EncoderInitBlock consists of three convolutional layers and an average pooling AvgPool. The three convolutional layers are used to extract different features from the input picture to form a multi-channel feature map, and the average pooling layer is used to reduce the size of the feature map and reduce the computational amount.
[0058] Step S52: As Figure 5 shown in (c), the EncoderConvBlock consists of two convolutional layers and an average pooling AvgPool. The convolutional layers are used to extract higher-order features from the previous-level features, and the average pooling AvgPool is used to reduce the size of the feature map and reduce the computational amount.
[0059] Step S6: As Figure 6 shown, use the samples without failures described in step S2 to train the encoding network described in step S5, so that the network described in step S3 can reconstruct the pictures based on the encoding.
[0060] Step S61: Initialize the parameters of the encoder network described in step S5. Input the picture x in the dataset into the encoding network for encoding to obtain the encoding z. Then input z into the generator trained in step S4 to generate the picture x'.
[0061] Step S62: Calculate the error using x and x' for backpropagation. The error function is L = L rec +L dis , where L rec represents the reconstruction error, calculated using MSE (mean square root error), and L dis represents the discriminator error, that is, use the discriminator to discriminate between x and x'. The judgment results of the two should be similar. L dis = MSE(D(x), D(x')), where D represents the discriminator.
[0062] Step S63: As Figure 6 shown, during the training process, fix the parameters of the generator and the discriminator. After training the entire dataset for 100 epochs, obtain the trained encoder network.
[0063] Step S7: Use the fault samples described in Step S2 to test the entire network. When an image without a fault is input into the network, the network encodes and then reconstructs the image, and the error between the reconstructed image and the input image is small; while when an image with a fault is input into the network, the network encodes and then reconstructs the image, and the error between the reconstructed image and the input image is large. Pixel-level fault localization can be performed based on the difference between the input and the reconstructed output images. At the same time, an anomaly score can be calculated from the input and the reconstructed output images to determine whether the image contains a fault. According to the test results, the AUC of this method can reach 98.07%, and faults can be well identified.
[0064] Step S71: The anomaly score calculation method is F-Score = MSE(x, x') + MSE(D(x), D(x')), where MSE represents the root mean square error function, x and x' represent the input image and the reconstructed image respectively, and D represents the discriminator.
[0065] Step S72: AUC represents the area enclosed by the ROC curve and the coordinate axes. The ROC curve is calculated from the false positive rate and the true positive rate. The closer the AUC is to 1, the better it can distinguish normal samples from abnormal samples.
[0066] It should be noted that for those of ordinary skill in the art, without departing from the principle described in the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A fault detection method for railway wagon parts based on deep anomaly detection, characterized in that: The specific steps include: Step S1: Obtain line array images; Step S2: Screen the obtained pictures and retain the images containing the target parts; there is no need to perform additional annotation on the images, just distinguish a large number of non-faulty samples from a small number of faulty samples; Step S3: Use deep learning technology to construct a generative adversarial network for generating normal target samples; Step S4: Use the non-faulty samples described in Step S2 to train the network described in Step S3 so that it can generate high-quality non-faulty sample pictures; Step S5: Use deep learning technology to construct an encoding network for encoding pictures, and the purpose of encoding is to enable the generative adversarial network described in Step S3 to regenerate the pictures according to the encoding; Step S6: Use the non-faulty samples described in Step S2 to train the encoding network described in Step S5 so that the network described in Step S3 can reconstruct the pictures well according to the encoding; Step S7: Use the faulty samples described in Step S2 to test the entire network; when a non-faulty picture is input into the network, the network encodes and then reconstructs the picture, and the error between the reconstructed picture and the input picture is small; while when a faulty picture is input into the network, the network encodes and then reconstructs the picture, and the error between the reconstructed picture and the input picture is large; perform pixel-level fault location according to the difference between the input and output pictures, and at the same time calculate the anomaly score through the input picture and the reconstructed picture to determine whether the picture contains a fault; In the said Step S3, in order to make the generated pictures contain more details for convenient fault location, use a proGAN-style structure to construct a generative adversarial network to generate high-quality and high-resolution sample pictures; the network is divided into a generator Generator and a discriminator Discriminator; The said generator Generator is mainly composed of three modules: GenInitBlock, GenConvBlock and Convert2RGB; GenInitBlock is used to convert the input random noise into an initialized feature map, GenConvBlock is used to increase the size of the feature map and enrich the quantity and details of the features through upsampling and convolution operations, and Convert2RGB is used to convert the multi-channel feature map into an RGB three-channel image; The GenInitBlock first performs PixelwiseNorm on the input random noise. Here, the dimension of the input noise is 512, and the PixelwiseNorm formula is: mean represents calculating the mean along the feature channel dimension; then through deconvolution operation, the noise is transformed into a 4x4 feature map. After activation by the LeakyRelu activation function, it goes through a 3x3 convolution and LeakyRelu, and finally goes through another PixelwiseNorm operation for output; The said GenConvBlock consists of an upsampling function and two convolutional layers. The upsampling function is used to enlarge the size of the feature map, and the two convolutional layers are used to supplement the features so that the upsampled feature map has richer detailed information; The said Convert2RGB module is implemented using a convolutional layer Conv(n,3), where n represents the number of channels of the feature map and 3 represents the RGB three channels; The discriminator mainly consists of three modules: FromRGB, DisConvBlock, and DisFinalBlock. The FromRGB module is used to convert an RGB three-channel image into a multi-channel feature map. The DisConvBlock is used to reduce the dimensionality of the feature map to reduce the number of parameters and extract higher-order features. The DisFinalBlock is used to score the image based on the features extracted by the previous modules to determine whether the image is real or forged. The FromRGB module is implemented using a convolutional layer Conv(3,n), where 3 represents the RGB three channels and n represents the number of channels of the feature map. The DisConvBlock consists of two convolutional layers and a downsampling function. The convolutional layers are used to extract higher-order features from the previous-level features, and the downsampling is used to reduce the size of the feature map and reduce the computational amount. The DisFinalBlock consists of a MinibatchStdDev function and three convolutional layers. The MinibatchStdDev is used to replace Batchnorm to improve the diversity of the data generated by the generator model. The formula is as follows: That is, the root mean square of the feature map is calculated in the Batch dimension, and then this value is extended to the size of the feature map and concatenated with the input x. Here, mean represents the mean value calculated in the Batch dimension. The three convolutional layers are used to convert the obtained feature map into a score. The higher the score, the more the discriminator believes that the input image is a real image.
2. A method for detecting faults in railway freight car parts based on deep anomaly detection according to claim 1, characterized in that: In step S2, according to the image obtained in step S1 and the position information of the photo, the pictures containing the position of the cock handle are screened; at the same time, it is judged whether a fault occurs in the target part, and the picture samples without faults are distinguished from the picture samples with faults, and no additional annotation of the image is required.
3. A method for detecting faults in railway freight car parts based on deep anomaly detection according to claim 1, characterized in that: In step S4, the training is carried out in an adversarial manner, and the generator and the discriminator are updated alternately, so that the final generator can generate pictures that the discriminator cannot distinguish between true and false; at the same time, the network is trained in a step-by-step manner, starting from generating pictures with a size of 4x4, and then gradually training to generate 8x8, 16x16... Finally, clear pictures with a size of 256x256 can be generated.
4. A method for detecting faults in railway freight car parts based on deep anomaly detection according to claim 1, characterized in that: In step S5, an encoder network Encoder is used. The encoder consists of an EncoderInitBlock and several EncoderConvBlocks, and is used to encode the picture and map a picture to a 512-dimensional vector.
5. A method for detecting faults in railway freight car parts based on deep anomaly detection according to claim 4, characterized in that: The EncoderInitBlock consists of three convolutional layers and an average pooling AvgPool. The three convolutional layers are used to extract different features from the input picture to form a multi-channel feature map, and the average pooling layer is used to reduce the size of the feature map and reduce the amount of computation; the EncoderConvBlock consists of two convolutional layers and an average pooling AvgPool. The convolutional layers are used to extract higher-order features from the previous-level features, and the average pooling AvgPool is used to reduce the size of the feature map and reduce the computational amount.
6. A method for detecting faults in railway freight car parts based on deep anomaly detection according to claim 1, characterized in that: In step S6, initialize the parameters of the encoder network described in step S5, input the image x in the dataset into the encoding network for encoding to obtain the encoding z; then input z into the generator trained in step S4 to generate the image x'; calculate the error using x and x' for backpropagation; the error function is L = L rec + L dis , where L rec represents the reconstruction error, calculated using MSE, and L dis represents the discriminator error, that is, use the discriminator to discriminate between x and x', and the two judgment results should be similar. L dis = MSE(D(x), D(x')), where D represents the discriminator; fix the parameters of the generator and the discriminator during the training process; after training the entire dataset for 100 epochs, obtain the trained encoder network.
7. A method for detecting faults in railway freight car parts based on deep anomaly detection according to claim 1, characterized in that: In the step S7, the anomaly score calculation method is F-Score = MSE(x, x') + MSE(D(x), D(x')), where MSE represents the root mean square error function, x and x' respectively represent the input image and the reconstructed image, and D represents the discriminator; Use the AUC measurement method to measure the performance of the method; AUC represents the area enclosed by the ROC curve and the coordinate axes, and the ROC curve is calculated from the false positive rate and the true positive rate ; the closer the AUC is to 1, the better it can distinguish normal samples from abnormal samples.
Citation Information
Patent Citations
Railway wagon manual brake shaft chain falling fault image recognition method
CN111080605A
Brake shoe breaking target detection method
CN111091555A