Fault detection model training method and fault detection method
By training a fault detection model that comprehensively considers multiple dimensions, it solves the problem of high leakage detection rate and low efficiency of crack faults in railway truck wall panels, and achieves fast and accurate fault detection and positioning, improving detection efficiency and safety.
Patent Information
- Application Number
- CN202510350415.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-24
AI Technical Summary
Manual inspection and investigation into the crack failure of railway truck wall panels has problems such as high leakage detection rate and low efficiency, which affects the safe operation of railway freight vehicles and cargo safety.
A pre-constructed neural network model is used to predict the preset fault point annotation image, and a multi-stage loss function is solved with real annotation. A fault detection model that comprehensively considers multiple dimensions such as position, size, confidence and classification is trained.
It realizes rapid detection of faults and positioning of fault locations, reduces missed detection rate, improves detection efficiency, reduces labor costs, and reduces errors caused by human factors.
Smart Images

Figure CN120198778A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of fault detection, and particularly to a method for training a fault detection model and a method for fault detection. Background Art
[0002] The materials of railway freight carriages are weathering steel. Due to reasons such as high-intensity mechanized unloading of railway freight transportation and non-standard operations of loading and unloading workers, faults such as cracks in the carriage wallboards are likely to occur. The current crack fault detection of railway freight car wallboards, etc. is carried out through manual inspection and investigation. Timely detection of crack faults and implementation of welding repair play an important role in ensuring the safe operation of railway freight vehicles and the safety of goods.
[0003] However, manual inspection and investigation have disadvantages such as high missed detection rate, large workload, and low inspection efficiency, seriously affecting the safe operation of railway freight vehicles and the safety of goods. Summary of the Invention
[0004] The purpose of the present invention is to provide at least one method for training a fault detection model and a method for fault detection, which can at least solve problems such as high missed detection rate and low efficiency in manual inspection and investigation, and can at least achieve the effect of quickly detecting faults and locating the fault positions to guide personnel for maintenance.
[0005] To solve the above technical problems, at least one embodiment of the present application provides a method for training a fault detection model, including:
[0006] Using a pre-constructed neural network model to predict a preset fault point annotation image to obtain a prediction result;
[0007] According to the prediction result and the true annotation, solving a multi-stage loss function, where the multi-stage loss function includes at least two of a position loss term, a size loss term, a confidence loss term, and a classification loss term;
[0008] Solving the target model parameters that minimize the value of the multi-stage loss function;
[0009] Adjusting the neural network model according to the target model parameters to obtain a trained fault detection model.
[0010] At least one embodiment of the present application further provides a method for fault detection, including using the fault detection model obtained by training through the method for training a fault detection model to identify a target image to be recognized to obtain a fault detection result.
[0011] At least one embodiment of the present application further provides a training device for a fault detection model, including:
[0012] A prediction module, configured to use a pre-built neural network model to predict a preset fault point marked image, and obtain a prediction result;
[0013] A loss value calculation module, configured to solve a multi-stage loss function according to the prediction result and the true annotation, where the multi-stage loss function includes at least two of a position loss term, a size loss term, a confidence loss term, and a classification loss term;
[0014] A model parameter calculation module, configured to solve target model parameters that minimize the value of the multi-stage loss function;
[0015] A model parameter adjustment module, configured to adjust the neural network model according to the target model parameters to obtain a trained fault detection model.
[0016] At least one embodiment of the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned fault detection model training method or the fault detection method.
[0017] At least one embodiment of the present application further provides a computer-readable storage medium, storing a computer program, where the computer program realizes the above-mentioned fault detection model training method or the fault detection method when executed by a processor.
[0018] The fault detection model training method and the fault detection method provided by the embodiments of the present application predict a preset fault point marked image by using a pre-built neural network model, and solve a multi-stage loss function in combination with the true annotation, so that the model not only focuses on the performance of a single aspect (such as position accuracy or classification correctness) during training, but comprehensively considers multiple dimensions such as position, size, confidence, and classification, and optimizes multiple objectives at the same time, so as to more accurately capture fault features and more finely evaluate the difference between the prediction result and the true situation. This comprehensive training method helps the model learn richer feature representations, thereby improving its generalization ability in different scenarios and conditions. The trained fault detection model can automatically predict the fault points of the input image without manual intervention. This automated and intelligent detection method can greatly improve the detection efficiency, reduce the labor cost, and reduce the errors caused by human factors.
[0019] In some optional embodiments, the calculation formula of the multi-stage loss function includes the following:
[0020]
[0021] Among them, L is the multi-stage loss function, S is the grid size, B is the number of detection boxes, indicates that the object appears in cell i, indicates that the object does not appear in cell i, indicates that the j-th bounding box predictor in cell i is responsible for this prediction, x i and y i are the coordinates of the predicted box, and are the coordinates of the labeled box, ω i and h i are the width and height of the predicted box, and are the width and height of the labeled box, C i is the confidence of the predicted box, p i p(c) represents the probability of belonging to class c, γ, μ, and λ respectively represent hyperparameters, and classes is the set of preset fault classes.
[0022] In this embodiment, the model directly receives the original image as input and outputs information such as the location, size, class, and confidence of the fault. In traditional training methods, it is necessary to calculate the features of the candidate regions, and finally a support vector machine classifier is also required for classification to identify relevant faults. This embodiment integrates these steps into a unified model and trains it through a unified loss function, enabling the trained fault detection model to simultaneously learn tasks such as feature extraction, target localization, and classification, realizing integrated optimization, simplifying the process, and reducing the computational complexity. The trained fault detection model can automatically complete the whole process from image input to fault detection without manual intervention. It significantly improves the detection efficiency and is applicable to large-scale and real-time application scenarios.
[0023] In some optional embodiments, the neural network model is a convolutional neural network model, and the training method further includes:
[0024] Using an activation function after each convolutional layer, and the calculation formula of the activation function is as follows:
[0025]
[0026] In this embodiment, by using the above activation function, the error signal can be propagated faster during the backpropagation process, thereby accelerating the convergence speed of the model. At the same time, it can increase the non-linear expression ability of the network, enabling it to handle more complex tasks.
[0027] In some optional embodiments, the neural network model is provided with a batch normalization layer, and the batch normalization layer is located between the convolutional layer and the activation function.
[0028] In this embodiment, the batch normalization layer can normalize the distribution of the intermediate data of the network as much as possible, significantly accelerate the training of the network, and have a certain regularization ability, which helps to improve the overfitting situation.
[0029] In some alternative embodiments, solving for the target model parameters that minimize the value of the multi-stage loss function includes:
[0030] Using the stochastic gradient descent algorithm, gradually iteratively solve for the minimum value of the multi-stage loss function along the direction of the gradient descent to obtain the minimized multi-stage loss function and the target model parameters that minimize the multi-stage loss function.
[0031] In this embodiment, using the stochastic gradient descent algorithm can alleviate the problem of uneven sample numbers among different categories in the dataset and effectively avoid local optimal solutions.
[0032] In some alternative embodiments, the training method further includes:
[0033] Adjust the learning rate of the neural network model in stages, where the training stage of the neural network model includes a first stage, an intermediate stage, and a later stage. In the first stage, the learning rate gradually increases from a first threshold to a second threshold. In the intermediate stage, the learning rate remains unchanged at the second threshold. In the later stage, train with the first threshold as the learning rate.
[0034] In this embodiment, gradually increasing the learning rate in the first stage can improve the situation of divergence caused by unstable gradients. In the middle of training, the model has found a relatively good parameter space. At this time, keeping the learning rate unchanged can ensure stable iteration and optimization of the model within this parameter space and avoid training instability caused by changes in the learning rate. In the later stage of training, the model is close to the optimal solution. At this time, reducing the learning rate can make the model make more delicate adjustments near the optimal solution, which helps to improve the final performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings, and these exemplary illustrations do not limit the embodiments.
[0036] Figure 1 is a flowchart of a method for training a fault detection model provided by an embodiment of the present application;
[0037] Figure 2 is a schematic diagram of a convolutional neural network model provided by an embodiment of the present application;
[0038] Figure 3 is a flowchart of a method for fault detection provided by another embodiment of the present application;
[0039] Figure 4 It is a schematic diagram of a training device for a fault detection model provided by another embodiment of the present application. Detailed implementation manners
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are presented for the readers to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented. The following division of each embodiment is for convenience of description and should not constitute any limitation on the specific implementation manner of the present application. Each embodiment can be combined and cross-referenced with each other on the premise of no conflict.
[0041] The material of the railway freight car body is weathering steel. Due to reasons such as high-intensity mechanized unloading in railway freight transportation and non-standard operation of loading and unloading workers, faults such as cracks in the car body wallboard are likely to occur. The current inspection and troubleshooting of cracks in railway freight car body wallboards, etc., are carried out manually through patrol inspections. Timely detection of crack faults and implementation of welding repairs play an important role in ensuring the safe operation of railway freight vehicles and the safety of goods. However, manual patrol inspections have disadvantages such as high missed inspection rates, large workloads, and low inspection efficiency, which seriously affect the safe operation of railway freight vehicles and the safety of goods.
[0042] To solve the above technical problems of high missed inspection rates, large workloads, and low inspection efficiency in manual patrol inspections, the present invention proposes a training method for a fault detection model. By using the image acquisition function of existing railway maintenance monitoring equipment, a large number of images of cracks in railway freight car body wallboards are collected, and a deep learning object detection algorithm model is used to automatically screen them, timely detect and locate the crack positions, and give an alarm, reducing the missed inspection rate of traditional manual patrols, saving labor costs, and ensuring freight safety. The trained model can be used to solve problems such as real-time detection of cracks in railway freight car body wallboards and guide the maintenance work of freight cars.
[0043] The implementation details of the training method for the fault detection model in this embodiment will be specifically described below. The following content is only implementation details provided for easy understanding and is not necessary for implementing this solution.
[0044] Embodiment 1:
[0045] The training method for the fault detection model in this embodiment can be applied to an electronic device with communication, computing, and data storage capabilities. Its specific process can be as Figure 1 shown and includes:
[0046] Step 110: Use a pre - constructed neural network model to predict a preset fault - point - labeled image to obtain a prediction result;
[0047] In this embodiment, the fault - point positions on the side - wall panel image of the railway freight car are labeled, such as the specific position, size, category, etc. of the fault. During the labeling process, tools like labelme can be used to make marks in the form of rectangles, polygons, circles, straight lines, and points. After selecting a preset number of fault images for marking, a training set is generated based on the marked images.
[0048] The pre - constructed neural network model can be built based on object - detection architectures such as Convolutional Neural Network (CNN), Region - based Convolutional Neural Network (R - CNN), YOLO, Faster R - CNN, etc., and receives a batch of preset fault - point - labeled images as input. When the model makes a prediction, it extracts image features and obtains the prediction result through forward - propagation calculation. The prediction result includes the position box, size, confidence level, and fault category of the fault point, etc.
[0049] Step 120: Solve a multi - stage loss function according to the prediction result and the true label, where the multi - stage loss function includes at least two of a position loss term, a size loss term, a confidence - level loss term, and a classification loss term;
[0050] In this embodiment, the prediction result is compared with the true label to calculate the multi - stage loss function. The multi - stage loss function contains at least two loss terms, such as a position loss term (measuring the difference between the predicted position box and the true position box), a size loss term (measuring the difference between the predicted size and the true size), a confidence - level loss term (measuring the difference between the predicted confidence level and the true probability of the existence of the fault), and a classification loss term (measuring the difference between the predicted category and the true category). By introducing multiple loss terms, the neural network model can learn more diverse features during the training process, so as to show stronger adaptability when facing unknown or complex situations. Correspondingly, the more dimensions the neural network model is evaluated in during the training process, the more conducive it is to optimizing the neural network model simultaneously in multiple aspects and improving the overall fault - detection performance.
[0051] Traditional fault - detection methods are generally divided into two stages:
[0052] 1) Candidate region extraction and feature calculation: Process the input image to extract the features of candidate regions (such as regions that may contain faults). For example, in R - CNN, candidate regions are generated through Selective Search, and then CNN is used to extract features.
[0053] 2) Classification and recognition: Input the extracted features into a support vector machine (SVM) classifier for classification (such as determining whether a fault exists). The role of SVM is to classify candidate regions based on features to determine whether it is the target fault.
[0054] The conventional training method has a complex process, requiring candidate region extraction, feature extraction, and classification to be completed step by step, with low efficiency. The error of each step will be transmitted to the subsequent stage, affecting the final result. Moreover, it relies on an additional classifier and requires separate training of SVM, increasing the model complexity and computational cost.
[0055] The multi-stage loss function method in this embodiment integrates candidate region extraction, feature extraction, and classification into a unified model through end-to-end training. The loss function includes multiple tasks (such as classification loss, localization loss, confidence loss, etc.) and is jointly optimized through weighted summation. Throughout the training process, no additional classifier (such as SVM) is required, and the classification task is completed by the neural network itself. The finally trained fault recognition model directly receives the input image and outputs prediction results such as the location, category, and confidence of the fault, without intermediate steps. For example, models such as YOLO or Faster R-CNN complete all tasks through a single forward propagation.
[0056] Step 130, solve the target model parameters that minimize the value of the multi-stage loss function;
[0057] In this embodiment, optimization algorithms such as Stochastic Gradient Descent (SGD), Adam, RMSprop, etc. can be used to iteratively update the parameters of the neural network model to minimize the value of the multi-stage loss function. By continuously adjusting the model parameters, the prediction results of the model are made closer to the true annotations, thereby improving the accuracy of fault detection.
[0058] Step 140, adjust the neural network model according to the target model parameters to obtain a trained fault detection model.
[0059] In this embodiment, after solving the target model parameters that minimize the value of the multi-stage loss function, these parameters are applied to a pre-constructed neural network model to obtain a trained fault detection model. Through the trained model, accurate detection of unknown fault images is achieved, including fault localization, size estimation, and category recognition, etc.
[0060] The trained model needs to be verified and tested to evaluate its generalization ability on unknown data. Finally, the trained model is deployed to the railway freight online repair information system to automatically detect crack faults in a timely manner through the images uploaded in real time by the image acquisition device, and the model parameters are updated regularly for iterative optimization.
[0061] In summary, for the training method of the fault detection model provided in this embodiment, by using a pre-constructed neural network model to predict a preset fault point annotation image and combining the true annotation to solve the multi-stage loss function, the model not only focuses on the performance of a single aspect (such as position accuracy or classification correctness) during training, but comprehensively considers multiple dimensions such as position, size, confidence, and classification, and optimizes multiple objectives simultaneously, so as to more accurately capture fault features and more precisely evaluate the difference between the prediction result and the true situation. This comprehensive training method helps the model learn richer feature representations, thereby improving its generalization ability in different scenarios and conditions. The trained fault detection model can automatically predict the fault points of the input image without manual intervention. This automated and intelligent detection method can greatly improve the detection efficiency, reduce labor costs, and reduce errors caused by human factors.
[0062] In some alternative embodiments, the calculation formula of the multi-stage loss function L is as follows:
[0063]
[0064] where S is the grid size, B is the number of detection boxes, indicates that the object appears in cell i, indicates that the object does not appear in cell i, indicates that the j-th bounding box predictor in cell i is responsible for this prediction, x i and y i are the coordinates of the prediction box, and are the coordinates of the annotation box, ω i and h i are the width and height of the prediction box, and are the width and height of the annotation box, C i is the confidence of the prediction box, p i (c) represents the probability of belonging to class c, and γ, μ, λ represent hyperparameters respectively, and classes is the preset set of fault categories.
[0065] In this embodiment, the multi-stage loss function includes the following 5 loss terms:
[0066] 1. Location loss:
[0067] This term is used to measure the location difference between the prediction box and the true box. γ is used to adjust the weight of the location loss.
[0068] 2. Size loss:
[0069] This item is used to measure the size difference between the predicted bounding box and the ground truth bounding box. μ is used to adjust the weight of the size loss.
[0070] 3. Confidence loss:
[0071] This item is used to measure the difference between the confidence of the predicted bounding box and the confidence of the ground truth bounding box.
[0072] 4. No-object confidence loss:
[0073] This item is used to handle cells without objects. λ is used to adjust the weight of the no-object confidence loss.
[0074] 5. Classification loss:
[0075] This item is used to measure the difference between the predicted class probability and the ground truth class probability.
[0076] In this embodiment, the model directly receives the original image as input and outputs information such as the location, size, class, and confidence of the fault. In traditional training methods, it is necessary to calculate the features of the candidate regions, and finally a support vector machine classifier is also required for classification to identify relevant faults. This embodiment integrates these steps into a unified model and trains it through a unified loss function, enabling the trained fault detection model to simultaneously learn tasks such as feature extraction, target localization, and classification, achieving integrated optimization, simplifying the process, and reducing the computational complexity. The trained fault detection model can automatically complete the entire process from image input to fault detection without manual intervention. It significantly improves the detection efficiency and is applicable to large-scale and real-time application scenarios. Especially applicable to the scenario of automatic detection of faults on the side walls of railway freight cars.
[0077] In some alternative embodiments, the neural network model is a convolutional neural network model, and the training method further includes: using an activation function after each convolutional layer, and the calculation formula of the activation function is as follows:
[0078]
[0079] Specifically, as shown in Figure 2, it is the convolutional neural network model in this embodiment. The backbone of this convolutional neural network model is composed of stacked convolutional layers (conv.layers) and 2 subsequent fully connected layers. The numbers in the figure represent the size of the feature map and the size of the convolutional kernel. The trained model inputs the real-time collected images of the railway wagon wall panels, and automatically annotates the fault positions after passing through the model. The convolutional layer is a network layer that contains convolutional calculations. Through the convolutional structure, a convolutional kernel of a given size slides on the input data at a certain stride, performs the dot product operation of matrices, forms feature extraction within the receptive field, and the calculated new tensor is used as the input of the next convolutional layer. Due to the introduction of the convolutional kernel, the computational amount of the network is greatly reduced. This phenomenon also plays a positive role in deepening the number of network layers.
[0080] Each convolutional layer is activated using the above activation function. By using the above activation function, the error signal can be propagated faster during the backpropagation process, thereby accelerating the convergence speed of the model. At the same time, it can increase the non-linear expression ability of the network, enabling it to handle more complex tasks. The last layer of the convolutional neural network model will predict the class probabilities and bounding box coordinates.
[0081] In some alternative embodiments, the neural network model is provided with a batch normalization layer, and the batch normalization layer is located between the convolutional layer and the activation function.
[0082] Specifically, the batch normalization layer (Batch Normalization, BN) represents the batch normalization process. The BN layer can make the distribution of the intermediate data in the network as normalized as possible, can significantly accelerate the training of the network, and has a certain regularization ability, which helps to improve the overfitting situation. The calculation formula of BN in this embodiment is as follows:
[0083]
[0084] where, m represents the size of the convolutional kernel, x i represents all trainable parameters on the same batch and channel, ε represents the variance estimation value constant (the value range is between 0 and 1, and can be adjusted according to the training results), τ represents the stretching parameter, and β represents the offset parameter.
[0085] The regularization effect of the BN layer provides a stable training environment for the optimization of the multi-stage loss function by reducing overfitting and stabilizing gradient propagation. The non-linear characteristics of the activation function enable the multi-stage loss function to drive the model to learn multi-task features more efficiently by enhancing the model expression ability and feature decoupling. The synergistic effect of the two enables the model to simultaneously optimize multiple loss terms in complex tasks (such as object detection, multi-label classification), significantly improving the overall performance.
[0086] In some alternative embodiments, solving for the target model parameters that minimize the value of the multi-stage loss function includes: using the stochastic gradient descent algorithm to gradually iteratively solve for the minimum value of the multi-stage loss function along the direction of gradient descent, obtaining the minimized multi-stage loss function and the target model parameters that minimize the multi-stage loss function.
[0087] Specifically, the stochastic gradient descent algorithm (SGD) calculates the gradient of the loss function with respect to the model parameters and gradually updates the parameters along the direction of gradient descent to minimize the loss function. In the multi-stage loss function, SGD needs to optimize multiple tasks simultaneously (such as classification loss, localization loss, confidence loss, etc.). The gradients of the multi-task loss function may come from different tasks and have inconsistent directions, resulting in an unstable optimization process. In this embodiment, using the stochastic gradient descent algorithm can alleviate the problem of uneven sample numbers between different categories in the dataset and effectively avoid local optima.
[0088] In some alternative embodiments, the training method further includes: adjusting the learning rate of the neural network model in stages, where the training stages of the neural network model include a first stage, an intermediate stage, and a later stage. In the first stage, the learning rate gradually increases from a first threshold to a second threshold. In the intermediate stage, the learning rate remains unchanged at the second threshold. In the later stage, training is performed using the first threshold as the learning rate.
[0089] Specifically, gradually increasing the learning rate in the first stage can improve the situation of divergence caused by unstable gradients. In the middle of training, the model has found a relatively good parameter space. At this time, keeping the learning rate unchanged can ensure that the model performs stable iteration and optimization within this parameter space, avoiding training instability caused by changes in the learning rate. In the later stage of training, the model is close to the optimal solution. At this time, reducing the learning rate can enable the model to make more refined adjustments near the optimal solution, which helps to improve the final performance of the model.
[0090] For example, the batch size used in training is 64, the momentum is 0.9, and the decay rate is 0.0005. The learning rate training process is set as follows: For the first epoch, the learning rate is slowly increased from 10 -3 to 10 -2 . If the learning rate is set too high at the beginning, divergence will occur due to unstable gradients. Then, enter the intermediate stage, where training is performed for 75 epochs at 10 -2 , and then in the later stage, training is performed for 30 epochs at a learning rate of 10 -3 .
[0091] Embodiment 2:
[0092] Based on the above embodiments, this embodiment provides an application example. This embodiment provides a method for fault detection, which can be applied to an electronic device with communication, computing, and data storage capabilities. The specific process can be as follows Figure 3 As shown, it is applied to the scenario of automatic detection of railway freight car wall panel faults. First, image data is collected through a monitoring device to form a fault image training data set; then a suitable deep neural network object detection model algorithm is constructed, and the model is optimized to improve the detection accuracy. Then, the collected fault image data set is used to train the model so that the model obtains the automatic detection ability. Finally, the trained model is deployed to the railway freight online repair information system, and the crack faults are automatically detected in a timely manner through the images uploaded in real time by the image acquisition device, and the model parameters are updated regularly for iterative optimization. The specific process is as follows:
[0093] The annotation of the fault points of the railway freight car wall panel image is a basic task and needs to be annotated with the help of tools. In this embodiment, the labelme tool is used, which supports the marking of rectangles, polygons, circles, lines, and points. It supports the export of label files for semantic and instance segmentation in VOC format and COCO format. It has the advantages of being cross-platform, supporting multiple operating systems such as Linux, Mac OS, and Windows, and being easy to install and use.
[0094] The structure of the deep neural network object detection model is as follows: The backbone of the deep neural network object detection model is composed of stacked convolutional layers (conv.layers) and 2 subsequent fully connected layers. The numbers in the figure represent the size of the feature map and the size of the convolutional kernel. The trained model inputs the railway freight car wall panel images collected in real time, and the fault positions are automatically annotated after passing through the model. A neural network is an algorithm model that simulates the functions of human brain neurons to store and process information and can perform distributed parallel processing on information. It has the characteristics of non-linearity, parallel processing, and fault tolerance, which enables it to obtain information about the production process from data contaminated by noise. Its basic structural unit is a neuron, and the network layer is composed of multiple neurons. A neuron is the computing unit of an artificial neural network and has input, output, and computing functions. A convolutional layer is a network layer that contains convolutional calculations. Through the convolutional structure, a convolutional kernel of a given size slides on the input data at a certain stride to perform the dot product operation of matrices, forming feature extraction within the receptive field, and the new tensor obtained by the operation is used as the input of the next convolutional layer. Due to the introduction of the convolutional kernel, the computational amount of the network is greatly reduced. This phenomenon also plays a positive role in deepening the number of network layers.
[0095] The leaky rectifified linear function is used to activate after each convolutional layer, as shown below.
[0096]
[0097] The last layer of the model will predict class probabilities and bounding box coordinates.
[0098] Add batch normalization layers. Batch normalization layers can make the distribution of intermediate data in the network as normalized as possible, significantly accelerate the training of the network, and have a certain regularization ability, which can improve overfitting.
[0099] The calculation formula of the batch normalization layer is as follows:
[0100]
[0101]
[0102] Among them, m represents the convolutional kernel size, x i represents all trainable parameters on the same batch and channel, ε represents the variance estimation value constant (the value range is between 0 and 1, and can be adjusted according to the training results), τ represents the stretching parameter, and β represents the offset parameter.
[0103] Use the following multi-stage loss function L during training:
[0104]
[0105] Among them, S is the grid size, B is the number of detection boxes, indicates that the object appears in cell i, indicates that the object does not appear in cell i, indicates that the j-th bounding box predictor in cell i is responsible for this prediction, x i and y i are the coordinates of the predicted box, and are the coordinates of the labeled box, ω i and h i are the width and height of the predicted box, and are the width and height of the labeled box, C i is the confidence of the predicted box, p i (c) represents the probability of belonging to class c, γ, μ, and λ respectively represent hyperparameters, and classes is the set of preset fault classes.
[0106] Compared with the industrial technology that first extracts candidate regions and then uses a support vector machine for classification to identify faults, the optimized multi-stage loss function increases the degree of automation of railway freight car wall panel fault detection, omits the step of candidate region extraction, and can achieve end-to-end automatic identification of fault locations in one step.
[0107] The batch size used during training is 64, the momentum is 0.9, and the decay rate is 0.0005. The learning rate training process in this embodiment is as follows: For the first epoch, the learning rate is slowly increased from 10 -3 to 10 -2 . If the learning rate is set too high at the beginning, divergence will occur due to unstable gradients. After that, enter the intermediate stage, where it is trained for 75 epochs at a learning rate of 10 -2 , and then in the later stage, it is trained for 30 epochs at a learning rate of 10 -3 .
[0108] Meanwhile, this model selects the Stochastic Gradient Descent (SGD) algorithm to alleviate the problem of uneven sample numbers among different categories in the dataset. This method iteratively solves for the minimum value of the function step by step along the direction of gradient descent, thereby obtaining the minimized loss function and the model parameters that minimize it.
[0109] Nearly 5000 pieces of data of railway freight car monitoring fault annotation images are used to train and test the deep neural network object detection model, achieving good prediction accuracy. Finally, the trained model is deployed in the actual railway freight online repair information system to detect faults such as cracks in the railway freight car wall panels in real time, and the model parameters are updated regularly for iterative optimization.
[0110] The existing railway freight car wall panel fault detection technology is manual inspection technology, while this embodiment proposes an end-to-end automated fault detection technology. Based on railway monitoring image data, an object detection algorithm is installed in the railway online repair information system, and ideal real-time detection accuracy for faults such as cracks in the railway freight car wall panels is achieved. Experiments show that the proposed method can replace the traditional manual inspection method, has better detection efficiency, and can better ensure the freight safety of railways.
[0111] Compared with the traditional manual inspection method, the fault detection method proposed in this embodiment is a fast and accurate model that can be connected to online monitoring devices such as network cameras and maintain its real-time performance capabilities, including the time to obtain images from the camera and display the detection results. The model architecture of this embodiment is simple and can be directly trained on the entire image. At the same time, this method can also be well generalized to new fields, making it an ideal choice for applications where fast and robust object detection methods are required.
[0112] Embodiment 3:
[0113] Another embodiment of this application relates to a training device for a fault detection model. The implementation details of the training device for the fault detection model in this embodiment are specifically described below. The following content is only the implementation details provided for convenience of understanding and is not necessary for implementing this solution. The schematic diagram of the training device for the fault detection model in this embodiment can be asFigure 4 As shown in the figure, it includes a prediction module 410, a loss value calculation module 420, a model parameter calculation module 430, and a model parameter adjustment module 440.
[0114] The prediction module 410 is used to predict a preset fault point annotation image by using a pre-constructed neural network model to obtain a prediction result;
[0115] The loss value calculation module 420 is used to solve a multi-stage loss function according to the prediction result and the true annotation, and the multi-stage loss function includes at least two of a position loss term, a size loss term, a confidence loss term, and a classification loss term;
[0116] The model parameter calculation module 430 is used to solve the target model parameters that minimize the value of the multi-stage loss function;
[0117] The model parameter adjustment module 440 is used to adjust the neural network model according to the target model parameters to obtain a trained fault detection model.
[0118] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0119] In an optional embodiment, the neural network model is a convolutional neural network model, and further includes:
[0120] An activation module, which is used to use an activation function after each convolutional layer, and the calculation formula of the activation function is as follows:
[0121]
[0122] In an optional embodiment, the model parameter calculation module includes:
[0123] A stochastic gradient calculation unit, which is used to use the stochastic gradient descent algorithm to gradually iterate and solve the minimum value of the multi-stage loss function along the direction of gradient descent, to obtain the minimized multi-stage loss function and the target model parameters that minimize the multi-stage loss function.
[0124] In an optional embodiment, it further includes:
[0125] A learning rate adjustment module for adjusting the learning rate of the neural network model in stages, where the training stage of the neural network model includes a first stage, an intermediate stage, and a later stage. In the first stage, the learning rate gradually increases from a first threshold to a second threshold. In the intermediate stage, the learning rate remains unchanged at the second threshold. In the later stage, training is performed with the first threshold as the learning rate.
[0126] Embodiment 4:
[0127] Another embodiment of the present application relates to an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the training method of the fault detection model or the fault detection method in the above embodiments.
[0128] Among them, the memory and the processor are connected by a bus. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be an element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.
[0129] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when performing operations.
[0130] Embodiment 5:
[0131] Another embodiment of the present application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the training method of the fault detection model or the method embodiments of fault detection.
[0132] That is, those skilled in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0133] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in actual applications, various changes can be made to them in form and details without departing from the spirit and scope of the present application.
Claims
1. A method for training a fault detection model, characterized in that: include: Use the pre-built neural network model to predict the preset fault point annotated image and obtain the prediction result; Solving a multi-stage loss function according to the prediction results and the true annotations, wherein the multi-stage loss function includes at least two of a position loss term, a size loss term, a confidence loss term, and a classification loss term; Solving the target model parameters that minimize the value of the multi-stage loss function; The neural network model is adjusted according to the target model parameters to obtain a trained fault detection model.
2. The method for training a fault detection model according to claim 1, characterized in that: The calculation formula of the multi-stage loss function includes the following: Among them, L is the multi-stage loss function, S is the grid size, B is the number of detection boxes, Indicates that the object appears in cell i, Indicates that the object does not appear in cell i, denotes the jth bounding box predictor in cell i that is responsible for the prediction, x i and i is the coordinate of the prediction box, and is the coordinate of the annotation box, ω i and h i is the width and height of the prediction box, and is the width and height of the annotation box, C i is the confidence of the prediction box, p i (c) represents the probability of belonging to category c, and γ, μ, and λ represent hyperparameters.
3. The method for training a fault detection model according to claim 1, characterized in that: The neural network model is a convolutional neural network model, and the training method further includes: An activation function is used after each convolutional layer. The activation function calculation formula is as follows:
4. The method for training a fault detection model according to claim 3, characterized in that: The neural network model is provided with a batch normalization layer, and the batch normalization layer is located between the convolution layer and the activation function.
5. The method for training a fault detection model according to any one of claims 1 to 4, characterized in that: The solving of the target model parameters that minimize the value of the multi-stage loss function includes: By using the stochastic gradient descent algorithm, the minimum value of the multi-stage loss function is solved step by step and iteratively along the direction of gradient descent, so as to obtain the minimized multi-stage loss function and the target model parameters that minimize the multi-stage loss function.
6. The method for training a fault detection model according to claim 1, characterized in that: The training method further comprises: The learning rate of the neural network model is adjusted in stages, wherein the training stages of the neural network model include a first stage, an intermediate stage and a later stage, in the first stage, the learning rate is gradually increased from a first threshold value to a second threshold value, in the intermediate stage, the learning rate maintains the second threshold value unchanged, and in the later stage, the first threshold value is used as the learning rate for training.
7. A method for fault detection, characterized in that: The method comprises using a fault detection model trained by the fault detection model training method according to any one of claims 1 to 6 to identify a target image to be identified, thereby obtaining a fault detection result.
8. A training device for a fault detection model, characterized in that: include: A prediction module is used to predict the preset fault point annotated image using a pre-built neural network model to obtain a prediction result; A loss value calculation module, used to solve a multi-stage loss function according to the prediction result and the real annotation, wherein the multi-stage loss function includes at least two of a position loss term, a size loss term, a confidence loss term, and a classification loss term; A model parameter calculation module, used to solve the target model parameters that minimize the value of the multi-stage loss function; The model parameter adjustment module is used to adjust the neural network model according to the target model parameters to obtain a trained fault detection model.
9. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the training method of the fault detection model as described in any one of claims 1 to 6, or the fault detection method as described in claim 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for training a fault detection model described in any one of claims 1 to 6 or the method for fault detection described in claim 7 is implemented.