Railway container flatcar loading and unloading state detection method based on deep learning
Through the improved deep learning convolutional neural network and YOLOv5 model, combined with the CBAM attention mechanism, the precise identification of railway container loading and unloading status is achieved, solving the problem of fast, accurate and safe locking and unlocking during container loading and unloading, and improving efficiency and safety.
Patent Information
- Application Number
- CN202510399351.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art cannot meet the needs of fast, accurate and safe locking and unlocking during the loading and unloading of railway containers, and there are problems such as time-consuming and safety hazards of manual judgment.
The improved deep learning convolutional neural network is adopted, combined with the YOLOv5 model and the CBAM attention mechanism, and the object detection model is trained by field acquisition and labeling of image and video data, so as to achieve accurate identification of container locking and locking states.
It improves the efficiency and safety of the container loading and unloading process, reduces manual intervention, reduces labor costs, and meets the rapid response needs of railway transportation.
Smart Images

Figure CN120278978A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of railway transportation, and particularly to a method for detecting the loading and unloading state of a railway container flatcar based on deep learning. Background Art
[0002] As an assembly station and central hub for container transportation, in a railway freight yard, large cranes such as container gantry cranes (also known as gantry cranes) or container forklifts are required to load and unload containers as shown. The crane needs to be operated by a driver. For the ground, there need to be staff on-site to give commands, and to provide real-time feedback on the position of the container to ensure that the keyhole of the container can be aligned with the lock head of the F-TR lock during the locking process, and to ensure that the container does not get hooked to the railway flatcar during the unlocking process, thereby ensuring the safety of railway containers throughout the loading and unloading process. In order to ensure the rapid operation of the railway as a transportation hub, higher requirements are put forward for the container loading and unloading capacity and operation efficiency, which has led to the development of railway container machinery towards high speed and intelligence. The container locking device has evolved from the original straight platform lock, commonly known as the mushroom head, to the current fully automatic rotary lock and the F-TR lock with a unique eagle head structure. Although the unique eagle head design of the F-TR lock ensures the safety of the container during transportation, it requires manual judgment to determine whether the unlocking is successful during the unlocking process, which undoubtedly increases the time cost in the logistics transportation process.
[0003] When a container is being loaded and locked, it is much simpler compared to the container unlocking process. First, the operator of the crane operates according to the specified safety regulations, moving the crane or the spreader, or both at the same time, to lift the container above the vehicle with the specified vehicle number. When approaching the F-TR lock of the vehicle and about to start locking, perform low-speed jogging, with each jog moving up and down a few centimeters. If there is an offset, cooperate with the ground staff to correct the position of the container to ensure that the four keyhole positions on the bottom of the container are completely aligned with the four lock heads of the F-TR locks on the vehicle, and then lock. Due to the unique eagle head structure design of the F-TR lock, when the container is locked, it will rotate horizontally about 0.34° along the inclined plane of the lock head and then enter the lock head, and then rotate horizontally and recover about 0.34° along the inclined plane of the back of the lock head to return to the original position. As Figure 1 shown.
[0004] Regarding the container loading and unloading process in the railway context, usually, the staff operating the crane first operates the spreader to lift the container from the designated stacking area. There are multiple types of stacking area divisions, such as the stacking area for containers without goods and the stacking area for containers loaded with goods. The common sizes of containers also include 20-foot standard containers, 40-foot standard containers, 45-foot standard containers, etc. After the staff operating the crane lifts the designated container, they will operate the container to move to the designated loading position. When the container is approaching the vehicle and about to be locked, they will communicate with the ground staff in real time via walkie-talkie to avoid the risk of the container falling off due to the deviation between the container lock hole and the position of the F-TR lock head on the flat car, which may lead to the failure of locking. However, this method has an obvious drawback, that is, it takes a long time. Moreover, if the corner fittings are not fully locked and the container is still in a suspended state, the ground safety personnel are relatively close to the container when observing whether the locking is successful nearby, which poses a safety hazard.
[0005] Therefore, in the current logistics and transportation context, relying solely on manual judgment cannot meet the requirements that railway containers can be accurately, quickly, and safely locked to the designated position during each loading, with the offset error maintained within a certain range. It also cannot meet the requirements that when unlocking, each corner fitting can be successfully unlocked without getting hooked, or being able to give a warning when a hooking situation occurs but no loss is caused, thus unable to ensure transportation safety.
[0006] The existing literature "Tang Xin. Research on the Visual Alignment Technology of the Spreader and the Container Based on Deep Learning [D]. Southwest Jiaotong University, 2022." proposed a two-step method for determining the pixel center of the container lock hole. The first step is to detect the lock hole border from the container image, the second step is to segment the lock hole from the lock hole border, then calculate the center of the lock hole based on the lock hole border, and then establish the matrix relationship between the camera coordinate system and the spreader coordinate system according to the camera deployment, so as to determine the orientation of the spreader and the container and complete the automatic alignment of the spreader and the container.
[0007] The existing literature "Zhang Jun, Diao Yunfeng, Cheng Wenming, et al. Tracking and Centering of Container Lock Holes Based on Video Streams [J]. Journal of Computer Applications, 2019(z2): 216–220." proposed a method for tracking and positioning container lock holes based on video streams. First, the foreground of the first frame image of the video is extracted, and the part containing the container is extracted from it. Then, the sliding window based on the support vector machine is used to locate and extract the lock hole area, and then track it. And a container loading and unloading simulation experiment is carried out in the laboratory according to a certain ratio. The experimental result shows that the accuracy rate is as high as 99.65%.
[0008] In the above prior art, there are relatively few application examples of using deep learning convolutional neural networks to identify and detect the keyholes of railway containers. Even when using a deep learning convolutional neural network model for keyhole detection, the detection results cannot meet the detection requirements of real-time rapid response at the railway freight site and have not been widely applied to the actual working environment. Summary of the Invention
[0009] In order to overcome or alleviate one or more of the above technical problems, the object of the present invention is to provide a method for detecting the loading and unloading state of railway container flatcars based on deep learning. Through an improved deep learning convolutional neural network, the locking and unlocking states can be identified more accurately and quickly and respond rapidly.
[0010] The present invention provides the following technical solutions:
[0011] A method for detecting the loading and unloading state of railway container flatcars based on deep learning, which includes the following steps:
[0012] S1: Field-collect images and / or video data of different loading and unloading states of railway container flatcars as basic data, and annotate the basic data to separately produce target detection data sets for different states when the railway container is locked and unlocked, and divide the target detection data sets into training sets, validation sets and test sets;
[0013] S2: Train a target detection model through the YOLOv5 model convolutional neural network and the produced target detection data set. The target detection data set is used to train the target detection model for the locking and unlocking processes of the container to obtain a trained target detection model;
[0014] S3: On the basis of the YOLOv5 model, add a CBAM attention mechanism and add an attention module to the Backbone; retrain the trained target detection model obtained in step S2, and use the generated weights to verify the effect of target detection to obtain detection result information;
[0015] S4: On the basis of the detection result information obtained in S3, improve the model inference detection speed through partial quantization to obtain the final detection result; the partial quantization is to quantize each layer of the network separately, compare the decline in the precision and recall rate of the network before and after quantization for each layer, and record the positions of the sensitive layers with a larger decline, and then skip this layer during the partial quantization process.
[0016] According to some embodiments, step S1 specifically includes the following steps:
[0017] S11: The collected object detection dataset includes multiple images, multiple videos of railway container locking, and multiple videos of railway container unlocking. The dataset includes various types of distances, angles, heights, and saturations.
[0018] S12: When making annotations, for the locking state, the situation where the container corner fitting does not contact the F-TR lock eagle head is marked as the approaching mark, the mark of locking in progress is marked when the corner fitting contacts the lock eagle head, the mark of successful locking is marked when the corner fitting contacts the flat car and the lock eagle head is inside the corner fitting, and the mark of failed locking is marked when the corner fitting contacts the flat car but the lock is not inside the corner fitting; for the unlocking state, the mark of not separated is marked when the container contacts the flat car in the initial state, the mark of unlocking in progress is marked when the F-TR lock is inside the corner fitting and gradually detaches from the flat car, and the mark of successful unlocking is marked when the container corner fitting completely detaches from the F-TR lock.
[0019] S13: The training set, validation set, and test set are divided in the ratio of 7:1.5:1.5.
[0020] According to some embodiments, in step S2, for the YOLOv5s network model, the critical training times of the training times Epoch are set according to the convergence situation during the training process, and the pre-trained weights are selected as yolov5s.pt.
[0021] According to some embodiments, step S4 specifically includes the following steps:
[0022] S41: In each layer of the network, find the maximum absolute value |max| of the weight values, determine the mapping interval as [-|max|, +|max|], and linearly map the weight values within this interval to [-127, 127].
[0023] S42: During the activation value quantization process, YOLOv5 introduces a saturation mapping strategy to solve the problem of inconsistent distributions in each layer; the SiLu activation function is adopted, and the activation values include positive and negative values; by selecting values less than the maximum positive value and greater than the minimum negative value as the interval thresholds, the activation values are mapped to [-127, 127], and the values outside the interval are directly mapped to ±127; the method for determining the interval thresholds is: using the validation set samples to input the network, recording the activation value distributions layer by layer, and searching for suitable thresholds based on the data.
[0024] Compared with the prior art, the deep learning-based railway container flat car loading and unloading state detection method provided by the present invention has the following beneficial effects:
[0025] 1) High efficiency: In the data collection stage, automated detection equipment is used to judge the loading and unloading status of containers through real-time system feedback and monitoring, and can work continuously and stably under various adverse conditions, thus greatly improving the efficiency of data collection operations. By replacing manual identification with machine recognition, the manual labor is greatly reduced, and the locking and unlocking status is calculated by machine operation, with higher accuracy.
[0026] 2) Safety and stability: By using the method for detecting the loading and unloading status of railway container flatcars provided by the present invention, the loading and unloading status of containers can be accurately monitored and judged in different environments, reducing the errors introduced by human factors and avoiding the safety hazards existing in the on-site observation by staff during the loading and unloading of containers, ensuring the stability of the loading and unloading process.
[0027] 3) Reduction of labor costs: The application of automated loading equipment can reduce the dependence on labor, thereby reducing labor costs and saving a large amount of human resources for enterprises.
[0028] 4) Accuracy and speed: Adding an attention mechanism improves the detection accuracy, and greatly improves the detection speed of the model through partial quantization, enabling the automated equipment to quickly and accurately detect the loading and unloading status of containers.
[0029] Using machine vision to replace manual detection can also avoid the influence of subjective factors. For example, after working for a long time, the staff may be tired and unable to accurately judge the observed images with the naked eye. Moreover, the machine vision technology takes less time to judge the locking and unlocking of containers, meeting the growing transportation requirements. In addition, for freight stations applying machine vision technology, the maintenance efficiency can be significantly improved during regular maintenance, and at the same time, it provides safety guarantees for railway freight transportation. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 The movement trajectory of the F-TR lock when loading a container provided for the background technology of the present invention.
[0031] Figure 2a The Faster R-CNN detection process provided for the embodiment of the present invention.
[0032] Figure 2b The principle of the RPN network provided for the embodiment of the present invention.
[0033] Figure 2c The ResNet50 network structure provided for the embodiment of the present invention.
[0034] Figure 2d The Bottleneck module structure provided for the embodiment of the present invention.
[0035] Figure 3Schematic diagram of the SSD algorithm structure provided by the embodiments of the present invention.
[0036] Figure 4a YOLOv5 network provided by the embodiments of the present invention.
[0037] Figure 4b Focus structure provided by the embodiments of the present invention.
[0038] Figure 4c CSP structure provided by the embodiments of the present invention.
[0039] Figure 4d SPP module provided by the embodiments of the present invention.
[0040] Figure 4e FPN+PAN structure provided by the embodiments of the present invention.
[0041] Figure 4f Loss curve graph of the YOLOv5s model training set provided by the embodiments of the present invention.
[0042] Figure 5 Detection result graph of the Faster R-CNN model provided by the embodiments of the present invention.
[0043] Figure 6 Detection result graph of the SSD model lock detection provided by the embodiments of the present invention.
[0044] Figure 7 Detection result graph of the YOLOv5 model lock detection provided by the embodiments of the present invention.
[0045] Figure 8 Comparison of the detection results of three convolutional networks provided by the embodiments of the present invention.
[0046] Figure 9 Detection result graph of the Faster R-CNN model unlocking detection provided by the embodiments of the present invention.
[0047] Figure 10 Detection result graph of the SSD model unlocking detection provided by the embodiments of the present invention.
[0048] Figure 11 Detection result graph of the YOLOv5 model unlocking detection provided by the embodiments of the present invention.
[0049] Figure 12 Comparison of the detection results of three convolutional networks provided by the embodiments of the present invention.
[0050] Figure 13 CBAM model structure provided by the embodiments of the present invention.
[0051] Figure 14Comparison of the detection results of five attention mechanisms provided by the embodiments of the present invention.
[0052] Figure 15 Sample images of the dataset provided by the embodiments of the present invention.
[0053] Figure 16 Annotation of the container locking state provided by the embodiments of the present invention (approaching).
[0054] Figure 17 Annotation of the container locking state provided by the embodiments of the present invention (locking successful).
[0055] Figure 18 Flowchart of the method for detecting the loading and unloading state of railway container flatcars based on deep learning provided by the embodiments of the present invention. Detailed implementation manners
[0056] The present invention will be described in detail below in conjunction with the embodiments and the drawings. However, it should be understood that the embodiments and the drawings are only used for exemplary description of the present invention, and do not constitute any limitation to the protection scope of the present invention. All reasonable transformations and combinations within the scope of the inventive concept of the present invention fall within the protection scope of the present invention.
[0057] As Figure 18 , in order to accurately identify the loading and unloading state of railway container flatcars, the following steps provided by this embodiment are used for detection:
[0058] First step, field-collect images and / or video data of different loading and unloading states of railway container flatcars as the dataset to be tested, and respectively make object detection datasets for railway container locking and unlocking based on the annotations of the data to be tested.
[0059] The datasets all come from on-site collection at a certain railway freight yard. The shooting subjects are the locking and unlocking processes of the F-TR locks of railway containers, and the shooting equipment is a mobile phone. The captured image datasets are collected from different positions such as close range, medium range, and long range. In some photos, the lock heads of the F-TR locks have different degrees of wear, rust, marks, etc.
[0060] S11: The collected data includes 1,616 images, 10 videos of railway container locking, and 10 videos of railway container unlocking. The dataset contains various types of distances, angles, heights, and saturations, and the information content of the dataset is high. It has high representativeness and integrity as a training dataset for detecting the locking and unlocking states of railway containers.
[0061] S12: When using LabelImg for annotation, for the locking state, when the container corner fitting does not touch the F-TR lock eagle head, it is marked as "approaching"; when the corner fitting touches the lock eagle head, it is marked as "locking"; when the corner fitting touches the flat car and the lock eagle head is inside the corner fitting, it is marked as "locked successfully"; when the corner fitting touches the flat car but the lock is not inside the corner fitting, it is marked as "locking failed". For the unlocking state, when the container touches the flat car in the initial state, it is marked as "not separated"; when the F-TR lock is inside the corner fitting and gradually detaches from the flat car, it is marked as "unlocking"; when the container corner fitting completely detaches from the F-TR lock, it is marked as "unlocked successfully".
[0062] S13: The training set, validation set, and test set are divided in a ratio of 7:1.5:1.5.
[0063] Specifically, the collected image data is as Figure 15 shown. The annotation categories include unlocking state detection and locking state detection. For example, unlocking state detection includes three categories: not separated during unlocking, unlocking, and unlocked successfully. As Figures 16 - 17 , it is the annotation for different locking states.
[0064] In the second step, three convolutional neural networks were used for comparison to train the object detection model with the prepared dataset. The three deep learning convolutional neural network models are the R-CNN model (Faster R-CNN), the SSD model, and the YOLO v5 model. By comparing and analyzing the detection effects of the three, the optimal network was selected as YOLOv5. Finally, the YOLOv5 model was used to train the object detection model for the container locking and unlocking processes with the prepared dataset to be measured.
[0065] As Figure 2a shown, the main detection process using Faster R-CNN is as follows:
[0066] First, the image is input into the convolutional network model. Through multi-layer filtering and dimensionality reduction operations in the convolutional layer, the image data is transformed into a rich multi-level feature map. Next, in the process of generating candidate regions from the feature map, the RPN (Region Proposal Network) is applied to generate a series of anchor points on the feature map. Then, the softmax classifier is used to judge the region where each anchor point is located to determine whether the position is a foreground object or a background. Each anchor point represents a potential target box with a predefined size and ratio. After the classifier's judgment, a large number of non-target background regions are eliminated, and foreground candidate boxes with higher confidence are selected. Then, the ROI pooling operation is performed to further process each foreground candidate box, generating a feature vector of a fixed size, providing a standardized input for subsequent classification and localization. Finally, the generated feature vector is input into the fully connected layer, and the softmax layer outputs the probability distribution of each category, thereby accurately judging the category of the object within the candidate region. Then, the position of the candidate box is fine-tuned using bounding box regression to make it closely fit the object's edge, thereby improving the detection accuracy.
[0067] Compared with the sliding window method used in the initial search features, Faster R-CNN uses an efficient RPN network to directly generate detection candidate boxes from the entire image. The principle of the RPN network is as Figure 2b shown. This method not only reduces the generation of many redundant and useless candidate regions but also greatly reduces the calculation of the number of parameters, which greatly improves the detection speed and detectability of the model.
[0068] To achieve the detection of locking and unlocking of railway containers, a relatively high detection speed is required. Therefore, the ResNet50 network structure with a relatively fast running speed is selected as the backbone feature extraction network. This structure is as Figure 2c shown:
[0069] As can be seen from the above figure, there is a Conv1 convolution in the ResNet50 model architecture, and the size of the convolution kernel is 3×7×7. After the first convolution, four groups of Bottleneck module groups are carried out. The structure of the Bottleneck module is as Figure 2d shown; the finally obtained feature map is input into the fully connected layer for classification and recognition.
[0070] After carefully configuring the environment required for the experiment, a series of specific parameters for model training also need to be set. These parameters will cover aspects of training such as network structure, input size, training batch size, and number of iterations. The specific parameters are shown in Table 1:
[0071] Table 1 Faster R-CNN Model Training Parameter Settings
[0072]
[0073] Four different datasets, namely train, val, test, and trainval, were produced in this training, and their corresponding ratios are 8:1:1:9. The size of the images input into the network is 640×640×3.
[0074] The SSD algorithm was adopted, as Figure 3 is the structural schematic diagram of the SSD algorithm. The execution steps of the SSD algorithm can be roughly divided into the following three steps: First is the input processing stage. The image to be detected is first input into the network model, and the backbone network of the model will perform feature extraction operations on the information of the image. Second is the generation stage of default boxes. During the process of feature pyramid processing, the SSD algorithm extracts feature maps from 6 specific convolutional layers and constructs multiple default boxes of different scales based on these feature maps. Each default box will detect and classify the objects within the box through confidence scores. Finally is the non-maximum suppression stage, which is a key component of the output detection results. It will filter out the default boxes with high overlap and low confidence scores from the candidate boxes obtained from different feature level mappings, so as to ensure that the final output bounding boxes are accurate identifications and accurate classifications of the targets.
[0075] The parameters during model training are different, as shown in Table 2 specifically. The network model of the SSD object detection algorithm uses vgg16_reducedfc as the pre-trained weights for training. The size of the images input into the network is 640×640. The dataset selected for training is the same as the one used in Faster R-CNN training, and the division method is also the same.
[0076] Table 2 SSD Model Training Parameter Settings
[0077]
[0078] The YOLOv5 algorithm was adopted, as Figure 4a , at the input end of the YOLOv5 model, in order to improve the efficient utilization of the dataset and the optimization of the network detection performance, a variety of strategic methods were adopted. First, Mosaic data augmentation is performed on the dataset. This method randomly stitches four images into a large image, which significantly increases the diversity and complexity of the training samples, thus enriching the features learned by the model during the training stage and avoiding the occurrence of overfitting. Second, YOLOv5 adopts an adaptive anchor box mechanism. Specifically, when the network model predicts the bounding boxes of the targets, it will dynamically generate a set of appropriate preset anchor box sizes according to the size and proportion of the targets. Compared with the fixed anchor box size setting, this improvement enables the model to achieve more accurate matching when detecting the targets, which undoubtedly improves the detection accuracy and recall rate.
[0079] In addition, YOLOv5 also uses the adaptive image scaling technology. Its specific function is to uniformly adjust the input images of different sizes to a fixed size, which not only meets the structural requirements of the network but also ensures the reasonable allocation of computing resources, contributing to improving the running speed of the model. The calculation method for adaptively shrinking the input image to a fixed size of 640×640 is shown in Equations (3–1) and (3–2):
[0080]
[0081] C l = min(C l , C w ) (0 - 2)
[0082] In the formula, l represents the length of the image; w represents the width of the image; C l represents the proportionality coefficient of l to the target size length of 640; C w represents the proportionality coefficient of w to the target size width of 640.
[0083] If the longer side size of the image is l, then set C l as the scaling coefficient. After setting the scaling coefficient, multiply the longer side l and the shorter side w of the image by the scaling coefficient respectively, then calculate the difference between the shorter side and 640, and divide this difference by 32 to obtain the quotient q and the remainder r. Then, the minimum amount of black edge to be filled is up to 640×r. Compared with the traditional image scaling method that does not consider the characteristics of the network structure, adaptive resizing can significantly reduce the unnecessary black edge filling pixels, specifically reducing the filling amount by 640×32×q pixels. The reason for choosing the specific value of 32 is that the YOLOv5 network adopts five downsampling operations in the Backbone part, and each downsampling will halve the size. Therefore, the multiples of 2 to the 5th power are used to adjust the image edges. Only in this way can the size of the image always maintain an integer multiple relationship with the stride of the convolutional layer, thus avoiding accuracy loss or other potential problems caused by size mismatch.
[0084] In the design of the YOLOv5 network architecture, its core part uses CSPDarknet53 as the backbone feature extraction network, and also innovatively integrates the Focus and CSP structures to further improve the performance of the model. The Focus module structure is as Figure 4bAs shown in the figure, it first slices the feature map, then stitches together the pixel grids with the same values in the same feature layer, and then stacks different layers. Therefore, the Focus structure reshapes the original RGB three-channel, 640×640 resolution image into a 12-channel feature map with a size halved to 320×320 resolution. Subsequently, the feature map will immediately undergo another convolution operation to expand the number of channels to 32. In this way, the finally generated feature map not only realizes effective downsampling of the original image, but also ensures highly concentrated and rich feature representation, which is crucial for subsequent object detection tasks.
[0085] As the core feature extraction model of the YOLOv5 network, the design principle and structure of the CSP structure are as Figure 4c shown. The CSP structure divides the input features into two branches. One branch uses the residual network to refine the input information and reduce the computational complexity, and the other branch performs convolution operations to further compress the number of feature channels, approximately reducing it to about half of the original. Subsequently, the output results of the two branches are merged. This design inspiration comes from the CSPNet architecture. After such processing, the entire network is lightweight in terms of calculation, greatly saving computational resources, reducing the memory burden, and at the same time enhancing the network model's learning ability and expression ability for features. Therefore, the CSP structure is being widely used in modern object detection frameworks such as YOLOv5.
[0086] The CSP structure is divided into two categories, which are respectively applied to the Backbone part or the Neck part of the network. The YOLOv5 series of network models are further divided into four different versions according to their number of parameters and complexity, namely s, m, l, and x. The network depth of these versions increases gradually. Since this article has relatively high requirements for detection speed when detecting the loading and unloading status of railway containers, the more lightweight YOLOv5s network is adopted to meet the requirements of accurate and real-time detection. In the Neck part of the YOLOv5 network structure, the SPP module (Spatial Pyramid Pooling module) and the FPN+PAN (Feature Pyramid Network+Path Aggregation Network) structure are adopted. The structure of the SPP module is as Figure 4d shown.
[0087] The structure combining the FPN and PAN networks. This innovative design concept can greatly improve the performance of object detection and recognition. Its structure is as Figure 4eAs shown in the figure. The FPN structure mainly performs convolution and upsampling operations on the feature maps of deep layers with rich semantic information. This process can closely combine the feature maps of different layers, thereby achieving cross-level feature aggregation and improving the comprehensiveness and robustness of features. The PAN structure, on the other hand, performs feature fusion in a top-down manner. The shallow features are downsampled twice, and the resulting results are then concatenated and fused with the middle-level features, and then layer by layer downward. This operation not only retains the information of the deep abstract features but also introduces the spatial details of the shallow features, thus enhancing the correlation and complementarity between different levels. Finally, the multi-level feature maps optimized by both FPN and PAN are used as inputs and passed to the Head part for prediction and classification.
[0088] The training environment of the YOLOv5 model is the same as that of the SSD. The object detection model is also trained on the Pytorch deep learning platform. The specific training parameters are shown in Table 3.
[0089] Table 3 Training Parameter Settings of the Faster R-CNN Model
[0090]
[0091]
[0092] For the YOLOv5s network model, a total of 550 Epochs are set. The pre-trained weights are selected as yolov5s.pt, and the image size input to the network is uniformly set to 640×640. The training set, validation set, and test set are set according to 7:1.5:1.5. The total loss function curve Loss of YOLOv5 consists of three parts, as shown in Equation (3–3) specifically:
[0093] Loss = λ0loss box + λ1loss cls + λ2loss obj (0 - 3)
[0094] In the formula, loss box represents the loss of the prediction box adjustment parameter; loss cls represents the loss of classification; loss obj represents the confidence loss of whether the prediction box contains the object to be detected; λ0 represents the weight of the loss of the prediction box adjustment parameter, which is set to 0.05 during training; λ1 represents the weight of the loss of classification, which is set to 0.5 during training; λ2 represents the weight of the confidence loss, which is set to 1.0 during training. Among them, the calculation formula of loss box is shown in Equation (3–4):
[0095]
[0096] Wherein, S 2 represents the number of partitions of the feature map; B represents the number of predicted bounding boxes in each region; represents whether the corresponding predicted bounding box is a positive sample. Wherein, Loss CIOU is calculated as shown in Equation (3–5):
[0097]
[0098] Wherein, D2 represents the distance between the center points of the predicted bounding box and the target bounding box; D c represents the diagonal distance of the minimum circumscribed rectangle; α represents a balance parameter and does not participate in the gradient calculation; υ represents a parameter for measuring the consistency of the aspect ratio, and the calculation method is as shown in Equation (3–6); IOU represents the intersection over union of the predicted bounding box A and the labeled bounding box B, and the calculation method is as shown in Equation (3–7).
[0099]
[0100] Wherein, w gt and h gt respectively represent the width and height of the labeled bounding box; w and h respectively represent the width and height of the predicted bounding box.
[0101]
[0102] For the confidence loss loss obj part, by counting the predicted bounding boxes remaining after adjustment, score screening and non-maximum suppression in the region where the center point of the labeled bounding box is located, calculating the intersection over union of these predicted bounding boxes and the labeled bounding box, the predicted bounding boxes that meet the intersection over union threshold are used as positive samples, and the remaining predicted bounding boxes are used as negative samples, and the cross-entropy loss is calculated, and the formula is as shown in Equation (3–8):
[0103]
[0104] Wherein, p0 represents the confidence score of the positive sample predicted bounding box; p IOU represents the intersection over union of this positive sample predicted bounding box and the labeled bounding box; w obj represents the weight of the positive sample, which is set to 1.0 during training.
[0105] For the loss loss cls of classification, based on the positive samples obtained in the calculation of the confidence loss, obtaining the predicted results of the types of the predicted bounding boxes corresponding to each positive sample, and calculating the cross-entropy loss, and the formula is as shown in Equation (3–9):
[0106]
[0107] Wherein, c pThe classification prediction result represented as the positive sample prediction box; c gt The classification corresponding to the annotation box of this positive sample prediction box; w cls Represents the weight of the positive sample, which is set to 1.0 during training.
[0108] The loss function curve during the training of the YOLOv5 model is as Figure 4f shown. Analyzing this figure, it can be known that in the first 50 Epochs of model training, the total training loss curve drops sharply, which means that the model is in the rapid learning stage. As the Epoch gradually increases, the model will have better and better detection effects on the locking and unlocking detection of railway containers. After the training process reaches 50 Epochs, the decline of the total training loss curve tends to flatten, and after the 500th Epoch, this curve tends to be horizontal, that is, the model completes the convergence process.
[0109] Overall evaluate the training results according to the changes in precision and recall combined with the mAP value.
[0110] The analysis and detection of the locking process are as follows:
[0111] The detection results are respectively as Figures 5 - 7 , from Figure 8 it can be seen that all the result indicators of the YOLOv5 convolutional network are better than the other two network models.
[0112] It can be seen from the data in Table 4 that the YOLOv5 convolutional network is not only superior to the Faster R-CNN and SSD convolutional models in terms of detection accuracy and recall, but also consumes the least time in inference detection. This is the most important measurement standard for working scenarios that require real-time detection. The Faster R-CNN model can detect 7 pictures within one second, but the SSD model can detect and recognize 19 pictures within one second, and the YOLOv5 model can even recognize and detect 22 result pictures within one second.
[0113] Comparative analysis of the detection results of three classic convolutional neural networks in Table 4
[0114]
[0115] The comparison of the analysis of the unlocking process is as follows:
[0116] The unlocking detection results are shown in Figures 9 - 11 , from Figure 12 it can be seen that the YOLOv5 convolutional network is not only superior to the other two networks in terms of detection effect during locking, but also superior to the other two networks in terms of detection effect during unlocking, and its detection accuracy and recall are the highest among the three.
[0117] From the data in Table 5, it can be seen that among the three neural networks, namely Faster R-CNN, SSD, and YOLOv5, the YOLOv5 model has obvious advantages in both the detection of container locking and unlocking, and its detection speed is also the fastest among the three. Therefore, it can be concluded that the YOLOv5 convolutional neural network is more suitable for the detection of the loading and unloading status of containers in railway freight yards. Therefore, in this embodiment, the YOLOv5 convolutional network is finally selected as the detection model for railway container locking and unlocking.
[0118] Comparative Analysis of Detection Results of Three Classic Convolutional Neural Networks in Table 5
[0119]
[0120] Step 3: Based on the selected YOLOv5 network model, add an attention mechanism, and then perform the training process respectively. Use the generated optimal weights to verify the effect of object detection.
[0121] The main feature of the CBAM (Convolutional Block Attention Module) architecture is to integrate two core components, namely the channel attention module and the spatial attention module, into one model. Its model structure is as Figure 13 shown. The channel attention module uniformly screens all input channels and then finds important feature information in each channel; the spatial attention module aims to obtain and strengthen the significant features in the two-dimensional spatial distribution, regardless of whether these features are distributed in the time series or the spatial structure. Through the synergistic effect of these two modules, each branch of CBAM can extract the most valuable information from the channel and spatial structures, which greatly improves the overall expressiveness of the network and the effectiveness of feature extraction. Also, because CBAM integrates both spatial and channel attention modules, compared with the SENet architecture that only focuses on channel information, CBAM can better improve the optimization processing of features and thus obtain more significant results.
[0122] The attention module is added to the Backbone. Five attention mechanisms, namely SE, ECA, CBAM, GAM, and CA, are respectively added to the Backbone of the original YOLOv5 model. By systematically optimizing and enhancing the information in the three dimensions of channel, space, and position of the feature map, the understanding and perception ability of the model for complex scenes are improved, thereby enhancing the performance of the model and the accuracy of the object detection algorithm. Experiments were conducted on the self-made railway container locking dataset for the five attention mechanism modules.
[0123] After detection and evaluation, as Figure 14Comparison of detection results. The YOLOv5 model with the attention mechanism added has a slightly increased detection accuracy and recall rate compared to the previous model without the attention mechanism, except that the improvement in the detection accuracy and recall rate of the model with the GAM attention mechanism is not obvious. Among the models with improved detection accuracy, the two models with the SE and CBAM attention mechanisms are more prominent in terms of detection accuracy and recall rate when detecting the locked and unlocked states of railway containers compared to the other models.
[0124] Table 6 Comparison of results of different attention mechanism modules
[0125]
[0126] As can be seen from the data in Table 6, the two models with the SE and CBAM attention mechanisms not only have improved detection accuracy but also the operation of adding the attention mechanism does not reduce the detection speed of the model. Comparing the two models with the SE and CBAM attention mechanisms added, the model with the CBAM attention mechanism takes the least time for detection and can detect more images per unit time.
[0127] The results show that the model with the CBAM attention mechanism added has the best detection effect. Therefore, the method of adding the CBAM attention mechanism is selected to improve the detection accuracy of the model.
[0128] Step 4: Further improvement is made on the basis of adding the CBAM attention mechanism to the YOLOv5s model. The detection speed of the model is improved by comparing the two methods of global quantization and partial quantization. Global quantization quantizes all layers in the network, while partial quantization quantizes each layer of the network separately. The decrease in the precision and recall rate of the network before and after quantization of each layer is compared, and the positions of the sensitive layers with a larger decrease are recorded. Subsequently, these layers are skipped during the partial quantization process.
[0129] The specific steps of partial quantization are as follows:
[0130] S41: In each layer of the network, find the maximum absolute value |max| of the weight values, determine the mapping interval as [-|max|, +|max|], and linearly map the weight values within this interval to [-127, 127];
[0131] S42: During the activation value quantization process, to address the issue of inconsistent distributions across layers, YOLOv5 introduced a saturation mapping strategy. The SiLu activation function is used, and the activation values include positive and negative values. By selecting values less than the maximum positive value and greater than the minimum negative value as interval thresholds, the activation values are mapped to [-127, 127], and values outside the interval are directly mapped to ±127. Method for determining interval thresholds: Use the validation set samples to input into the network, record the activation value distributions layer by layer, and search for appropriate thresholds based on the data.
[0132] Experiments have shown that the detection speed of the globally quantized model has increased significantly, but the detection accuracy has decreased significantly. The detection speed of the partially quantized model has increased by approximately 70%, and its detection accuracy has only decreased slightly, achieving a balance between detection speed and detection accuracy. Therefore, the optimization result of the partial quantization is much better than that of the global quantization in terms of performance, not only improving the detection efficiency but also ensuring the detection accuracy.
[0133] The final detection results will be presented to the crane operator, facilitating the crane operator to better debug and control the on-site environment in the cab and accurately assist in determining the loading and unloading status.
[0134] The above embodiments are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.
Claims
1. A method for detecting the loading and unloading status of railway container flat cars based on deep learning, characterized in that: It includes the following steps: S1: Field-collect images and / or video data of different loading and unloading states of railway container flatcars as basic data, and label the basic data to separately produce object detection data sets for different states when railway containers are locked and unlocked, and divide the object detection data sets into training sets, validation sets and test sets; S2: Train an object detection model through the convolutional neural network of the YOLOv5 model and the produced object detection data sets. The object detection data sets are used to train the object detection model for the process of container locking and unlocking to obtain the trained object detection model; S3: On the basis of the YOLOv5 model, add the CBAM attention mechanism and add the attention module to the Backbone; Retrain the trained object detection model obtained in step S2, and use the generated weights to verify the effect of object detection to obtain detection result information; S4: On the basis of the detection result information obtained in S3, improve the model inference detection speed through partial quantization to obtain the final detection result; the partial quantization is to quantize each layer of the network separately, compare the decrease in the precision and recall rate of the network before and after quantization of each layer, record the positions of the sensitive layers with a large decrease, and then skip this layer during the partial quantization process.
2. The method for detecting the loading and unloading state of a railway container flatcar based on deep learning according to claim 1, wherein: Step S1 specifically includes the following steps: S11: The collected object detection data set contains multiple images, multiple videos of railway container locking, and multiple videos of railway container unlocking. The data set contains various types of distances, angles, heights, and saturations; S12: When making annotations, for the locked state, mark the situation where the container corner fitting does not contact the F-TR lock eagle head as the approaching mark, mark the contact between the corner fitting and the lock eagle head as the locking mark, mark the successful locking mark when the corner fitting contacts the flatcar and the lock eagle head is inside the corner fitting, and mark the failed locking mark when the corner fitting contacts the flatcar but the lock is not inside the corner fitting; for the unlocking state, mark the non-separation mark when the container contacts the flatcar in the initial state, mark the unlocking mark when the F-TR lock is inside the corner fitting and gradually detaches from the flatcar, and mark the successful unlocking mark when the container corner fitting completely detaches from the F-TR lock; S13: The training set, validation set and test set are divided in a ratio of 7:1.5:1.
5.
3. The method for detecting the loading and unloading state of railway container flat cars based on deep learning according to claim 1, wherein: In step S2, for the YOLOv5s network model, set the critical training times of the training times Epoch according to the convergence situation during the training process, and select yolov5s.pt as the pre-training weight.
4. The method for detecting the loading and unloading state of a railway container flatcar based on deep learning according to claim 1, wherein: Step S4 Specifically includes the following steps: S41: In each layer of the network, find the absolute value maximum |max| of the weight value, determine the mapping interval as [-|max|, +|max|], and linearly map the weight values within this interval to [-127, 127]; S42: During the activation value quantization process, YOLOv5 introduces a saturation mapping strategy to solve the problem of inconsistent distributions across layers; the SiLu activation function is adopted, and the activation values include positive and negative values; by selecting a value less than the maximum positive value and greater than the minimum negative value as the interval threshold, the activation values are mapped to [-127, 127], and values outside the interval are directly mapped to ±127; the method for determining the interval threshold is as follows: use the validation set samples to input the network, record the activation value distribution layer by layer, and search for appropriate thresholds based on the data.