Train open wagon door slot identification method
By improving the YOLOv5 network, replacing the Conv in the backbone network with SCConv, adding the RepViT module and using the SIoU loss function, the problem of difficulty in identifying door gaps in open trains is solved, and efficient and accurate door gap identification effect is achieved.
Patent Information
- Application Number
- CN202510252512.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-29
AI Technical Summary
In the prior art, it is difficult to identify the door cracks of open trains, resulting in leakage during coal transportation, affecting the accurate positioning and identification effect of the rubber-spraying sealing robot.
The improved YOLOv5 network is adopted, and by replacing the Conv convolution in the backbone network into the SCConv module, adding the RepViT module in the neck network, and using the SIoU loss function in the head network, building a train open door slot recognition model, performing image preprocessing and feature fusion, and improving recognition accuracy and robustness.
Effectively compress the calculation cost, improves the recognition efficiency and accuracy of open door cracks in trains, reduces calculation redundancy, enhances the recognition ability in complex environments, and improves the recognition speed and accuracy.
Smart Images

Figure CN120388333A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of identifying the gaps of open train car doors, and specifically to a method for identifying the gaps of open train car doors. Background Art
[0002] As an important member of energy sources, there are significant geographical differences in the supply and demand of coal in China, which determines that the main coal supply in China is mainly by railway transportation. Due to the poor sealing performance of the gaps of open train car doors, during the transportation of coal, the vibration of the carriages causes coal powder to leak from the edges of the gaps, resulting in a large amount of waste of coal resources. Therefore, it is crucial to ensure the accurate positioning of the gap by the spray adhesive sealing robot.
[0003] However, there are also many problems during the spray adhesive process of the spray adhesive sealing robot. First, when the train enters the coal loading area of the coal mine, it will move forward at a uniform speed to load coal, and there are relatively strict requirements for the recognition speed of the open car door gaps by the robot; second, since the color of the side door of the open car is the same as that of the open car, and there will be residual sealant around the gap after the sealant treatment of the gap, both will affect the direct recognition of the gap. Summary of the Invention
[0004] In order to solve the above problems, the present invention provides a method for identifying the gaps of open train car doors.
[0005] The present invention adopts the following technical solutions: A method for identifying the gaps of open train car doors, including: S1: Collect images of the side doors of open cars in different states, preprocess the images, label the door shafts in the images to form the required data set, and divide it into a training set and a test set; S2: Construct a model for identifying the gaps of open train car doors; S3: Adjust the parameters of the model for identifying the gaps of open train car doors, and set the size of the input image, the number of recognition types, and the number of iterations; S4: Test and evaluate the trained model for identifying the gaps of open train car doors; S5: Use the trained model for identifying the gaps of open train car doors to identify the gaps in the image, and calculate the position information of the gaps and draw the position of the gaps through the relative position information between the door shaft and the gaps and the predicted position information of the door shaft.
[0006] In some embodiments, step S1 includes: S11: Collect data: Use an industrial camera to collect the overall image of the side door of the open car and images of the door shaft at different angles; S12: Image preprocessing: Screen and remove incomplete and unclear images, and adjust the image size; S13: Image annotation: Label all the pictures of the door hinge at various angles in the pre - processed image; S14: Perform image enhancement; S15: Dataset division: After image enhancement, divide all the images into a training set and a test set according to a certain proportion.
[0007] In some embodiments, the train open - door gap recognition model includes: A backbone network, and the backbone network includes: A replaced SCConv, four C3 layers connected alternately with the replaced SCConv, and an SPPF layer; the input of the backbone network is the dataset formed after the previous pre - processing, then multi - level features are extracted from the input data and transformed into three different - level feature maps; A neck network, and the neck network includes: Four one - dimensional convolutional layers Conv, four up - sampling layers upsample, four C3 layers, four Concat layers, and a RepViT module; the inputs of the neck network are the three different - level feature maps output by the backbone network. After fusing the feature maps of different sizes in the neck network, the transmission of different - size information is enhanced, and the accuracy and robustness of object detection are improved. Finally, three new different - level feature maps are output to the head network; A head network, and the head network includes: The head network includes three convolutional layers and three detection heads. The inputs of the head network are the three different - level feature maps passed in by the neck network; the convolutional layers adapt to different detection tasks by adjusting the number of channels and kernel size, and the detection heads implement object classification and regression tasks, retain the bounding box with the highest confidence, and output the processed bounding box and confidence score; The loss function in the detection head is the SIoU loss function.
[0008] In some embodiments, the SCConv module includes: A spatial reconstruction unit SRU, which inputs the feature X and outputs a spatially refined feature ; A channel reconstruction unit CRU, which operates on the spatially refined feature to obtain a channel - refined feature Y, so as to reduce the redundancy between intermediate feature maps and improve the feature representation of the convolutional neural network.
[0009] In some embodiments, the calculation process of the spatial reconstruction unit SRU is: Normalize the input feature X by subtracting the mean μ and dividing by the standard deviation σ; Pass the result through the sigmoid function with relevant weights The weight values of the reweighted feature map are mapped to the range (0, 1) and gated by a threshold. The weights above the threshold are set to 1 to obtain the informative weight W1, while they are set to obtain the non-informative weight W2; Multiply the input feature X by W1 and W2 respectively to obtain two weighted features: the feature-rich feature and the feature with less information ; has rich and expressive spatial content, while has little information and is regarded as redundant.
[0010] In some embodiments, the calculation process of the channel reconstruction unit CRU is as follows: For a given spatially refined feature , divide the channels of X W into two parts according to the segmentation ratio, where one part has αC channels and the other part has (1 - α)C channels, where is the segmentation ratio; Use 1×1 convolution to compress the channels of the feature map to improve the calculation efficiency, and use the squeezing ratio r to spatially refine the feature X W the upper part X UP and the lower part X low ; X UP is input into the up-conversion stage. For X UP , use the convolution operations GWC and PWC to extract high-level representative information and reduce the calculation cost; at the same time, use the k×k GWC and 1×1 PWC operations on the same X UP , and then add the outputs to obtain the combined representative feature map Y1; X low is input into the bottom conversion stage. For X low , use the 1×1 PWC operation to generate a feature map with shallow hidden details, and connect the generated feature map and the reused feature X low to form the output Y2 of the bottom stage; Combine the output features Y1 and Y2 of the up-conversion stage and the down-conversion stage to obtain the channel-refined feature Y.
[0011] In some embodiments, combining the output features Y1 and Y2 of the up-conversion stage and the down-conversion stage includes: Collect the height, width, and convolution kernel of the output features Y1 and Y2 through global average pooling, and calculate the global spatial information , m = 1, 2; Use the calculated global spatial information to calculate the attention weights β1 and β2 of the output features Y1 and Y2 under the channel attention operation; Merge the upper feature Y1 and the lower feature Y2 in a channel manner to obtain the channel-refined feature Y.
[0012] In some embodiments, the calculation process of the RepViT module is as follows: Divide the input image into several uniformly sized P×P blocks, which represent different parts of the image. Then, transform each image block into an embedding vector of a fixed dimension through a linear transformation. These vectors form a token sequence, and position encoding is added to the embedding vectors to obtain the positions of the token sequence.
[0013] In some embodiments, step S3 includes: Before training, configure the parameters of the model. The image size imgsz = [640, 640], the confidence threshold conf_thres = 0.25, the IoU threshold iou_thres = 0.45, the initial learning rate lr0 = 0.01, the weight decay coefficient weight_decay = 0.0005, and the number of iterations epochs = 200.
[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention identifies the side door shaft of the gondola car, obtains the minimum and maximum coordinates of the side door shaft, and then finally obtains the position of the door gap by using the relative position relationship between the door gap and the side door shaft.
[0015] 2. The present invention replaces the Conv convolution with the SCConv module in the backbone network of the traditional YOLOv5 to compress the redundant features in the convolutional neural network, reduce the computational load, and improve the model performance. First, SCConv can effectively compress the model and reduce the computational cost by reducing the spatial and channel redundancy between features, making the identification of the side door shaft more efficient. Second, SCConv can improve the performance of the neural network and enhance the ability to identify the door gap in complex environments by optimizing the feature representation.
[0016] 3. The present invention adds the RepVit module to the neck network of the traditional YOLOv5. By introducing parametric mapping, it not only maintains the global modeling ability but also utilizes the local perception characteristics of the CNN, significantly improving the model's understanding of image features and reducing the computational cost to a certain extent.
[0017] 4. The present invention replaces the traditional loss function IoU in the head network of the traditional YOLOv5 with SIoU. SIoU only requires one regression coordinate (x or y), which effectively reduces the total number of degrees of freedom and can effectively improve the training speed and inference accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a step description diagram of the method of the present invention; Figure 2 is the overall framework flowchart of the present invention; Figure 3 is the structural diagram of the SCConv module of the present invention; Figure 4 is the structural diagram of the RepVit module of the present invention; Figure 5 is the detection result of the present invention. Detailed implementation manners
[0019] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than limiting the invention. Additionally, it should be noted that only the parts related to the relevant invention are shown in the drawings for the convenience of description.
[0020] It should be noted that, without conflict, the features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below with reference to the drawings and embodiments.
[0021] The specific implementation of the train open - door gap recognition method includes the following steps: S1. Collect the side - door images of open - top wagons in different states at the nearby coal - mine loading area, and pre - process the images. Screen out the images with poor quality, enhance the remaining images, label the side - door hinges in the images, form the required data set, and divide it into a training set and a test set; The specific situation of constructing the side - door hinges is as follows: S11: Data collection: Use an industrial camera to collect the overall images of the open - top wagon side - doors and the images of the hinges at various angles at the nearby coal - mine loading area; including the overall images and the images of the hinges at various angles collected at different time periods due to different lighting conditions, increasing the diversity of the data and improving the robustness of the network.
[0022] S12: Image pre - processing: After data collection, screen out the incomplete and blurred images, such as severely over - exposed and severely blurred images, to prevent images with too poor quality from having a negative impact on the network, and adjust the image size to 640×640 pixels; S13: Image annotation: In this embodiment, only the side - door hinges are annotated. The method is to use the LabelImg software to annotate all the images with hinges in the pre - processed images and name them spindle; S14: Image enhancement: Image rotation, image translation, image flipping, image stretching in geometric transformation, random cropping, central cropping in image cropping, image noise injection, and image brightness enhancement; S15: Dataset Division: After image enhancement, all the images formed after image processing and the corresponding annotation files are divided into a training set and a test set according to the ratio of 8:2.
[0023] S2: Construct a recognition model for the gap of train car doors.
[0024] The recognition model for the gap of train car doors includes: A backbone network, which includes: A replaced SCConv, four C3 layers connected alternately with the replaced SCConv, and an SPPF layer; the input of the backbone network is the dataset with a size of 640×640 formed after the previous preprocessing, and then multi-level features are extracted from the input data (the features include low-level detailed information such as edges and textures; high-level semantic information such as the shape and position of objects), and are transformed into three different-level feature maps. A neck network, which includes: Four one-dimensional convolutional layers Conv, four upsampling layers upsample, four C3 layers, four Concat layers, and a RepViT module; the inputs of the neck network are three different-level feature maps output by the backbone network. After fusing the feature maps of different sizes in the neck network, the transmission of information of different sizes is enhanced, and the accuracy and robustness of object detection are improved. Finally, three new different-level feature maps are output to the head network. A head network, which includes: The head network includes three convolutional layers and three detection heads. The inputs of the head network are three different-level feature maps passed in by the neck network; the convolutional layers adjust the number of channels and kernel size to adapt to different detection tasks, and the detection heads implement object classification and regression tasks, retain the bounding box with the highest confidence, and output the processed bounding box and confidence score. The loss function in the detection head is the SIoU loss function.
[0025] Among them, the SCConv module includes: A spatial reconstruction unit SRU, which inputs the feature X and outputs a spatially refined feature ; A channel reconstruction unit CRU, which operates on the spatially refined feature to obtain a channel-refined feature Y, so as to reduce the redundancy between intermediate feature maps and improve the feature representation of the convolutional neural network.
[0026] The calculation process of the spatial reconstruction unit SRU is: Normalize the input feature X by subtracting the mean μ and dividing by the standard deviation σ; The weight values of the feature map reweighted by the relevant weights are mapped to the range (0, 1) through the sigmoid function and gated by a threshold. The weights above the threshold are set to 1 to obtain the informative weight W1, while they are set to obtain the non-informative weight W2; Multiply the input feature X by W1 and W2 respectively to obtain two weighted features: the feature-rich feature and the feature with less information ; ; has rich and expressive spatial content, while has little information and is regarded as redundant.
[0027] The calculation process of the channel reconstruction unit CRU is as follows: For the given spatially refined feature , divide the channels of X W into two parts according to the splitting ratio, where one part has αC channels and the other part has (1 - α)C channels, where is the splitting ratio; Use 1×1 convolution to compress the channels of the feature map to improve the calculation efficiency, and use the squeezing ratio r to spatially refine the feature X W the upper part X UP and the lower part X low ; X UP is input into the up-conversion stage. For X UP , use the convolution operations GWC and PWC to extract high-level representative information and reduce the calculation cost; at the same time, use the k×k GWC and 1×1 PWC operations on the same X UP , and then add the outputs to obtain the combined representative feature map Y1; X low is input into the bottom-conversion stage. For X low , use the 1×1 PWC operation to generate a feature map with shallow hidden details, and connect the generated feature map and the reused feature X low to form the output Y2 of the bottom stage; Combine the output features Y1 and Y2 of the up-conversion stage and the bottom-conversion stage to obtain the channel-refined feature Y.
[0028] Combining the output features Y1 and Y2 of the up-conversion stage and the bottom-conversion stage includes: Collect the height, width, and convolution kernels of the output features Y1 and Y2 through global average pooling to calculate the global spatial information , m = 1, 2; Use the calculated global spatial information to calculate the attention weights β1 and β2 of the output features Y1 and Y2 under the channel attention operation; Combine the upper feature Y1 and the lower feature Y2 in a channel manner to obtain the channel-refined feature Y.
[0029] Among them, the calculation process of the RepViT module is as follows: Divide the input image into several uniformly sized P×P blocks, which represent different parts of the image. Then, transform each image block into an embedded vector of a fixed dimension through a linear transformation. These vectors form a token sequence. Add positional encoding to the embedded vectors to obtain the positions of the token sequence.
[0030] Based on the standard Transformer, RepViT introduces reparameterization, changing multiple matrix multiplications in the original calculation of sub-attention to multiple weight matrices in the combined attention calculation. Therefore, the calculation of the self-attention mechanism in the encoder of RepViT reduces the computational complexity and speeds up the non-linear transformation of each token.
[0031] After being processed by the encoder layer, RepViT obtains the global representation of the image through pooling and finally sends it to the task-specific head.
[0032] The head network consists of three convolutional layers and three detection heads. The inputs of the head network correspond to three different levels of feature maps input from the neck network respectively. The convolutional layers adapt to different detection tasks by adjusting the number of channels and the kernel size. The detection heads implement the object classification and regression tasks, and measure the overlap degree between the bounding boxes more precisely through the SIoU operation, retain the bounding box with the highest confidence, and output the processed bounding boxes and confidence scores. The SloU loss function redefines the angular penalty metric considering the angle of the vector between the expected regressions on the basis of the IoU loss function, enabling the prediction box to quickly drift to the nearest axis. Subsequently, only one coordinate needs to be regressed, effectively reducing the total degrees of freedom and improving the inference accuracy.
[0033] This embodiment discloses a method for identifying the door gap of a train open car based on the improved YOLOv5. The YOLOv5 network in it is divided into a backbone network, a neck network, and a head network; the Conv convolution in the backbone network is replaced by the SCConv module, the RepVit module is added to the neck network, and the loss function IoU in the head network is replaced by SIoU to form an improved YOLOv5 network conducive to the identification of the train open car door gap; The spatial and channel reconstruction convolution SCConv for feature redundancy is introduced to compress CNN using the spatial and channel redundancy between features, obtain better performance by reducing redundant features, and significantly reduce complexity and computational cost. SCConv consists of two sequentially placed units: a spatial reconstruction unit SRU and a channel reconstruction unit CRU. Specifically, for an input feature X, the spatially refined feature is first obtained through the SRU operation , and then the channel-refined feature Y is obtained using the CRU operation to reduce the redundancy between intermediate feature maps and improve feature representation.
[0034] The aforementioned spatial reconstruction unit SRU uses a split-reconstruction operation to separate the feature map with rich information from the feature map with less information corresponding to the spatial content, and uses the scaling factor in group normalization to evaluate the information content of different feature maps.
[0035] Its algorithm process is as follows: First, an intermediate feature map is given , where N is the batch axis, C is the channel axis, and H and W are the spatial height and width axes.
[0036] First, the input feature X is normalized by subtracting the mean μ and dividing by the standard deviation σ as follows: where μ and σ are the mean and standard deviation of X, is a small positive number added to stabilize the division, and γ and β are trainable affine transformations.
[0037] The relevant weights of normalization are given by the following formula: Then, the weight values of the feature map reweighted by are mapped to the range (0, 1) through the sigmoid function and gated by a threshold. The formula for obtaining W is as follows: Finally, the input feature X is multiplied by W1 and W2 respectively to obtain two weighted features: the feature with rich information and the feature with less information. has rich and expressive spatial content, while has little information and is regarded as redundant.
[0038] The channel reconstruction unit (CRU) further reduces the redundancy of the spatial refinement feature map in the channel dimension by using the "split - transform - fuse" strategy. In addition, the CRU extracts rich representative features through lightweight convolution operations, while using inexpensive operations and feature reuse schemes to process redundant features. It is divided into three steps: Split, Transform, and Fuse.
[0039] First, for the given spatial refinement feature , first divide the channels of X W into two parts, one part has αC channels, and the other part has (1 - α)C channels. Subsequently, 1×1 convolution is used to compress the channels of the feature map to improve computational efficiency, and the spatial refinement feature X W the upper part X UP and the lower part X low .
[0040] Next, X UP is input into the up - conversion stage. Efficient convolution operations GWC and PWC are used to replace k×k convolution to extract high - level representative information while reducing the computational cost. Due to sparse convolutional connections, GWC reduces the number of parameters and computations, but cuts off the information flow between channel groups. While PWC makes up for the loss of information and helps information flow between feature channels. Therefore, k×k GWC and 1×1 PWC operations are used on the same X UP , and then the outputs are added together to obtain the combined representative feature map Y1.
[0041] The up - conversion stage can be expressed as: where, , are the learnable weight matrices of GWC and PWC; and are the input and output feature maps of the up - conversion respectively.
[0042] Meanwhile, X low is input into the bottom - conversion stage. 1×1 PWC operation is applied to generate a feature map with shallow hidden details as a supplement to the rich feature extractor. Finally, the generated and reused feature X low are concatenated to form the output Y2 of the bottom stage, and the calculation formula is as follows: where, is the learnable weight matrix of PWC, ∪ is the concatenation operation, and are the bottom - input and output feature maps respectively.
[0043] Finally, the simplified SKNet method is used to adaptively merge the output features Y1 and Y2 of the up-conversion stage and the down-conversion stage. First, global average pooling is applied to collect the global spatial information S m Then, the global channel descriptors S1 and S2 of the upper part and the lower part are stacked together, and channel attention operations are used to generate the feature importance vectors β1 and β2. Then, the upper feature Y1 and the lower feature Y2 are merged in a channel-wise manner to obtain the channel-refined feature Y. The calculation formula is as follows: The RepViT module introduced in the neck network of the yolov5 network can provide stronger visual feature expression ability than before. And the optimization of the computational efficiency of the RepViT module helps to speed up the object detection, thus improving the performance of object detection. The RepViT module has four stages, and its algorithm process is as follows: Given an intermediate feature map with dimensions B×3×H×W, where B is the batch size, 3 refers to the channel dimension, and H and W are the width and height of the input image. First, the input image is processed through a continuously stacked 3×3 convolution as the first stage (i.e., stem) of the input processing stage, reducing the latency of the backbone; both the second stage and the third stage adopt the method of a single and deeper downsampling layer. First, a 1×1 convolution is used to adjust the channel dimension of the output of the first stage, and then the input and output of two 1×1 convolutions are connected through a residual connection to form a feed-forward network. At the same time, a RepViT block is added in front to further deepen the downsampling layer. The fourth stage adds a RepViTSE block in front of the downsampling stage adopted in the second and third stages; after the four stages are run, it passes through a classifier composed of a global average pooling layer followed by a linear layer.
[0044] Using SloU instead of the traditional IoU as the loss function of the head network, the SloU loss function, based on the IoU loss function, considers the angle between the vectors of the expected regression, redefines the angle penalty metric, can make the prediction box quickly drift to the nearest axis, and then only needs to regress one coordinate, effectively reducing the total number of degrees of freedom, and can improve the training speed and inference accuracy.
[0045] S3. Adjust the parameters of the train open door crack recognition model, set the size of the input image of the convolutional neural network, the number of recognition categories, the number of iterations, etc.; Specific process: Adjust the parameters of the improved network model. According to the size of the computer's memory and video memory, the required recognition effect and training speed, set the size of the input image of the convolutional neural network, the number of recognition categories, and the number of iterations. And the user needs to use a graphics card type that supports CUDA acceleration.
[0046] This invention conducts experiments based on the versions of PyTorch 1.10.1 and Python 3.8. When conducting experiments, it uses the GPU of NVIDIA GeForce RTX3060 to participate in the operation, with 16G of memory, 8G of video memory, and the CUDA version is 11.3. Before training, first configure the parameters of the model. The image size imgsz = [640, 640], the confidence threshold conf_thres = 0.25, the Iou threshold iou_thres = 0.45, the initial learning rate lr0 = 0.01, the weight decay coefficient weight_decay = 0.0005, and the number of iterations epochs = 200.
[0047] S4. Use the evaluation metrics Precision, Recall, mPA, and the number of parameters to test and evaluate the trained improved network model; The evaluation metrics are precision (Precision), recall (Recall), mean average precision mAP_0.5, and the number of parameters. P represents the proportion of correctly predicted positive samples to the actual number of positive samples. R represents the proportion of correctly predicted positive samples to the total number of predicted samples. mAP represents the comprehensive weighted average of the average precision of all category detections. Among them, the higher the values of precision (Precision), recall (Recall), and mean average precision mAP_0.5, the better the object detection. And the larger the value of the number of parameters, the slower the object detection speed. The specific calculation formulas are as follows: Among them, TP represents true positives, that is, positive samples predicted as positive by the model; FP represents false positives, that is, negative samples predicted as positive by the model; FN represents false negatives, that is, positive samples predicted as false by the model; K represents K categories, and AP i represents the average precision of the i-th category.
[0048] After the test and evaluation, the ablation experiment results of the model are as follows in the table: The results in the above table show that when improving the YOLOv5 network, while detecting the speed of the loss part, the accuracy rate has increased by 5.7%, the recall rate has increased by 7.0%, and mAP_0.5 has increased by 6.9%.
[0049] S5. Use the trained improved YOLOv5 to identify the door gap in the image. Based on the relative position information between the door hinge and the door gap and the predicted door hinge position information, calculate the position information of the door gap and draw the position of the door gap.
[0050] Prepare the image to be detected and the network model weight file best.pt obtained after training. Modify the model weight file and image path in detect.pt to complete the detection of the door hinge. Extract the center coordinates of the door hinge bounding box. Based on the known relative distance between the door hinge and the door gap, finally obtain the relevant coordinates of the door gap and draw the position of the door gap.
Claims
1. A method for identifying the gap of a train open car door, characterized in that: Including: S1: Collect the images of the side doors of open-top wagons in different states, preprocess the images, label the door hinges in the images to form the required dataset, and divide it into a training set and a test set; S2: Construct a recognition model for the door gaps of train open-top wagons; S3: Adjust the parameters of the recognition model for the door gaps of train open-top wagons, and set the size of the input image, the number of recognition categories, and the number of iterations; S4: Test and evaluate the trained recognition model for the door gaps of train open-top wagons; S5: Use the trained recognition model for the door gaps of train open-top wagons to recognize the door gaps in the images. Based on the relative position information between the door hinge and the door gap and the predicted door hinge position information, calculate the position information of the door gap and draw the position of the door gap.
2. The method for identifying the gap of the open door of a train according to claim 1, wherein The step S1 includes: S11: Data collection: Use an industrial camera to collect the overall images of the side doors of open-top wagons and the images of the door hinges at different angles; S12: Image preprocessing: Screen and remove incomplete and blurred images, and adjust the image size; S13: Image annotation: Label all the pictures of the door hinges at various angles in the preprocessed images; S14: Perform image enhancement; S15: Dataset division: After image enhancement, divide all the images into a training set and a test set according to a ratio.
3. The method for identifying the gap of the open door of a train according to claim 1, wherein The recognition model for the door gaps of train open-top wagons includes: Backbone network, and the backbone network includes: A replaced SCConv, four C3 layers alternately connected with the replaced SCConv, and an SPPF layer; the input of the backbone network is the dataset formed after the previous preprocessing. Then, multi-level features are extracted from the input data and transformed into three different-level feature maps; Neck network, and the neck network includes: Four one-dimensional convolutional layers Conv, four upsampling layers upsample, four C3 layers, four Concat layers, and a RepViT module; the inputs of the neck network are the three different-level feature maps output by the backbone network. Through the fusion of feature maps of different sizes in the neck network, the transmission of information of different sizes is enhanced, and the accuracy and robustness of object detection are improved. Finally, three new different-level feature maps are output to the head network; Head network, and the head network includes: The head network includes three convolutional layers and three detection heads. The inputs of the head network are the three different-level feature maps transmitted by the neck network; the convolutional layers adapt to different detection tasks by adjusting the number of channels and the kernel size. The detection heads implement object classification and regression tasks, retain the bounding box with the highest confidence, and output the processed bounding box and confidence score; The loss function in the detection head is the SIoU loss function.
4. The method for identifying the gap of the open door of a train according to claim 3, characterized in that, The SCConv module includes: Spatial Reconstruction Unit SRU, the Spatial Reconstruction Unit SRU takes the input feature X and outputs a spatially refined feature ; Channel Reconstruction Unit (CRU), which performs operations on the spatially refined features to obtain the channel-refined feature Y, reduce the redundancy between intermediate feature maps, and improve the feature representation of the convolutional neural network.
5. The train open door gap recognition method according to claim 4, characterized in that, The calculation process of the spatial reconstruction unit SRU is: Normalize the input feature X by subtracting the mean μ and dividing by the standard deviation σ; The weight values of the feature map reweighted by the relevant weights are mapped to the range (0, 1) through the sigmoid function and gated by a threshold. The weights above the threshold are set to 1 to obtain the information weight W1, while they are set to obtain the non-information weight W2; The weight values of the feature map reweighted by the relevant weights are mapped to the range (0, 1) through the sigmoid function and gated by a threshold. The weights above the threshold are set to 1 to obtain the information weight W1, while they are set to obtain the non-information weight W2; Multiply the input feature X by W1 and W2 respectively to obtain two weighted features: the feature rich in information and the feature with less information ; has rich and expressive spatial content, while has little information and is regarded as redundant.
6. The train open door gap recognition method according to claim 4, characterized in that The calculation process of the channel reconstruction unit CRU is: For a given spatial refinement feature , divide the channels of X W into two parts according to the segmentation ratio, where one part has αC channels and the other part has (1-α)C channels, where is the segmentation ratio; Use 1×1 convolutions to compress the channels of the feature map to improve computational efficiency, and use the squeezing ratio r to spatially refine the feature X W The upper part X UP and the lower part X low ; X UP Input into the up - conversion stage for X UP Use convolution operations GWC and PWC to extract high - level representative information and reduce computational costs; meanwhile, on the same X UP Use k×k GWC and 1×1 PWC operations, and then add the outputs to obtain the combined representative feature map Y1; X low Enter the bottom conversion stage for X low Use 1×1 PWC operations to generate a feature map with shallow hidden details, and connect the generated feature map and the reused feature X low to form the output Y2 of the bottom stage; Merge the output features Y1 and Y2 of the up-conversion stage and the down-conversion stage to obtain the channel-refined feature Y.
7. The method for identifying the gap of the train open door according to claim 6, characterized in that, The merging of the output features Y1 and Y2 of the up-conversion stage and the down-conversion stage includes: Collect the height, width, and convolution kernel of the output features Y1 and Y2 through global average pooling, and calculate the global spatial information , where m = 1, 2; Calculate the attention weights β1 and β2 of the output features Y1 and Y2 under the channel attention operation using the calculated global spatial information; Merge the upper feature Y1 and the lower feature Y2 in a channel-wise manner to obtain the channel-refined feature Y.
8. The method for identifying the gap of the open door of a train according to claim 3, characterized in that, The calculation process of the RepViT module is as follows: Divide the input image into several uniformly sized P×P blocks, which represent different parts of the image. Then, transform each image block into an embedded vector of a fixed dimension through a linear transformation. These vectors form a token sequence, and position encoding is added to the embedded vectors to obtain the positions of the token sequence.
9. The train open door gap recognition method according to claim 1, characterized in that The step S3 includes: Before training, configure the parameters of the model. The image size imgsz = [640, 640], the confidence threshold conf_thres = 0.25, the Iou threshold iou_thres = 0.45, the initial learning rate lr0 = 0.01, the weight decay coefficient weight_decay = 0.0005, and the number of iterations epochs = 200.