Chain embracing identification method, system and device and storage medium
By improving the target detection and instance segmentation model of the YOLOv8m network, the bending degree of railway chains was detected and evaluated, and accidents such as stoppage, conflict and derailment caused by the chain were solved, and timely warning and prevention of the risk of the brakes was achieved.
Patent Information
- Application Number
- CN202510484004.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-17
AI Technical Summary
During the railway operation, the vehicle was carrying the gate due to the locking of the brake shoe during the hump slip, causing accidents such as stoppage, conflict and derailment, which seriously affected the efficiency of railway operations and posed safety hazards.
The object detection model and instance segmentation model based on the improved YOLOv8m network are used to detect whether there is a chain in the image and segment and bending degree evaluation of the chain. If the bending degree is greater than the preset threshold, a brake risk warning will be issued.
It realizes efficient detection and positioning of railway chains, accurately evaluates the bending degree of the chain, and timely issues a risk warning of the brake holding risk, effectively avoiding railway accidents caused by chain holding.
Smart Images

Figure CN120071019A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of railway brake detection, and particularly relates to a chain embrace recognition method, system, device and storage medium. Background Art
[0002] In the hump shunting operation of railways, if the brake shoes hold the wheels to generate braking, resulting in the vehicle operating with brakes, the shunted vehicle may stop midway during the shunting process, and even cause collisions and derailments, seriously affecting the railway operation efficiency and even resulting in safety accidents.
[0003] Intelligent recognition of railway brakes is an important technology in the field of railway transportation safety. Braking means that during the disintegration and shunting operations of the hump, due to problems such as the vehicle to be shunted not exhausting air as required or the manual brake not releasing the brake as required, all the brake shoes of the vehicle still tightly adhere to the wheel tread. Braking can cause the vehicle to stop midway, leading to frequent accidents such as collisions and derailments. The braking used by railway freight cars during operation to stop and control speed is divided into two categories: one is air braking, and the other is manual braking. Air braking is to inject air into the air cylinder to push the piston, so that the piston rod extends to drive the brake shoes to hold the wheels, and the brake shoes rub against the wheels to perform railway braking; manual braking is to manually rotate the manual brake, so that the manual brake chain pulls the brake shoes to hold the wheels, and the brake shoes rub against the wheels to perform railway braking. Chain embrace means that during the shunting operation of the hump, when the vehicle is shunted, the manual brake is not released as required, resulting in the brake shoes holding the wheels and the vehicle operating with brakes.
[0004] In recent years, with the continuous emergence of new models in the field of deep learning, the vehicle intelligent detection technology based on computer vision has also developed rapidly. In the field of object detection, the mainstream algorithms are mainly divided into two categories: (1) Single-stage models: Such models simplify the object detection task into a regression and classification problem, and directly predict the category and bounding box of the object through the input image. Due to the simple structure, single-stage models have a high detection speed and are suitable for real-time detection scenarios and device operations with limited computing power. (2) Two-stage models: Different from single-stage models, two-stage models divide the detection task into two stages. The first stage generates candidate regions, and the second stage further classifies and regresses these candidate regions. By gradually refining the object features, this model performs better in small object detection and complex scenarios. However, the multi-stage computational characteristics make the model structure more complex, correspondingly increasing the computational cost.
[0005] Therefore, how to use deep learning to detect and warn the railway brake risk in real time and dynamically is worthy of research. Summary of the Invention
[0006] In view of the above problems, the present invention proposes a chain embrace recognition method, system, device and storage medium, in an attempt to solve or alleviate one or more of the above problems.
[0007] According to one aspect of the present invention, a chain hug recognition method is proposed, and the method includes: Input the image to be detected into a trained object detection model based on the first improved YOLOv8m network for detection to determine whether there is a chain in the image to be detected; wherein, the improvements of the first improved YOLOv8m network include: adding a spatial pyramid pooling module and a partial self-attention module to the original backbone network of YOLOv8m; adding a coordinate attention mechanism after the neck network and before the head network; If there is a chain in the image to be detected, extract the image of the chain detection frame area, and input the image of the chain detection frame area into a trained instance segmentation model based on the second improved YOLOv8m-Seg network for chain segmentation to obtain a chain mask image; wherein, the improvements of the second improved YOLOv8m-Seg network include: adding a spatial pyramid pooling module and a partial self-attention module to the original backbone network of YOLOv8m-Seg; adding a multi-scale sequence fusion module between the backbone network and the neck network; adding a feature fusion module between the neck network and the head network; for the chain mask image, use a parabola fitting algorithm to extract the center line of the chain, and evaluate the bending degree of the chain according to the ratio of the arc depth of the chain to the chord length of the chain; If the bending degree is greater than a preset bending threshold, a brake risk warning is issued.
[0008] Further, the step of inputting the image to be detected into a trained object detection model based on the first improved YOLOv8m network for detection to determine whether there is a chain in the image to be detected includes: Input the image to be detected into the backbone network for feature extraction, including: using multiple C2f feature extraction modules in the backbone network to extract local features of the image and perform feature fusion; subsequently, aggregate features of different scales through the spatial pyramid pooling module, including: processing features of different scales through a convolutional layer, batch normalization, and SiLU activation function to obtain a feature map with halved number of channels; sequentially passing through multiple max pooling layers to gradually reduce the size of the feature map; splicing the feature maps after pooling at different scales in the channel dimension; subsequently, through the partial self-attention module, evenly divide the feature map aggregated by the spatial pyramid pooling module into two parts, one part of the feature map enters the self-attention module to perform global information modeling through matrix operations between query vectors, key vectors, and value vectors; the other part of the feature map is fused with the output of the self-attention module through a skip connection; Use the neck network to fuse, enhance, and process the features extracted by the backbone network; Generate a feature map with enhanced position information using a coordinate attention mechanism, including: First, perform global average pooling on the features output by the neck network in the horizontal and vertical directions respectively to extract compressed features in different directions; Subsequently, these features are fused through a group of convolutions with shared weights; The fused features are re-divided into two parts in the horizontal and vertical directions, and attention matrices are generated through 1×1 convolutions and Sigmoid activation functions respectively; Finally, the features output by the neck network are weighted using the attention matrix through element-wise multiplication to obtain a feature map with enhanced position information; Use the head network to convert the feature map with enhanced position information into the final object detection result.
[0009] Furthermore, the dimensions of the query vector and the key vector in the partial self-attention module are half of the value vector; Batch normalization is used instead of layer normalization for normalization processing.
[0010] Furthermore, the step of inputting the chain detection box region image into the trained instance segmentation model based on the second improved YOLOv8m-Seg network for chain segmentation includes: Input the chain detection box region into the backbone network for feature extraction, including: Use multiple C2f feature extraction modules in the backbone network to extract local features of the image and perform feature fusion; Subsequently, aggregate features at different scales through a spatial pyramid pooling module, including: Process features at different scales through a convolutional layer, batch normalization, and SiLU activation function to obtain a feature map with halved number of channels; Pass through multiple max pooling layers in sequence to gradually reduce the size of the feature map; Concatenate the feature maps after pooling at different scales in the channel dimension; Subsequently, through the partial self-attention module, evenly divide the aggregated feature map into two parts, one part of the feature map enters the self-attention module to perform global information modeling through matrix operations between the query vector, key vector, and value vector; The other part of the feature map is fused with the output of the self-attention module through a skip connection; Use the multi-scale sequence fusion module to fuse the features extracted by the backbone network, including: Extract feature maps at levels P2 to P5 from the backbone network, and initially fuse low-level local detail information and high-level global semantic information through element-wise concatenation after size adjustment to generate a comprehensive description of the target features; Subsequently, perform channel dimensionality reduction through 1×1 convolutions and map the features to a higher-level feature space; Use RepVGG modules with convolutional kernels of 3 and 5 to extract edge details and context information of the target under different receptive fields; Use the neck network to fuse, enhance, and process the features extracted by the backbone network; The fused features extracted by the multi-scale sequence fusion module are fused with the feature map at the P3 level of the neck network using a feature fusion module, including: concatenating the two parts of the feature maps in the channel dimension; generating a set of feature weights through 3×3 convolution and the Sigmoid activation function; adjusting the features to be fused respectively according to the set of feature weights; and fusing the two adjusted parts of the features through element-wise addition. The head network is used to convert the feature map fused by the feature fusion module into a final instance segmentation result; wherein, the head network includes a segmentation head and a prediction head, the segmentation head is used to generate a high-resolution native mask of the target, and the prediction head is used to predict target attributes, and the target attributes include the position of the target box, class information, and mask coefficients related to segmentation.
[0011] Further, the formula for evaluating the degree of bending of the chain according to the ratio of the depth of the chain arc to the chord length of the chain is:
[0012] Wherein, is the curvature of the center line of the chain, represents the chord length of the center line of the chain, represents the maximum depth of the center line of the chain compared to the chord length.
[0013] According to another aspect of the present invention, a chain hug recognition system is proposed, and the system includes: A chain detection module configured to input an image to be detected into a trained object detection model based on the first improved YOLOv8m network for detection to determine whether there is a chain in the image to be detected; wherein, the improvements of the first improved YOLOv8m network include: adding a spatial pyramid pooling module and a partial self-attention module to the original backbone network of YOLOv8m; adding a coordinate attention mechanism after the neck network and before the head network; A chain segmentation module configured to, if there is a chain in the image to be detected, extract the image of the chain detection box area, input the image of the chain detection box area into a trained instance segmentation model based on the second improved YOLOv8m-Seg network for chain segmentation, and obtain a chain mask image; the improvements of the second improved YOLOv8m-Seg network include: adding a spatial pyramid pooling module and a partial self-attention module to the original backbone network of YOLOv8m-Seg; adding a multi-scale sequence fusion module between the backbone network and the neck network; adding a feature fusion module between the neck network and the head network; A chain hug recognition module configured to, for the chain mask image, calculate the center line of the chain using a parabola fitting algorithm and evaluate the degree of bending of the chain according to the ratio of the depth of the chain arc to the chord length of the chain; if the degree of bending is greater than a preset bending threshold, a brake risk warning is issued.
[0014] According to another aspect of the present invention, an electronic device is also provided, which includes a memory, a processor, and a computer program; wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the above-mentioned chain hug recognition method.
[0015] According to another aspect of the present invention, a computer-readable storage medium is also provided, and the storage medium stores a computer program; the computer program is executed by a processor to implement the above-mentioned chain hug recognition method.
[0016] The beneficial technical effects of the present invention are: The present invention provides a chain hug recognition method, system, device, and storage medium. The image to be detected is input into a trained object detection model based on the first improved YOLOv8m network for detection to determine whether there is a chain in the image to be detected; if there is a chain in the image to be detected, the image of the chain detection frame area is extracted, and the image of the chain detection frame area is input into a trained instance segmentation model based on the second improved YOLOv8m-Seg network for chain segmentation to obtain a chain mask image; for the chain mask image, a parabola fitting algorithm is used to calculate the center line of the chain, and the bending degree of the chain is evaluated according to the ratio of the arc depth of the chain to the chord length of the chain; if the bending degree is greater than a preset bending threshold, a brake risk warning is issued. Among them, a spatial pyramid pooling module and a partial self-attention module are introduced in the fourth stage of the original backbone network of the first improved YOLOv8m network; a coordinate attention mechanism is introduced in the third stage of the neck network; by adaptively adjusting the feature weights, the model can more accurately focus on the target boundary, thereby realizing the efficient detection and positioning of the chain; a spatial pyramid pooling module and a partial self-attention module are added to the original backbone network of the second improved YOLOv8m-Seg network; a multi-scale sequence fusion module is added between the backbone network and the neck network; a feature fusion module is added between the neck network and the head network. By effectively fusing the global context information and the position information, the analysis ability of the model for complex shapes at different scales is enhanced, thereby realizing the accurate segmentation of the target. Description of the Drawings
[0017] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, wherein: Figure 1 is a flowchart of a chain hug recognition method according to an embodiment of the present invention;
[0018] Figure 2It is another flowchart of a chain embrace recognition method according to an embodiment of the present invention;
[0019] Figure 3 It is a schematic structural diagram of the first improved YOLOv8m network in an embodiment of the present invention;
[0020] Figure 4 It is a schematic structural diagram of the second improved YOLOv8m-Seg network in an embodiment of the present invention; Figure 5 It is a schematic structural diagram of a chain embrace recognition system according to an embodiment of the present invention. Detailed implementation manners
[0021] Next, the principles and spirit of the present invention will be described with reference to several exemplary embodiments. It should be understood that these embodiments are given only to enable those skilled in the art to better understand and then implement the present invention, rather than limiting the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.
[0022] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, equipment, method, or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software. In this article, it should be understood that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0023] Chain embrace recognition needs to automatically identify and locate components such as the manual brake chain and brake cylinder chain at the bottom of different vehicle models, judge the morphological changes of the brake chain target, calculate the bending degree of the manual brake chain and brake cylinder chain, and determine the braking and holding vehicles by setting conditions and rules such as brake chain fitting scores and recognition areas, and combining the set thresholds for early warning. Considering that there are a large number of irrelevant regions in high-resolution images, and the task target has a long and strip-shaped morphological structure, which only accounts for a very small proportion of the total pixels of the image, resulting in the dilution of target features and making it difficult for a single model to accurately judge the task target. Therefore, the chain embrace recognition process is divided into two parts: target detection and instance segmentation. In the target detection part, the first improved YOLOv8m network structure is used to obtain the ROI region of the chain; in the instance segmentation part, based on the chain region located in the target detection stage, the second improved YOLOv8m-Seg network is used to complete the accurate segmentation of the brake chain.
[0024] Thus, an embodiment of the present invention proposes a chain embrace recognition method, as Figures 1 - 2 shown, the method includes: S1. Input the image to be detected into the trained object detection model based on the first improved YOLOv8m network for detection to determine whether there is a chain in the image to be detected; S2. If there is a chain in the image to be detected, extract the image of the chain detection box area, input the image of the chain detection box area into the trained instance segmentation model based on the second improved YOLOv8m-Seg network for chain segmentation, and obtain the chain mask image; S3. For the chain mask image, calculate the center line of the chain using the parabola fitting algorithm, and evaluate the bending degree of the chain according to the ratio of the arc depth of the chain to the chord length of the chain; S4. If the bending degree is greater than the preset bending threshold, issue a brake risk warning.
[0025] The method starts from S1. In S1, the image to be detected is input into the trained object detection model based on the first improved YOLOv8m network for detection to determine whether there is a chain in the image to be detected.
[0026] According to the embodiments of the present invention, the training process of the object detection model based on the first improved YOLOv8m network includes: S11. Obtain the training data set. The training and testing of the object detection model use a self-made data set, which is aggregated from the on-site collected data of multiple stations and divided into a training set and a testing set according to a ratio of 8:2. The data set is carefully labeled by a professional team according to strict labeling standards, ensuring the high quality of the data and the accuracy of the labeling. During the data collection process, the actual environmental differences of different stations are fully considered to enhance the adaptability of the model to diverse scenarios. For example, the collection tasks cover a variety of lighting conditions, including strong daylight, low light at night, and shadow areas; the diversity of equipment; and the complexity of the background, including situations such as raindrops and leaf debris occlusion.
[0027] S12. Preprocess the training data. To further improve the generalization ability of the model to diverse scenarios, data augmentation operations are performed on the training data. These augmentations include rotation, flipping, scaling, brightness adjustment, etc., enabling the model to adapt to complex actual environments.
[0028] S13. Input the preprocessed training data set into the object detection model based on the first improved YOLOv8m network for training to obtain the trained object detection model.
[0029] To balance the accuracy and inference speed of the model, YOLOv8m, which is more mature in technology among single-stage models, is selected as the basic framework of the model. As a member of the YOLO series, YOLOv8m is also composed of three parts of the network: the backbone, the neck, and the head. Specifically, the backbone network effectively captures local and global information of the image by stacking C2f feature extraction modules designed with multi-gradient flows; the neck network adopts the idea of a path aggregation network, and extracts effective information from feature maps at multiple levels through shrinking and expanding paths, thereby enhancing the detection ability for small objects and multi-scale targets; the head network is responsible for completing the final object detection task, including class prediction, bounding box regression, and object confidence estimation.
[0030] In complex scenarios, since the background and the target share similar texture or color features, the feature extraction process of the model is easily interfered with, leading to false detections and missed detections. To better meet the detection requirements of the task objectives, the object detection model proposed in the present invention has been specifically improved and optimized based on YOLOv8m. The specific architecture of the improved model is as Figure 3 shown. The structure of the improved model is still composed of three parts: the backbone, the neck, and the head. The improvements include: introducing a spatial pyramid pooling module and a partial self-attention module in the fourth stage of the backbone network; introducing a coordinate attention mechanism in the third stage of the neck network. The purpose of doing this is to further enhance the precise positioning ability of the model for the target boundary, making it more adaptable to the actual requirements of the system for detection accuracy while maintaining a high inference speed.
[0031] The backbone network is responsible for extracting multi-level and multi-scale general features from the input image, forming a feature map with high-level semantic information, and providing support for subsequent detection tasks. First, the C2f feature extraction module, as the basic unit of the network, is responsible for extracting local features at each stage and enhancing the expression of multi-scale information through feature fusion. This module adopts a CSP structure design and divides the input feature map into two branches through a 1×1 convolution: one branch is processed through multiple residual modules and the other branch directly performs a convolution operation. Subsequently, the outputs of the two branches are concatenated through a Concat operation, processed through batch normalization (BN) and the SiLU activation function, and finally the features are sorted through a convolution operation to obtain the final output. By stacking multiple C2f modules, the network can learn rich feature expressions at different levels and optimize the fusion of high-level semantic information and low-level detail information at different stages. The formula of the C2f feature extraction module can be expressed as follows:
[0032] Among them, is the input feature, is the output feature, Represents a 1×1 convolution, represents the channel splitting operation, is the residual module, represents the feature concatenation operation. is the intermediate feature of the module processing process and is the output obtained after N times of residual module processing.
[0033] Subsequently, different scale features are aggregated through the Spatial Pyramid Pooling - Fast (SPPF) module, enabling the network to better retain semantic information and generate richer feature representations when processing complex scenes. The SPPF module first processes the input feature map through a convolutional layer, batch normalization (BN), and SiLU activation function to obtain a feature map with halved number of channels. Then, the feature map sequentially passes through three 5×5 max - pooling layers, and the size of the feature map gradually decreases after each pooling. The output of each pooling layer serves as the input for the next pooling. This serial pooling process can effectively capture information at different scales and enhance the model's ability to capture details and global information in the image. Finally, by concatenating the feature maps pooled at different scales along the channel dimension, the SPPF module can integrate feature information from different scales to form a multi - scale feature representation. The formula of the spatial pyramid pooling module can be expressed as follows:
[0034] where, is the input feature, is the output feature of the module, represents the 5×5 max - pooling operation, represents the output obtained after k times of max - pooling processing.
[0035] When the model separates target and background features, it often inevitably introduces unnecessary background information to interfere with the target features, resulting in false detections and missed detections. To solve this problem, the common approach is to introduce an attention mechanism in the deep layers of the model to highlight the target features. However, most of the existing mainstream attention mechanisms are implemented through convolutional operators, and the inherent limitations of convolutional operators make it difficult to establish long - range dependencies between features, and this relationship is precisely the key to accurately locating the target. In contrast, the Transformer structure can efficiently capture global long - range dependencies with its self - attention mechanism. However, the computational complexity and memory occupancy of the Transformer are relatively high, often bringing a large time overhead in real - time inference tasks.
[0036] Therefore, a Partial Self-Attention (PSA) module is introduced into the backbone network. The design of the PSA module aims to effectively enhance the network's ability to model long-range dependencies while avoiding the computational overhead problem brought by the global self-attention mechanism. Specifically, the input feature map is evenly divided into two parts, thus controlling the complexity of self-attention calculation within a relatively low range. One part of the feature map is input into the self-attention module for global information modeling to capture long-range dependencies and enhance the context understanding of the target. Global information modeling is carried out through matrix operations between the query vector, key vector, and value vector to capture the long-range correlations between features; the other part of the feature map is fused with the output of the self-attention module through a skip connection. In addition, to further improve the inference efficiency, the PSA module optimizes the dimensions of the query vector and key vector in the self-attention mechanism, setting their dimensions to half of the value vector, thus reducing the computational amount. At the same time, the module uses batch normalization (BN) instead of layer normalization (LN) for normalization processing, thus improving the running speed and stability of the model. This module is placed after the fourth stage with the lowest resolution in the model, focusing on extracting and sorting out the key information in the high-level abstract features. Since the feature resolution in the fourth stage is relatively low, the quadratic complexity of self-attention calculation is significantly reduced, so the overall inference speed can still meet the real-time requirements. The formula of the partial self-attention module can be expressed as follows:
[0037] where, is the input feature, is the output feature, represents the Transformer module. Through the above operations, the long-range dependencies between features can be effectively captured without significantly increasing the computational overhead, so as to better process the semantic information in the deeper layers of the network.
[0038] The neck network is used to integrate the features extracted from the backbone network for fusion, enhancement, and processing to help the model better detect objects of different scales. Specifically, the neck network receives the features at the P3, P4, and P5 levels in the backbone network as inputs. First, it enhances the low-level features through a bottom-up information transmission path, and then fuses them with the high-level features to ensure that the network can obtain rich multi-level feature representations. Subsequently, the network spreads the detailed information to deeper layers of the network through a top-down information transmission path, enhancing the interaction between high-level semantic features and low-level detailed features, and further improving the model's sensitivity to details. Combining the bottom-up and top-down information flows ensures that the network still maintains high detection performance in complex scenarios, especially showing significant advantages in multi-scale object detection tasks.
[0039] To further improve the model's ability to extract target location information, a Coordinate Attention (CA) mechanism is introduced before the output part of the third stage of the neck network and before the head network to enhance the accurate positioning ability of the target boundary. The Coordinate Attention mechanism extracts compressed features in different directions by performing global average pooling in the horizontal and vertical directions respectively; subsequently, these features undergo information interaction and integration through a group of convolutional layers with shared weights to enhance the feature expression ability; the fused features are re-divided into two parts in the horizontal and vertical directions, and corresponding attention matrices are generated through 1×1 convolution and the Sigmoid activation function respectively; finally, the original features are weighted using the attention matrices through element-wise multiplication to enhance the position information in the feature map, effectively improving the model's positioning accuracy for the target boundary. The formula of the Coordinate Attention mechanism can be expressed as follows:
[0040] Among them, is the input feature, is the output feature, and represent the average pooling operations in the horizontal and vertical directions of the feature, is the Sigmoid activation function.
[0041] The head network is responsible for converting the feature map processed by the backbone and neck into the final object detection result. Specifically, the network consists of 3 detection heads, and each detection head can be divided into two parts, both of which are composed of 2 3×3 convolutions and one 1×1 convolution, respectively used to predict the coordinate regression information and the class confidence information.
[0042] Furthermore, to prevent the model from overfitting, regularization techniques such as weight decay and Dropout are introduced during the model training process to improve the model's robustness.
[0043] Then, the image to be detected is input into the trained object detection model based on the first improved YOLOv8m network for detection to determine whether there is a chain in the image to be detected. If the target is detected in the current image frame to be detected, the ROI region is extracted according to the output of the detection model and used as the input of the segmentation model.
[0044] Then, S2 is executed. In S2, if there is a chain in the image to be detected, the image of the chain detection box region is extracted, and the image of the chain detection box region is input into the trained instance segmentation model based on the second improved YOLOv8m-Seg network for chain segmentation to obtain the chain mask image.
[0045] According to the embodiments of the present invention, the process of training an instance segmentation model based on the second improved YOLOv8m-Seg network includes: S21. Obtain a training data set. Multiple data sets collected on-site are used for training and testing.
[0046] S22. Preprocess the training data. In order to further improve the generalization ability of the model for diverse scenarios, data augmentation operations are performed on the training data.
[0047] S23. Input the preprocessed training data set into the instance segmentation model based on the second improved YOLOv8m-Seg network for training to obtain a trained instance segmentation model.
[0048] Since the local area of slender objects is extremely narrow, occupying only a small number of pixels, it is extremely easy to be interfered by complex backgrounds and submerged during the feature transfer process. In addition, the aspect ratio of slender objects far exceeds that of conventional objects. Limited by the receptive field of the detector, it is difficult to generate complete object features, affecting the accurate segmentation of slender objects by the model. Aiming at the special shape of slender objects, the instance segmentation model proposed by the present invention is improved and optimized on the basis of the YOLOv8m-Seg network. The specific model structure is as Figure 4 shown. Specifically, a spatial pyramid pooling module and a partial self-attention module are added to the original backbone network of YOLOv8m-Seg; a multi-scale sequence fusion module is added between the backbone network and the neck network; a feature fusion module is added between the neck network and the head network. Among them, the multi-scale sequence fusion module aggregates feature maps of different scales in the backbone network to generate a global information representation containing more details, thereby enhancing the sensitivity of the network to target edge details and morphological changes; in the third stage of the neck network, the global context information and position information are effectively fused through the feature fusion module and passed down layer by layer, enabling the downstream segmentation head to use more accurate information to complete fine segmentation.
[0049] The backbone network undertakes the core function in the task. By extracting multi-level and multi-scale general features from the input image, it generates a feature map with high-level semantic information, providing a solid foundation for subsequent tasks. First, the C2f feature extraction module, as the basic unit of the network, is responsible for extracting local features at each stage and enhancing the expression of multi-scale information through feature fusion.
[0050] Subsequently, the Spatial Pyramid Pooling-Fast (SPPF) module is used to aggregate features from different scales to enhance the network's semantic understanding ability when dealing with complex scenes. The SPPF module first processes the input feature map through a convolutional layer, batch normalization (BN), and the SiLU activation function to obtain a feature map with half the number of channels. Then, the feature map sequentially passes through three 5×5 max-pooling layers, and the size of the feature map gradually decreases after each pooling. The output of each pooling layer serves as the input for the next pooling. This serial pooling process can effectively capture information at different scales and improve the model's ability to capture details and global information in the image. Finally, by concatenating the feature maps pooled at different scales along the channel dimension, the SPPF module can integrate feature information from different scales to form a multi-scale feature representation.
[0051] Subsequently, the backbone network introduces the Partial Self-Attention (PSA) module. The input feature map is evenly divided into two parts to control the complexity of self-attention calculation within a lower range. One part of the features is input into the self-attention module for global information modeling to capture long-range dependencies and enhance the context understanding of the target. Global information modeling is performed through matrix operations between the query vector, key vector, and value vector to capture the long-range correlations between features; the other part of the features is fused with the output of the self-attention module through a skip connection.
[0052] To further improve the segmentation accuracy, the Multi-scale sequence fusion (MSF) module is designed. By aggregating multi-scale features, it generates global context information to effectively compensate for the defects caused by the limited receptive field. Specifically, this module extracts feature maps at levels P2 to P5 from the backbone network. After size adjustment, it initially fuses low-level local detail information and high-level global semantic information through element-wise concatenation to generate a comprehensive description of the target features. To avoid information loss during excessive upsampling or downsampling of features, all levels of feature maps are adjusted to the same size as the P3-level feature map. Subsequently, 1×1 convolution is used for channel dimensionality reduction, and the features are mapped to a higher-level feature space to simplify the computational complexity and retain key information. On this basis, the RepVGG module with convolution kernels of 3 and 5 is used to extract the edge details and context information of the target under different receptive fields, further enhancing the expression ability of the global context and ensuring that the network can accurately capture the complete features of the target in a complex background. The formula for the multi-scale sequence fusion module can be expressed as follows:
[0053] Among them, represents the feature map of the i-th level, is the output feature of the module, represents the RepVGG module.
[0054] The neck network is composed of a bottom-up expansion path and a top-down contraction path, which is used to integrate the features extracted from the backbone network for fusion, enhancement, and processing to help the model better detect objects at different scales. The neck network receives the features at the P3, P4, and P5 levels in the backbone network as inputs. First, it enhances the low-level features through the bottom-up expansion path, and then fuses them with the high-level features to ensure that the network can obtain rich multi-level feature representations. Through the top-down contraction path, the neck network diffuses the detailed information to deeper levels of the network, enhancing the interaction between high-level semantic features and low-level detailed features, and further improving the model's sensitivity to details. Combining the bottom-up and top-down information flows ensures the detection performance of the neck network in the segmentation of slender objects.
[0055] Subsequently, the feature fusion module (Feature fusion module, FFM) fuses the feature maps from the multi-scale sequence fusion module and the P3-level feature map of the neck network expansion path to enhance the network's accurate segmentation ability. Specifically, the module first concatenates the two parts of the feature maps in the channel dimension to initially achieve information complementarity. Then, it generates a set of feature weights through 3×3 convolution and the Sigmoid activation function to reflect the importance of each feature. According to the set of feature weights, the features to be fused are adjusted respectively to enhance the response of key features and suppress redundant information. Finally, the two parts of the features are fused by element-wise addition to ensure that the final output has stronger feature expression ability, thereby improving the network's segmentation effect on slender objects. The formula of the feature fusion module is:
[0056] Among them, and are the features to be fused, is the output feature, is the Sigmoid activation function.
[0057] The head network is responsible for converting the feature map fused by the feature fusion module into the final instance segmentation result. The head network consists of a segmentation head and three prediction heads in total. The main task of the segmentation head is to generate a high-resolution native mask of the target, which is composed of multiple convolutional layers. By gradually extracting and restoring spatial information, accurate segmentation results are generated. First, a 3×3 convolution is used to process the input feature map to extract spatial information and enhance the detailed expressiveness of the features. Subsequently, the feature map is upsampled through transposed convolution to restore the spatial resolution of the image, thereby retaining the fine-grained information of the target. Finally, a 1×1 convolutional layer is used to reduce the dimension of the channels to generate a high-resolution target mask. The prediction heads are designed for feature maps at different levels and are mainly used to predict the attributes of the target, including the position of the target box, class information, and mask coefficients related to segmentation. First, the perception ability of the target area is enhanced through 3×3 convolution, and then the target class classification, bounding box regression, and mask coefficient prediction are completed through 1×1 convolution. Finally, according to the outputs of the segmentation head and the prediction heads, the segmentation mask of the target is generated.
[0058] Then, S3 is executed. In S3, a parabolic fitting algorithm is used to calculate the center line of the chain for the chain mask image, and the bending degree of the chain is evaluated based on the ratio of the arc depth of the chain to the chord length of the chain.
[0059] According to the embodiments of the present invention, the segmentation result will be approximately fitted by a parabola, and the curvature of the brake chain will be calculated based on the ratio of the arc depth of the brake chain to the chord length of the chain. By analyzing the segmentation result, the curvature calculated from the segmentation result is compared with a preset threshold, and the curvature change is used to identify potential risk situations. All segmentation images under the same regression box are regarded as the same target to avoid the influence of occlusion. To accurately obtain the geometric characteristics of the target, a parabolic fitting algorithm based on the least squares method is used to obtain the center line of the chain: . Subsequently, a mathematical derivation method is used to calculate the curvature of the center line of the chain to quantitatively reflect the deformation degree of the chain. The specific formula is:
[0060] Among them, is the curvature of the center line of the chain, represents the chord length of the center line of the chain, represents the maximum depth of the center line of the chain compared to the chord length.
[0061] Then, S4 is executed. In S4, if the bending degree is greater than the preset bending threshold, a brake risk warning is issued.
[0062] The technical effect of the present invention is further verified through experiments.
[0063] Build a chain embrace recognition system, with the acquisition device preset on the side of the train. When the magnetic steel detects the train signal, the system takes the continuously acquired video frame data as input. Subsequently, the video frame data is processed through two parts: the main program and the intelligent detection module. During the processing, the main program receives the curvature information of the brake chain in the image transmitted back by the intelligent detection module. The main program further calculates the proportion of all curvature information that is less than 0.035. If this proportion exceeds 50%, it is determined that there is a risk of brake embrace, and a warning prompt is issued in a timely manner.
[0064] Multiple on-site collected datasets are used for training and testing. Among them, the chain detection dataset contains 16,051 images, and the chain segmentation dataset contains 14,832 images. To ensure the effectiveness of the model in practical applications, these datasets are strictly labeled and screened, covering targets under different scenarios and lighting conditions. Both the detection and segmentation models are built based on the PyTorch deep learning framework. The batch size is set to 16, and the SGD optimizer with an initial learning rate of 0.01 and a momentum of 0.937 is used, and a total of 300 epochs are trained. To ensure the stability and convergence speed of the training process, a learning rate scheduling strategy is adopted, reducing the learning rate by 10 times every 100 epochs to achieve a balance between model exploration and convergence.
[0065] To achieve the optimal detection accuracy, detailed adjustments and optimizations are made for data augmentation. The final data augmentation parameters of the detection model and the segmentation model are shown in Table 1 and Table 2 respectively.
[0066] Table 1 Data Augmentation Parameters of the Detection Model
[0067] Table 2 Data Augmentation Parameters of the Segmentation Model
[0068] To verify the advancement of the segmentation model, the method proposed in the present invention is comprehensively compared with various mainstream methods in the field, including Mask R-CNN, YOLACT, YOLOv5m-Seg, and YOLOv8m-Seg, etc. To ensure the fairness and authority of the comparison, the public codes of these methods are used to reproduce their network structures, and the models are trained and evaluated in the same training environment. The hyperparameter settings, datasets, and evaluation metrics are strictly unified in the experiment to ensure the comparability of the results. As shown in Table 3, the proposed segmentation model is superior to other instance segmentation methods in all evaluation metrics, fully demonstrating its advantages in instance segmentation tasks.
[0069] Table 3 Comparison of the Segmentation Model of the Present Invention with Other Instance Segmentation Methods
[0070] To verify the effectiveness of the introduced modules, ablation experiments were conducted. The experiments used YOLOv8m-Seg as the baseline model, and on this basis, a multi-scale sequence fusion module and a feature fusion module were successively added to verify the contribution and effectiveness of each module. The experimental results are shown in Table 4. Among them, the multi-scale sequence fusion module significantly enriches the diversity of feature expressions by modeling and fusing information at different scales, enabling the model to have higher recognition ability and robustness when dealing with complex scenes or diverse targets. While the feature fusion module effectively fuses features from different sources and passes key information downward through the shrinking path, further enhancing the transmission and representation ability of semantic information, thus optimizing the performance of the model in the segmentation task. These results indicate that the design of each module has practical value, and their combination can significantly improve the overall model performance, verifying the effectiveness and rationality of the method of the present invention.
[0071] Table 4 Experimental Results
[0072] To improve the inference speed, the brake recognition model was converted from PyTorch to TensorRT to make full use of hardware acceleration. The inference GPU deployed on-site is NVIDIA GeForce RTX 4070, which supports FP16 and FP32 calculations, and both provide a computing performance of 29.15 TFLOPS. Through TensorRT optimization, the inference speed of the model has been significantly improved, and the inference time for a single image (including pre- and post-processing) is about 25 ms. To further improve the processing efficiency, multi-threaded inference was adopted during the deployment process to enhance the concurrent processing ability. The optimized deployed brake recognition system can process image data captured by multiple cameras in real-time at a speed of 50 frames per second during on-site actual operation, meeting the requirements of real-time monitoring and detection.
[0073] After comprehensive testing, the brake recognition system demonstrated excellent performance, with its detection accuracy exceeding 95%. Whether in complex environments or under variable conditions, the system can efficiently and accurately identify and detect brake risks, providing a strong guarantee for the safe operation of trains.
[0074] The present invention takes YOLOv8m as the core algorithm and makes customized improvements according to task requirements. For object detection, a spatial pyramid pooling module and some self-attention modules are added to the original backbone network of YOLOv8m; a coordinate attention mechanism is added after the neck network and before the head network to meet the high requirements for precise positioning in the task. By adaptively adjusting the feature weights, the model can more accurately focus on the target boundary, thus achieving efficient detection and positioning of the chain. For instance segmentation, for the target of the manual brake chain of a manpower brake which has a special shape, a spatial pyramid pooling module and some self-attention modules are added to the original backbone network of YOLOv8m-Seg; a multi-scale sequence fusion module is added between the backbone network and the neck network; a feature fusion module is added between the neck network and the head network. By effectively fusing the global context information and the position information, the analysis ability of the model for complex shapes at different scales is enhanced, thus achieving precise segmentation of the target.
[0075] Another embodiment of the present invention provides a chain hugging recognition system, as Figure 5 shown. The system includes: A chain detection module 510 configured to input an image to be detected into a trained object detection model based on the first improved YOLOv8m network for detection to determine whether there is a chain in the image to be detected; A chain segmentation module 520 configured to, if there is a chain in the image to be detected, extract the image of the chain detection box area, and input the image of the chain detection box area into a trained instance segmentation model based on the second improved YOLOv8m-Seg network for chain segmentation to obtain a chain mask image; A chain hugging recognition module 530 configured to, for the chain mask image, calculate the center line of the chain by using a parabola fitting algorithm, and evaluate the bending degree of the chain according to the ratio of the arc depth of the chain to the chord length of the chain; if the bending degree is greater than a preset bending threshold, a brake risk warning is issued.
[0076] For the parts not described in detail in a chain hugging recognition system according to an embodiment of the present invention, please refer to the above specific description of the method embodiment.
[0077] The method of the present invention can be executed in an electronic device. The electronic device can be any device with storage and computing capabilities, which can be implemented as, for example, a server, a workstation, etc., or can be implemented as a personal configured computer such as a desktop computer, a notebook computer, or can be implemented as a terminal device such as a mobile phone, a tablet computer, a smart wearable device, an Internet of Things device, etc., but is not limited thereto.
[0078] An electronic device may include: a processor, a memory, an input / output interface, a communication interface, and a bus. Among them, the processor, the memory, the input / output interface, and the communication interface are communicatively connected to each other inside the electronic device through the bus. The processor may be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification. The memory may be implemented in the form of ROM, RAM, a static storage device, a dynamic storage device, etc. The memory may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory and are called and executed by the processor. The input / output interface is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the electronic device or may be externally connected to the electronic device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc. The communication interface is used to connect to a communication module to implement communication interaction between this electronic device and other devices. Among them, the communication module may implement communication in a wired manner or in a wireless manner. The bus includes a path for transmitting information between various components of the electronic device.
[0079] An embodiment of the present invention further provides a non-transitory readable storage medium that stores instructions for causing the electronic device to execute the method according to the embodiments of the present invention. The readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of the readable storage medium include, but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage, etc.
[0080] It should be noted that the terms used in this invention are only for describing specific embodiments and do not limit the scope of this application. As shown in the specification of this invention, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method or device comprising the said element.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A chain hugging identification method, characterized in that: include: Input the image to be detected into the trained target detection model based on the first improved YOLOv8m network for detection to determine whether there is a chain in the image to be detected; wherein the improvements of the first improved YOLOv8m network include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m backbone network; adding a coordinate attention mechanism after the neck network and before the head network; If there are chains in the image to be detected, extract the chain detection box area image, input the chain detection box area image into the trained instance segmentation model based on the second improved YOLOv8m-Seg network for chain segmentation, and obtain a chain mask image; wherein the improvements of the second improved YOLOv8m-Seg network include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m-Seg backbone network; adding a multi-scale sequence fusion module between the backbone network and the neck network; adding a feature fusion module between the neck network and the head network; For the chain mask image, a parabola fitting algorithm is used to extract the center line of the chain, and the curvature of the chain is evaluated according to the ratio of the chain arc depth to the chain chord length; If the bending degree is greater than the preset bending threshold, a brake risk warning is issued.
2. A chain hugging identification method according to claim 1, characterized in that: The step of inputting the image to be detected into the trained target detection model based on the first improved YOLOv8m network for detection to determine whether there is a chain in the image to be detected includes: The image to be detected is input into the backbone network for feature extraction, including: using multiple C2f feature extraction modules in the backbone network to extract local features of the image and perform feature fusion; then, the features of different scales are aggregated through the spatial pyramid pooling module, including: processing the features of different scales through convolutional layers, batch normalization and SiLU activation functions to obtain feature maps with half the number of channels; passing through multiple maximum pooling layers in sequence to gradually reduce the size of the feature map; the feature maps after pooling at different scales are spliced in the channel dimension; then, through the partial self-attention module, the feature maps aggregated by the spatial pyramid pooling module are evenly divided into two parts, one part of the feature map enters the self-attention module to perform global information modeling through matrix operations between query vectors, key vectors and value vectors; the other part of the feature map is fused with the output of the self-attention module through jump connections; The neck network is used to fuse, enhance and process the features extracted by the backbone network; The coordinate attention mechanism is used to generate a feature map with enhanced position information, including: first, performing global average pooling on the features output by the neck network in the horizontal direction and the vertical direction respectively, and extracting compressed features in different directions; then, these features are fused through a set of convolutions with shared weights; the fused features are re-divided into two parts in the horizontal direction and the vertical direction, and the attention matrix is generated through 1×1 convolution and Sigmoid activation function respectively; finally, the features output by the neck network are weighted by the attention matrix through element-wise multiplication to obtain a feature map with enhanced position information; The head network is used to transform the feature map enhanced with position information into the final object detection result.
3. A chain hugging identification method according to claim 2, characterized in that: The dimension of the query vector and the key vector in the partial self-attention module is half of the value vector; batch normalization is used instead of layer normalization for standardization.
4. A chain hugging identification method according to claim 1, characterized in that: The step of inputting the chain detection frame area image into a trained instance segmentation model based on the second improved YOLOv8m-Seg network to perform chain segmentation includes: The chain detection frame area image is input into the backbone network for feature extraction, including: using multiple C2f feature extraction modules in the backbone network to extract local features of the image and perform feature fusion; then, the features of different scales are aggregated through the spatial pyramid pooling module, including: processing the features of different scales through convolutional layers, batch normalization and SiLU activation functions to obtain feature maps with half the number of channels; passing through multiple maximum pooling layers in sequence to gradually reduce the size of the feature map; the feature maps after pooling at different scales are spliced in the channel dimension; then, through the partial self-attention module, the aggregated feature map is evenly divided into two parts, one part of the feature map enters the self-attention module to perform global information modeling through matrix operations between query vectors, key vectors and value vectors; the other part of the feature map is fused with the output of the self-attention module through jump connections; The features extracted by the backbone network are fused using a multi-scale sequence fusion module, including: extracting feature maps of levels P2 to P5 from the backbone network, initially fusing low-level local detail information and high-level global semantic information through element splicing after resizing, and generating a comprehensive description of the target features; then, performing channel dimension reduction through 1×1 convolution and mapping the features to a higher-level feature space; using the RepVGG module with convolution kernels of 3 and 5, extracting edge details and contextual information of the target under different receptive fields; The neck network is used to fuse, enhance and process the features extracted by the backbone network; The fusion features extracted by the multi-scale sequence fusion module are fused with the feature map of the neck network P3 level by using the feature fusion module, including: splicing the two parts of the feature map in the channel dimension; generating a feature weight set by 3×3 convolution and Sigmoid activation function; adjusting the features to be fused respectively according to the feature weight set; fusing the adjusted two parts of the features by element addition; The head network is used to convert the feature map fused by the feature fusion module into the final instance segmentation result; wherein the head network includes a segmentation head and a prediction head, the segmentation head is used to generate a high-resolution native mask of the target, and the prediction head is used to predict the target attributes, the target attributes including the target box position, category information and mask coefficients related to segmentation.
5. A chain hugging identification method according to claim 1, characterized in that: The formula for evaluating the degree of chain bending based on the ratio of the chain arc depth to the chain chord length is: in, is the curvature of the chain centerline, represents the chord length of the center line of the chain, Indicates the maximum depth of the centerline of the chain compared to the chord length.
6. A chain hugging identification system, characterized in that: include: A chain detection module is configured to input the image to be detected into a trained target detection model based on the first improved YOLOv8m network for detection to determine whether there is a chain in the image to be detected; wherein the improvements of the first improved YOLOv8m network include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m backbone network; adding a coordinate attention mechanism after the neck network and before the head network; A chain segmentation module is configured to extract a chain detection frame area image if there is a chain in the image to be detected, and input the chain detection frame area image into a trained instance segmentation model based on the second improved YOLOv8m-Seg network to perform chain segmentation to obtain a chain mask image; wherein the improvements of the second improved YOLOv8m-Seg network include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m-Seg backbone network; adding a multi-scale sequence fusion module between the backbone network and the neck network; and adding a feature fusion module between the neck network and the head network; The chain holding identification module is configured to calculate the chain centerline using a parabola fitting algorithm for the chain mask image, and evaluate the degree of bending of the chain according to the ratio of the chain arc depth to the chain chord length; if the degree of bending is greater than a preset bending threshold, a brake risk warning is issued.
7. An electronic device, characterized in that: include: A memory, a processor and a computer program; wherein the computer program is stored in the memory and is configured to be executed by the processor to implement the chain holding identification method described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that The storage medium stores a computer program; the computer program is executed by a processor to implement the chain holding identification method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Train brake chain state recognition method and device, equipment and storage medium
CN114758150A
Optical fiber surface defect detection method and device
CN116071294A
Light-weight automobile tail lamp speech real-time identification method
CN116129400A
Rail surface defect detection method and device, electronic equipment and storage medium
CN117952946A
Segmentation method for rail defect detection based on improved YOLOv10 and SETR
CN119671959A